Best Practices for Reproducible Research Data Work

Author

Cecilia Regueira

Reproducible Pipeline

Anyone should reproduce all results by running master and changing one path only.

project/
├── data/
│   ├── raw/
│   └── clean/
├── code/
│   ├── 01_clean_data.*
│   ├── 02_construct_vars.*
│   └── 03_analysis.*
├── outputs/
│   ├── tables/
│   └── figures/
├── master.*
├── README.md

Reproducible Workflow in R

In R, reproducibility is commonly achieved through scripted pipelines and dynamic documents (e.g., Quarto or R Markdown). A typical workflow is coordinated through a master script:

Notemaster.R

Purpose: Reproduce all analysis outputs end to end

source(“01_clean_data.R”) source(“02_construct_variables.R”) source(“03_analysis.R”)

Each script performs a single, well-defined task. For example, a data‑cleaning script:

01_clean_data.R

Note01_clean_data.R

library(dplyr)

raw <- read.csv(“data/raw/survey.csv”)

clean <- raw |> filter(!is.na(outcome)) |> mutate( treatment = as.integer(treatment == “yes”), age_centered = age - mean(age, na.rm = TRUE) )

write.csv(clean, “data/clean/survey_clean.csv”, row.names = FALSE)

02_construct_vars.R

Note

df <- read.csv(“data/clean/survey_clean.csv”)

df <- df |> mutate( treat_bin = ifelse(treatment == “yes”, 1, 0), age_centered = age - mean(age, na.rm = TRUE) )

write.csv(df, “data/clean/survey_model.csv”, row.names = FALSE)

03_analysis.R

Note

df <- read.csv(“data/clean/survey_model.csv”)

model <- lm(outcome ~ treat_bin + age_centered, data = df)

sink(“outputs/tables/regression.txt”) summary(model) sink()

Final tables and figures are generated directly in a Quarto document, ensuring that any code or data change automatically updates the paper.

Reproducible Workflow in Python

In Python, reproducibility is often structured using modular scripts or notebooks, combined with environment management (e.g., requirements.txt or conda environments).

Notemain.py

Reproduce all results

from clean_data import clean_data from construct_vars import construct_vars from analysis import run_model

df = clean_data(“data/raw/survey.csv”) df = construct_vars(df) run_model(df)

The data‑cleaning module:

Noteclean_data.py

import pandas as pd

def clean_survey(path): df = pd.read_csv(path) df = df.dropna(subset=[“outcome”]) df[“treatment”] = (df[“treatment”] == “yes”).astype(int) df[“age_centered”] = df[“age”] - df[“age”].mean() return df

construct_vars.py

Note

def construct_vars(df): df[“treat_bin”] = (df[“treatment”] == “yes”).astype(int) df[“age_centered”] = df[“age”] - df[“age”].mean() return df

analysis.py

Note

import statsmodels.formula.api as smf

def run_model(df): model = smf.ols(“outcome ~ treat_bin + age_centered”, data=df).fit() with open(“outputs/tables/regression.txt”, “w”) as f: f.write(model.summary().as_text())

By running python main.py, another researcher can regenerate all outputs with minimal configuration, supporting computational reproducibility and reuse.

Reproducible Workflow in Stata

In Stata, reproducibility relies heavily on master do‑files that control execution order, file paths, and settings.

Note
  • MASTER.do
  • Purpose: Reproduce all data work and outputs

version 17 clear all set more off

global root “C:/project”

do “\({root}/code/01_clean_data.do" do "\){root}/code/02_construct_variables.do” do “${root}/code/03_analysis.do”

A cleaning script might include explicit documentation of inputs and outputs:

Note
  • 01_clean_data.do
  • INPUT: raw/survey.dta
  • OUTPUT: clean/survey_clean.dta

use “${root}/data/raw/survey.dta”, clear drop if missing(outcome)

gen treatment = treatment_var == “yes” egen age_mean = mean(age) gen age_centered = age - age_mean rename treat treatment

save “${root}/data/clean/survey_clean.dta”, replace

02_construct_vars.do

Note

use “${ROOT}/data/clean/survey_clean.dta”, clear

gen treat_bin = treatment == “yes” egen age_mean = mean(age) gen age_centered = age - age_mean

save “${ROOT}/data/clean/survey_model.dta”, replace

03_analysis.do

Note

use “${ROOT}/data/clean/survey_model.dta”, clear

reg outcome treat_bin age_centered esttab using “${ROOT}/outputs/tables/regression.tex”, replace

Because all steps are executed from the master file, the entire workflow can be rerun by modifying only the root directory path.

Principle Implementation
Single entry point MASTER.do, master.R, main.py
Explicit data flow Raw → Clean → Model
No manual edits All outputs generated automatically
Portable One configurable path or environment setting
Reviewable Scripted, readable, version‑controlled code