Best Practices for Reproducible Research Data Work
Reproducible Pipeline
Anyone should reproduce all results by running master and changing one path only.
project/
├── data/
│ ├── raw/
│ └── clean/
├── code/
│ ├── 01_clean_data.*
│ ├── 02_construct_vars.*
│ └── 03_analysis.*
├── outputs/
│ ├── tables/
│ └── figures/
├── master.*
├── README.md
Reproducible Workflow in R
In R, reproducibility is commonly achieved through scripted pipelines and dynamic documents (e.g., Quarto or R Markdown). A typical workflow is coordinated through a master script:
Purpose: Reproduce all analysis outputs end to end
source(“01_clean_data.R”) source(“02_construct_variables.R”) source(“03_analysis.R”)
Each script performs a single, well-defined task. For example, a data‑cleaning script:
01_clean_data.R
library(dplyr)
raw <- read.csv(“data/raw/survey.csv”)
clean <- raw |> filter(!is.na(outcome)) |> mutate( treatment = as.integer(treatment == “yes”), age_centered = age - mean(age, na.rm = TRUE) )
write.csv(clean, “data/clean/survey_clean.csv”, row.names = FALSE)
02_construct_vars.R
df <- read.csv(“data/clean/survey_clean.csv”)
df <- df |> mutate( treat_bin = ifelse(treatment == “yes”, 1, 0), age_centered = age - mean(age, na.rm = TRUE) )
write.csv(df, “data/clean/survey_model.csv”, row.names = FALSE)
03_analysis.R
df <- read.csv(“data/clean/survey_model.csv”)
model <- lm(outcome ~ treat_bin + age_centered, data = df)
sink(“outputs/tables/regression.txt”) summary(model) sink()
Final tables and figures are generated directly in a Quarto document, ensuring that any code or data change automatically updates the paper.
Reproducible Workflow in Python
In Python, reproducibility is often structured using modular scripts or notebooks, combined with environment management (e.g., requirements.txt or conda environments).
Reproduce all results
from clean_data import clean_data from construct_vars import construct_vars from analysis import run_model
df = clean_data(“data/raw/survey.csv”) df = construct_vars(df) run_model(df)
The data‑cleaning module:
import pandas as pd
def clean_survey(path): df = pd.read_csv(path) df = df.dropna(subset=[“outcome”]) df[“treatment”] = (df[“treatment”] == “yes”).astype(int) df[“age_centered”] = df[“age”] - df[“age”].mean() return df
construct_vars.py
def construct_vars(df): df[“treat_bin”] = (df[“treatment”] == “yes”).astype(int) df[“age_centered”] = df[“age”] - df[“age”].mean() return df
analysis.py
import statsmodels.formula.api as smf
def run_model(df): model = smf.ols(“outcome ~ treat_bin + age_centered”, data=df).fit() with open(“outputs/tables/regression.txt”, “w”) as f: f.write(model.summary().as_text())
By running python main.py, another researcher can regenerate all outputs with minimal configuration, supporting computational reproducibility and reuse.
Reproducible Workflow in Stata
In Stata, reproducibility relies heavily on master do‑files that control execution order, file paths, and settings.
- MASTER.do
- Purpose: Reproduce all data work and outputs
version 17 clear all set more off
global root “C:/project”
do “\({root}/code/01_clean_data.do" do "\){root}/code/02_construct_variables.do” do “${root}/code/03_analysis.do”
A cleaning script might include explicit documentation of inputs and outputs:
- 01_clean_data.do
- INPUT: raw/survey.dta
- OUTPUT: clean/survey_clean.dta
use “${root}/data/raw/survey.dta”, clear drop if missing(outcome)
gen treatment = treatment_var == “yes” egen age_mean = mean(age) gen age_centered = age - age_mean rename treat treatment
save “${root}/data/clean/survey_clean.dta”, replace
02_construct_vars.do
use “${ROOT}/data/clean/survey_clean.dta”, clear
gen treat_bin = treatment == “yes” egen age_mean = mean(age) gen age_centered = age - age_mean
save “${ROOT}/data/clean/survey_model.dta”, replace
03_analysis.do
use “${ROOT}/data/clean/survey_model.dta”, clear
reg outcome treat_bin age_centered esttab using “${ROOT}/outputs/tables/regression.tex”, replace
Because all steps are executed from the master file, the entire workflow can be rerun by modifying only the root directory path.
| Principle | Implementation |
|---|---|
| Single entry point | MASTER.do, master.R, main.py |
| Explicit data flow | Raw → Clean → Model |
| No manual edits | All outputs generated automatically |
| Portable | One configurable path or environment setting |
| Reviewable | Scripted, readable, version‑controlled code |