Best Practices for Coding

Author

Cecilia Regueira

Executive Summary

Note

After more than ten years of hands‑on coding and applied research experience, this article presents a rigorous and comprehensive set of best practices for reproducible, transparent, and collaborative data work.

This summary synthesizes and integrates key principles from Clean Code: A Handbook of Agile Software Craftsmanship by Robert C. Martin—which emphasizes readable, maintainable, and professional code—and Development Research in Practice: The DIME Analytics Data Handbook by the World Bank, which formalizes end‑to‑end standards for reproducible, ethical, and collaborative research workflows.

Best practices for code and analysis include:

1. Reproducibility and a Clean Code

Reproducibility is achieved by structuring projects so that all results can be regenerated from raw data using clearly defined master scripts. Computational reproducibility is treated as a core research output, enabling independent verification and reuse of methods across contexts and over time.

Robert C. Martin emphasizes that code is a primary communication artifact, not merely an execution mechanism. Core contributions of Clean Code include:

  • Readability and Intentionality – Code should be easy to read and self‑explanatory
  • Single Responsibility and Modularity – Each function or script should do one thing well
  • Meaningful Comments and Documentation – Comments explain why decisions were made, not just what the code does
  • Refactoring as a Continuous Practice – Code quality improves through regular, incremental cleanup
✅ Click to read more

Reproducibility ensures that research findings can be independently verified and consistently regenerated using the same data and code. In practice, this requires structuring analytical workflows so that all results—tables, figures, and statistics—are produced directly from raw or well‑documented intermediate data using executable code. Reproducibility is therefore not an afterthought, but a design principle applied throughout the project lifecycle.

A core practice is the use of a master script, which defines software versions, directory paths, and execution order, and runs all data processing and analysis steps end to end. By changing only a single file path or configuration parameter, another researcher should be able to rerun the complete workflow on a different machine and reproduce identical results.

Well‑written reproducible code is modular, readable, and well documented. Tasks are split into separate scripts (for example, 01_clean_data, 02_construct_variables, 03_analysis), each with clear headers describing inputs, outputs, and purpose. Version control systems such as Git track changes over time, enabling transparent review of how results evolve.

Reproducibility is further strengthened through dynamic documents, such as Quarto, R Markdown, or Jupyter notebooks, which integrate code, outputs, and narrative text in a single source file. Any change in data or code is immediately reflected in the final report.

Importantly, computational reproducibility is treated as a research output in its own right. Reproducibility packages typically include executable code, minimal working data (or detailed metadata when access is restricted), documentation, and a README file with clear instructions.

Reproducible research requires clearly structured, documented, and executable code. While the principles of reproducibility are consistent across disciplines, their implementation varies by programming environment.

📘 View Code Examples

2. Transparency

Transparency requires thorough documentation of data, code, and research decisions, with public release whenever legally and ethically feasible. Using trusted external repositories—such as GitHub, OSF, or institutional data catalogs—ensures long‑term access, accountability, and opportunities for reuse.

Where transparency requires openness, Clean Code ensures that what is open is also understandable.

3. Ethical and Secure Data Handling

Ethical and secure data handling protects research participants by enforcing informed consent, separating and encrypting sensitive information, and complying with applicable legal and institutional standards. Careful data governance throughout the project lifecycle ensures privacy, minimizes risk, and maintains trust while still enabling scientific transparency.

In higher education settings, researchers and analysts must comply with the Family Educational Rights and Privacy Act (FERPA), which governs access to student education records. Compliance requires limiting access to personally identifiable information, applying appropriate de‑identification practices, and ensuring data are used solely for legitimate educational or research purposes.

4.Data Work as a Collaborative Process

Data work is inherently collaborative and depends on shared standards, clear communication, and coordinated workflows across research teams. Treating data work as a collective responsibility—rather than an individual or ad hoc activity—reduces errors, improves efficiency, and strengthens institutional memory.

Effective collaborative data work relies on standardization and transparency. Common folder structures, naming conventions, coding styles, and version control systems allow team members to contribute with minimal friction. Documentation of decisions and data transformations ensures continuity across staff transitions and supports meaningful peer review.

Equally important is the use of shared tools and workflows, such as version‑controlled repositories, task‑tracking systems, and dynamic documents that integrate code and outputs. Regular code review and clear role definition foster accountability while encouraging collective learning.

Where collaborative data work requires shared standards, Agile software development provides the team practices that sustain those standards over time.

✅ Click to read about Agile practices

Agile Software Development emphasizes team‑based workflows, rapid feedback, and adaptability:

  • Iterative development supports early validation of assumptions and outputs
  • Shared ownership of code reduces risk from staff turnover
  • Continuous integration and testing prevent silent errors
  • Professional responsibility emphasizes correctness, clarity, and maintainability

Together, these principles reinforce the view that reproducibility is a continuous process and that collaboration must be intentionally designed and maintained.