I maintain 176+ GitHub repositories.
Here's my confession: my Git hygiene for ML was terrible for the first two years.
Code versioned? Yes. Data versioned? No. Model weights tracked? Sometimes. Experiment configs? Scattered across Slack messages and sticky notes. Reproducibility? "Let me try to remember what I changed..."
Sound familiar?
ML needs version control for EVERYTHING, not just code. Data (use DVC). Models (MLflow Model Registry). Experiments (Weights & Biases or MLflow). Configurations (Hydra or OmegaConf).
My current workflow: feature branch per experiment. DVC tracks data changes. MLflow logs every metric. Automated testing in CI catches regressions. Every repository has a model card.
The moment that convinced me to change: a client asked me to reproduce results from six months earlier. I couldn't. Not because the code was lost — because I didn't know which data version and which config produced those results.
Reproducibility isn't a nice-to-have. In regulated industries, it's legally required. For everyone else, it's the foundation of trustworthy AI.
Set up proper versioning once. Save yourself hundreds of hours of "which experiment was that?" forever.