The single practice that improved my team's ML code quality the most wasn't a tool or a process.
It was making code reviews mandatory and specific.
Before: PRs sat for days. Reviews were rubber stamps. "LGTM" with no comments. Bugs made it to production weekly.
After: every PR reviewed within 24 hours. Reviewers required to leave at least two substantive comments. ML-specific checklist: is the train/test split correct? Is there data leakage? Are the metrics appropriate? Is the preprocessing identical for training and inference?
The ML-specific checklist was the game changer. Traditional code review catches syntax and logic bugs. ML code review catches the insidious bugs — data leakage, wrong evaluation, preprocessing mismatches — that produce code that runs perfectly but gives wrong results.
I've caught more production-impacting bugs in code review than in testing. Because ML tests often pass even when the logic is wrong (the code runs, it just produces a bad model).
If your ML team doesn't have a review checklist that includes data quality, evaluation methodology, and train-serve consistency — create one this week. It's the cheapest quality improvement you'll ever make.