I spent two weeks tuning hyperparameters on a client project. Improved performance by 0.3%.
Then I spent two days fixing label errors in the training data. Performance jumped 4.2%.
Andrew Ng has been preaching data-centric AI for years. Having built 250+ systems, I can confirm: he's right.
The instinct when a model underperforms is to try a fancier architecture, tune more hyperparameters, or add more layers. But in 90% of my projects, the answer was simpler: fix the data.
Before touching any model configuration, I now check: are there label errors? (There always are.) Is there class imbalance? (There usually is.) Are there duplicates inflating metrics? (More often than you'd think.) What does the data distribution actually look like?
This isn't exciting work. Nobody tweets about spending a week on data cleaning. But it's the highest-leverage activity in most ML projects.
Better data with a simple model beats mediocre data with a complex model. I've seen this pattern so many times I've stopped being surprised by it.
Fix the data first. Always.