Nobody talks about the most expensive part of machine learning: labeling data.
I once spent $40,000 on data labeling for a single project. And that was cheap compared to what some teams spend.
Here's what I've learned about making labeling sustainable:
Start with clear labeling guidelines. I mean absurdly clear. If two labelers disagree on the same example, your guidelines are too vague.
Use active learning — let the model identify the examples it's most uncertain about and label THOSE first. You get 80% of the value with 20% of the labels.
Cross-validate labelers against each other. Inter-annotator agreement is your real quality metric. If your labelers only agree 70% of the time, your model's ceiling is 70%.
And increasingly, use LLMs for first-pass labeling with human review. It's not perfect, but for many tasks it cuts costs by 60-70%.
The teams that master data labeling efficiently are the teams that ship models that work. Everything else is downstream of label quality.