The quality of your fine-tuning dataset matters more than the quantity.
I learned this by wasting two weeks on a bad dataset.
We had 50,000 examples for fine-tuning. The model trained fine. Metrics looked decent. But the outputs were inconsistent and sometimes bizarre.
The problem: our 50,000 examples were machine-generated, low quality, and full of contradictions. Some labeled the same input differently. Others had responses that were technically correct but stylistically wrong for the use case.
We curated 2,000 high-quality examples instead. Hand-verified, consistent, representative of the actual use case. The fine-tuned model was dramatically better.
Rules for fine-tuning data that actually work:
Quality over quantity. 1,000 excellent examples beat 50,000 mediocre ones.
Consistency is crucial. If similar inputs get wildly different response styles, the model learns noise.
Diversity within quality. Cover the range of inputs your model will see, but make sure every example is high quality.
Include edge cases deliberately. The model learns from what it sees. If your training data doesn't include weird inputs, the model won't handle them.
Format matters. If you want JSON output, every training example should have clean JSON. If you want concise answers, don't include verbose ones.
Treat your fine-tuning dataset like a product. Curate it. Version it. Improve it over time. It's the single highest-leverage asset in your fine-tuning pipeline.