I tripled a model's accuracy on rare classes without collecting a single new data point.
The trick: aggressive, domain-informed data augmentation.
For an image classification task with 5,000 examples of common defects and 47 examples of rare ones: geometric transforms (rotation, flipping), color jittering, cutout/random erasing, mixup between rare class examples, and synthetic generation using a fine-tuned diffusion model.
47 real examples became 5,000 augmented examples. The rare class accuracy went from 31% to 89%.
For text tasks: back-translation (English → French → English gives you a natural paraphrase), synonym replacement with constraints, LLM-generated paraphrases, and entity swapping.
For tabular data: SMOTE is the baseline, but CTGAN produces more realistic synthetic samples for complex distributions.
The key insight: augmentation works best when it's domain-informed. Random noise rarely helps. But transforms that represent realistic variation (images taken at different angles, text written by different people, data with seasonal variation) help enormously.
Think about what variation your production data will have that your training data doesn't. Then augment to fill that gap.
It's the cheapest way to make your model more robust. And it should be the first thing you try before collecting more data.