A client once asked me to build an image classifier for their specific product line.
They had 500 labeled images.
A few years ago, I'd have said "we need at least 50,000 images." Instead, I took a pre-trained ResNet, fine-tuned the last few layers on their 500 images, and had a working classifier in an afternoon. 93% accuracy.
Transfer learning is so established now that training from scratch almost never makes sense. Why learn what a cat looks like when someone's already spent millions in compute teaching a model that?
The mental model is simple: start with a pre-trained model (ResNet for images, BERT for text, Whisper for audio, Llama for language tasks, CLIP for image-text). Adapt it to your specific domain with your (potentially small) dataset.
10x less data needed. 100x less compute. Often better results because you're building on robust foundations.
The most productive ML engineers in 2026 don't train models from scratch. They adapt existing ones. It's not lazy — it's smart engineering. Standing on the shoulders of giants isn't just a metaphor. It's a strategy.