I fine-tuned a Llama model on my company's internal documentation last month.
The whole thing — data prep, training, evaluation — took 6 hours on a single GPU.
Two years ago, that would've cost thousands in compute. Now it costs less than a nice dinner.
QLoRA changed everything. You're training 0.1% of the model's parameters while keeping 99% of the performance. Unsloth makes it even faster. Axolotl handles the config headaches.
But here's the nuance most tutorials skip: knowing WHEN to fine-tune.
Need the model to know new facts? Use RAG, not fine-tuning.
Need the model to behave differently — match a tone, follow a specific format, reason in a domain-specific way? That's when you fine-tune.
I see teams wasting weeks fine-tuning to add knowledge when they should've built a RAG pipeline. And vice versa.
The best results come from doing both. But knowing which tool fits which problem is the real skill.