The most impactful optimization I've ever done wasn't about the model. It was adding a cache.
The system was an LLM-powered customer service bot. Processing every query through the LLM at $0.03 per request. 10,000 queries per day. That's $300/day just in API costs.
But here's the thing — customer service questions are repetitive. "What's your return policy?" gets asked 200 times a day in slightly different ways.
I added a semantic cache: embed every query, check if a similar query (cosine similarity > 0.92) has been answered recently, and return the cached response if so.
Cache hit rate: 43%. Cost dropped from $300/day to $170/day. Latency for cached responses: 50ms instead of 2 seconds.
Caching strategies for LLM applications that work: exact match cache for identical prompts (simplest), semantic cache for similar queries (most impactful), prompt prefix caching for queries that share system prompts, and response streaming cache for repeated structured outputs.
Redis for the cache store. A lightweight embedding model for similarity computation. Total implementation time: one afternoon.
If you're running an LLM in production without caching, you're probably overpaying by 30-50%. Low-hanging fruit doesn't get lower than this.