All posts
// / Blog

The system that worked perfectly for 100 users completely fell apart at 10,000.

It was an LLM-powered search feature. At 100 users, response times were under a second. At 10,000, some users waited 30+ seconds. Others got timeout errors.

The bottleneck wasn't the model. It was everything around the model.

Embedding computation was synchronous — each request waited for the previous one. The vector database had a single connection, creating a queue. The LLM calls had no timeout or retry logic, so one slow response held up everything behind it.

What I learned about scaling ML systems:

Make everything async. Embedding computation, retrieval, and LLM calls should all be non-blocking.

Connection pooling for databases. A single connection is a bottleneck that's invisible at low load.

Set aggressive timeouts. A 30-second response is worse than a "please try again" message.

Add a queue for non-real-time requests. Not everything needs an instant response.

Cache aggressively. At scale, cache hit rates go up because query diversity is less than you think.

Load test BEFORE you launch. Not after users complain. Use locust or k6 to simulate realistic traffic patterns.

The model is rarely the bottleneck. The infrastructure around it usually is.

#Scalability#SystemDesign#MachineLearning#Performance#Engineering#MLOps