All posts
// / Blog

Not everything needs to be real-time. And that's one of the most expensive lessons in ML.

A client wanted real-time recommendations. The engineering complexity and infrastructure cost was enormous — streaming pipeline, low-latency serving, real-time feature computation.

I asked: "How often do your users' preferences actually change?"

Turns out, their user behavior was remarkably stable. Weekly batch updates would have captured 95% of the recommendation quality at 10% of the infrastructure cost.

The decision framework I use:

Real-time is necessary when: the cost of stale predictions is high (fraud detection), the environment changes rapidly (dynamic pricing), or user experience demands instant feedback (search results).

Batch is sufficient when: predictions are consumed in non-time-sensitive contexts (email recommendations), underlying data changes slowly (content matching), or the cost of real-time infrastructure outweighs the quality gain.

The sweet spot for many applications is "near real-time" — update every few minutes or hours. You get most of the freshness benefit without the full complexity of streaming.

Every time someone says "we need this in real-time," my first question is "do you? Really?" Half the time, the honest answer is no. And the cost difference between real-time and batch is typically 5-10x.

Choose the right architecture for your actual latency requirements, not the coolest one.

#SystemDesign#MachineLearning#Architecture#RealTime#Engineering#MLOps