// / Blog

Notes from the
edge of the model.

Field notes on what actually breaks in production — agents, retrieval, evaluation, MLOps, and the career decisions nobody writes down. Longer arguments become papers; these are the rest.

353 posts

Voice AI is about to have its ChatGPT moment.

OpenAI's Whisper made speech recognition nearly free. Real-time translation is now possible. Voice cloning requires 3 seconds of audio. Text-to-speech is indistinguishable from human…

Read post

I used to dread meetings with non-technical stakeholders about ML projects.

What changed: I stopped presenting accuracy numbers and started presenting stories. Before: "The model achieves 92% precision and 87% recall on the test set with an F1 of 0.89." After…

Read post

Not every query deserves a GPT-4 response. Most don't.

I built a routing system that sends simple queries to a tiny model (1.5B params), medium-complexity queries to a 7B model, and only routes genuinely complex queries to a frontier model…

Read post

I've worked remotely on AI projects with teams across 7 countries.

The big challenge: ML experiments are harder to communicate asynchronously. "I tried X and it didn't work" tells your teammate nothing. They need to see the data, the metrics, the failed…

Read post

The day I switched from free-text LLM outputs to structured outputs, my production error rate…

Free-text: "The sentiment is probably positive, about 0.8 confidence." Structured: {"sentiment": "positive", "confidence": 0.82, "entities": ["product_x"]} Parsing free text from an LLM…

Read post

We deployed a model update on a Thursday afternoon. Monday morning, customer complaints tripled.

The new model was objectively better on our test set. But it was worse on a specific segment of high-value customers that our test set underrepresented. Rollback took 4 hours because we…

Read post

The most underused testing technique in ML: synthetic adversarial examples.

Most engineers test their model on a held-out test set and call it a day. But the test set comes from the same distribution as the training set. It doesn't test what happens when reality…

Read post

Standard RAG failed spectacularly on a project last quarter.

The chunks containing engineer profiles, project descriptions, and skill mappings were all in the vector store. But the semantic search retrieved fragments that were individually…

Read post

Every ML project I've estimated has taken longer than I predicted. Every single one.

After 250+ projects, here's why and what I do about it: ML projects have fundamental uncertainty that software projects don't. You don't know if the data is good enough until you try…

Read post

I ran a blind evaluation last month: gave 50 domain-specific queries to GPT-4, Claude, and a…

The fine-tuned Llama won on 31 of 50 queries for our specific use case. Not because open source models are "better" in general — they're not. But for a well-defined, specific domain…

Read post

Hot take: most ML projects have zero tests. And most ML engineers don't know what to test.

Testing ML systems is fundamentally different from testing software. You're not just checking "does this function return the right output?" You're checking "does this statistical system…

Read post

I tripled a model's accuracy on rare classes without collecting a single new data point.

The trick: aggressive, domain-informed data augmentation. For an image classification task with 5,000 examples of common defects and 47 examples of rare ones: geometric transforms…

Read post