// / Blog

Notes from the
edge of the model.

Field notes on what actually breaks in production — agents, retrieval, evaluation, MLOps, and the career decisions nobody writes down. Longer arguments become papers; these are the rest.

354 posts

I inherited an ML project with no documentation.

It took me 3 weeks just to understand what the system did. Another 2 weeks to figure out how to retrain the model. Another week to discover there was a critical data preprocessing step…

Read post

Standard RAG is "retrieve then answer." Agentic RAG is "think about what to retrieve, retrieve…

The difference in practice is enormous. Standard RAG: user asks a complex question. System retrieves top-5 documents by similarity. Feeds them to the LLM. If the right information wasn't…

Read post

A company asked me to build a model using data scraped from social media profiles without users'…

I said no. Not because it was technically difficult. Because it was wrong. The data was public (people had posted it online). The use case wasn't malicious (predicting consumer…

Read post

Training a model on a single GPU: straightforward.

I spent two weeks debugging a distributed training job that worked on 1 GPU but produced garbage on 4. The culprit: a subtle issue with batch normalization statistics not being properly…

Read post

The most impactful AI application I've seen isn't a chatbot, a recommendation engine, or an…

It's a real-time captioning system for deaf and hard-of-hearing users. Whisper + a lightweight LLM for formatting + a low-latency streaming pipeline = live captions that are more…

Read post

We set up 43 monitoring alerts for our ML system. Within a week, we turned off 38 of them.

Alert fatigue is real. When everything is alerting, nothing is alerting. The mistake: alerting on every metric deviation. A 2% change in prediction distribution at 3 AM is not an…

Read post

Not every AI user is a Fortune 500 company.

A local restaurant wanted to predict daily ingredient needs to reduce food waste. We built a simple time-series model using their POS data and local event calendars. Food waste dropped…

Read post

We started with a monolithic ML application.

It worked great until we needed to update the embedding model without redeploying the entire system. Or scale the inference service independently of the data processing pipeline. Or use…

Read post

A bank rejected someone's loan application. They asked why. The answer: "The model said so."

This is not acceptable. Not legally (under GDPR and the EU AI Act), not ethically, and not practically — because an unexplainable decision can't be appealed or corrected. Explainability…

Read post

A marketing team I worked with was spending $15,000/month on content creation.

Not by replacing writers. By augmenting them. The pipeline: AI generates first drafts based on briefs and brand guidelines. Human writers edit, add voice, and ensure quality. AI handles…

Read post

I've used 6 different vector databases in production. Here's my honest, no-sponsor comparison.

Pinecone: Managed, zero-ops, works immediately. Best for teams that don't want to manage infrastructure. Expensive at scale. Great for getting started and staying in production…

Read post

Every Saturday morning I build something small with a technology I've never used before.

Last month's weekend projects: Week 1: Built a simple MCP server to understand the protocol. Took 3 hours. Now I understand what all the hype is about. Week 2: Trained a small LoRA…

Read post