// / Blog

Notes from the
edge of the model.

Field notes on what actually breaks in production — agents, retrieval, evaluation, MLOps, and the career decisions nobody writes down. Longer arguments become papers; these are the rest.

353 posts

I started posting about my AI work publicly two years ago. It was terrifying.

My first technical post got 12 views. My second got 8. My imposter syndrome screamed: "Who are you to teach anyone anything?" I kept posting anyway. Mostly about things I was learning…

Read post

The most impactful change I made to a RAG system wasn't the model, the embeddings, or the…

Default chunking (500 tokens with 50 overlap) is the "Hello World" of RAG. It works for demos. It fails in production. What I've learned through painful iteration: Semantic chunking…

Read post

A law firm asked me to build an AI that reviews contracts.

The lawyers' response: "What about the 6% it missed?" Fair point. In legal work, a missed clause can cost millions. The standard isn't "usually right." It's "never misses anything…

Read post

The most expensive mistake in AI isn't a bad model. It's leaving GPUs idle.

I audited a company's AI infrastructure last month. They had 8 A100 GPUs running 24/7. Average utilization: 23%. They were effectively burning $15,000/month on idle compute. GPU…

Read post

Your RAG system is only as good as your evaluation. And most RAG evaluations are terrible.

The standard approach: throw 10 questions at it, eyeball the answers, say "looks good." That's not evaluation. That's vibes. Here's the evaluation framework I use for every RAG system…

Read post

The system that worked perfectly for 100 users completely fell apart at 10,000.

It was an LLM-powered search feature. At 100 users, response times were under a second. At 10,000, some users waited 30+ seconds. Others got timeout errors. The bottleneck wasn't the…

Read post

A junior engineer asked me: "Should I learn LangChain or LlamaIndex?"

My answer: "What are you building?" They didn't have an answer. They wanted to learn a framework for the sake of having it on their resume. This is backwards. Frameworks are tools. You…

Read post

The most fragile system I ever built wasn't the model. It was the data pipeline feeding it.

17 data sources. 3 different formats (JSON, CSV, XML). Updated at different frequencies (real-time, hourly, daily). Different quality levels (some clean, some a mess). The model was…

Read post

I replaced a 5-person manual process with an AI agent pipeline.

The process: research a topic, collect information from multiple sources, synthesize findings, draft a report, review for accuracy, format and deliver. The agent pipeline: a research…

Read post

Our LLM application was working fine. Until it wasn't. And we had no idea what changed.

The problem: LLM behavior is non-deterministic. The same input can produce different outputs. So when quality degrades gradually, it's incredibly hard to pinpoint when and why. This is…

Read post

A general-purpose model scored 85% on our finance benchmark. After domain adaptation, it scored 96%.

The 11% difference? That's the difference between a demo and a production system in a specialized industry. Domain adaptation is the most reliable way to dramatically improve model…

Read post

An AI system I built was used in a way I never intended.

I designed a productivity analysis tool for a manufacturing client. It identified bottlenecks in production workflows. Good use case. Clear value. Six months later, I discovered they'd…

Read post