// / Blog

Notes from the
edge of the model.

Field notes on what actually breaks in production — agents, retrieval, evaluation, MLOps, and the career decisions nobody writes down. Longer arguments become papers; these are the rest.

353 posts

I've interviewed at about 15 companies for AI roles.

Red flags I wish I'd recognized earlier: "We use AI for everything." — If they can't tell you specifically what problem AI solves, they're in hype-driven development mode. "The model is…

Read post

Not everything needs to be real-time. And that's one of the most expensive lessons in ML.

A client wanted real-time recommendations. The engineering complexity and infrastructure cost was enormous — streaming pipeline, low-latency serving, real-time feature computation. I…

Read post

I taught an ML course last semester.

You know what's really hard? Explaining backpropagation to someone who last did calculus in high school. Making gradient descent intuitive without oversimplifying. Helping students debug…

Read post

The most impressive demo I saw this year: an AI system for insurance claims processing.

Take a photo of car damage. Record a voice description. Upload the police report PDF. The system processes all three inputs simultaneously, cross-references them for consistency…

Read post

The best model I ever deployed was a logistic regression.

A client wanted a churn prediction system. I started with logistic regression as a baseline while the "real" model (a gradient boosted ensemble with 200 features) was being developed…

Read post

ML team standups should be different from regular engineering standups. Here's why ours changed.

Traditional standup: "Yesterday I worked on the feature extraction module. Today I'll continue. No blockers." That tells the team almost nothing useful for ML work. Our ML standup…

Read post

A client looked at my model's predictions and said: "This is completely wrong for our use case."

My first instinct was defensive. The metrics were good. The architecture was sound. The training data was curated. But they were right. The model predicted customer behavior based on…

Read post

Your evaluation dataset is lying to you. Here's how to catch it.

Three ways eval datasets mislead: Data leakage between train and eval. More common than you think, especially when both come from the same source. A duplicate example in both sets…

Read post

I spent a year as the bridge between an ML research team and a product engineering team.

The researchers: brilliant, creative, published at top conferences. Their code worked in notebooks. Their experiments were fascinating. Their prototypes couldn't handle 100 concurrent…

Read post

Context window length is the most misunderstood feature in LLMs.

"Our model supports 200K tokens!" Great. But can it actually USE 200K tokens effectively? I ran an experiment: buried a specific fact at different positions in a long context and asked…

Read post

The hardest part of deploying AI in a traditional company isn't the technology. It's the trust.

I deployed a demand forecasting model at a retail company. The model was 30% more accurate than the existing spreadsheet-based approach. But the procurement team refused to use it for 4…

Read post

I needed 100,000 product descriptions for a classification model. Budget for data labeling: zero.

Web scraping saved the project. But it also taught me that scraping for ML has its own set of challenges. Data quality from scraping is wildly inconsistent. HTML structures change…

Read post