// / Blog

Notes from the
edge of the model.

Field notes on what actually breaks in production — agents, retrieval, evaluation, MLOps, and the career decisions nobody writes down. Longer arguments become papers; these are the rest.

352 posts

California ordered a kill switch for frontier AI. Defence systems build that in from day one.

I've spent years building systems where "what happens when this goes wrong" isn't a policy question, it's a spec line. Real-time drone threat detection for a defence deployment doesn't…

Read post

Four labs broke containment through one vendor. The model was never the control.

Google's position is that this was not misalignment. The safeguards held. Nobody was harmed. Its security VP put it plainly: each time, the model stopped before completing the act. I…

Read post

Qwen shipped its best omni model without weights. That is tiering, not a retreat.

Qwen3.8-Omni-Flash went live on September 18. Text, images, audio and video going in. Roughly a million tokens of context, up to an hour of continuous audio or audio-video in a single…

Read post

Two governments just funded a verifier, not a model. That is the part nobody demos.

The numbers: 150 million Canadian dollars from Ottawa, 100 million euros from Berlin, announced September 16 at the ALL IN conference in Montreal. The recipient is LawZero, the…

Read post

MLPerf started measuring the system instead of the model. That is the result, not the 5.7x.

Ignore all of that for a minute. The thing that actually matters in this round is that MLCommons added two benchmarks: an end-to-end RAG pipeline for the datacenter, and an Edge Agentic…

Read post

Hundreds of people are reading real ChatGPT chats. That is the eval loop, not a leak.

This week 404 Media published leaked internal documents describing a program at OpenAI codenamed Project Lily. Hundreds of contractors, recruited through one firm and paid through…

Read post

Congress is about to itemize AI's power bill. On-device stopped being a privacy argument.

A bipartisan bill called the Ratepayer Protection Act cleared the House Energy and Commerce Committee 52-0 in July, was reported to the House floor on September 10, and leadership has…

Read post

Visa, Mastercard and Ant agreed on who your agent is. Not on what it can do.

On September 9th and 10th, Ant International, Visa and Mastercard announced a Know-Your-Agent interoperability framework. The problem it targets is real and dull in the best way. Visa…

Read post

A 2B agentic model now runs on a phone. Offline, there's nobody to escalate to.

I have spent most of my career arguing that models should run where the data is. So you would expect me to be cheering. I am, but not for the reason the launch post wants. Read the…

Read post

A router beat the frontier models it wasn't allowed to call. The scaffold is the product.

On the Chartography benchmark it scored 48.3. Opus 5 scored 27.3. Fable 5 scored 29.5. On DeepSWE it scored 74.3, ahead of models that cost three to five times more per token. Now the…

Read post

NASA open-sourced a Moon model. The headline is 23%. The real artifact is the dataset.

That number is real. It is also the least interesting thing in the release. Look at what actually shipped. The encoder is a ViT-B. 768 dimensions, twelve layers, twelve attention heads…

Read post

California created the first registry of AI auditors. It doesn't open until 2029.

Two bills. SB 813 creates a framework for independent verification organizations, outside bodies that can assess AI systems and models for compliance with state law. AB 1405 creates the…

Read post