Notes from the
edge of the model.
Field notes on what actually breaks in production — agents, retrieval, evaluation, MLOps, and the career decisions nobody writes down. Longer arguments become papers; these are the rest.
352 posts
California ordered a kill switch for frontier AI. Defence systems build that in from day one.
I've spent years building systems where "what happens when this goes wrong" isn't a policy question, it's a spec line. Real-time drone threat detection for a defence deployment doesn't…
By Pranay Mahendrakar Read postFour labs broke containment through one vendor. The model was never the control.
Google's position is that this was not misalignment. The safeguards held. Nobody was harmed. Its security VP put it plainly: each time, the model stopped before completing the act. I…
By Pranay Mahendrakar Read postQwen shipped its best omni model without weights. That is tiering, not a retreat.
Qwen3.8-Omni-Flash went live on September 18. Text, images, audio and video going in. Roughly a million tokens of context, up to an hour of continuous audio or audio-video in a single…
By Pranay Mahendrakar Read postTwo governments just funded a verifier, not a model. That is the part nobody demos.
The numbers: 150 million Canadian dollars from Ottawa, 100 million euros from Berlin, announced September 16 at the ALL IN conference in Montreal. The recipient is LawZero, the…
By Pranay Mahendrakar Read postMLPerf started measuring the system instead of the model. That is the result, not the 5.7x.
Ignore all of that for a minute. The thing that actually matters in this round is that MLCommons added two benchmarks: an end-to-end RAG pipeline for the datacenter, and an Edge Agentic…
By Pranay Mahendrakar Read postHundreds of people are reading real ChatGPT chats. That is the eval loop, not a leak.
This week 404 Media published leaked internal documents describing a program at OpenAI codenamed Project Lily. Hundreds of contractors, recruited through one firm and paid through…
By Pranay Mahendrakar Read postCongress is about to itemize AI's power bill. On-device stopped being a privacy argument.
A bipartisan bill called the Ratepayer Protection Act cleared the House Energy and Commerce Committee 52-0 in July, was reported to the House floor on September 10, and leadership has…
By Pranay Mahendrakar Read postVisa, Mastercard and Ant agreed on who your agent is. Not on what it can do.
On September 9th and 10th, Ant International, Visa and Mastercard announced a Know-Your-Agent interoperability framework. The problem it targets is real and dull in the best way. Visa…
By Pranay Mahendrakar Read postA 2B agentic model now runs on a phone. Offline, there's nobody to escalate to.
I have spent most of my career arguing that models should run where the data is. So you would expect me to be cheering. I am, but not for the reason the launch post wants. Read the…
By Pranay Mahendrakar Read postA router beat the frontier models it wasn't allowed to call. The scaffold is the product.
On the Chartography benchmark it scored 48.3. Opus 5 scored 27.3. Fable 5 scored 29.5. On DeepSWE it scored 74.3, ahead of models that cost three to five times more per token. Now the…
By Pranay Mahendrakar Read postNASA open-sourced a Moon model. The headline is 23%. The real artifact is the dataset.
That number is real. It is also the least interesting thing in the release. Look at what actually shipped. The encoder is a ViT-B. 768 dimensions, twelve layers, twelve attention heads…
By Pranay Mahendrakar Read postCalifornia created the first registry of AI auditors. It doesn't open until 2029.
Two bills. SB 813 creates a framework for independent verification organizations, outside bodies that can assess AI systems and models for compliance with state law. AB 1405 creates the…
By Pranay Mahendrakar Read post