OpenAI's Astra solved 10 open math problems. The number that matters is the zero.
OpenAI's unreleased model Astra just produced machine-checked proofs for ten problems that had sat unsolved for a decade or more, including a construction of a non-sofic group that closed a question open since 1999, and a new upper bound on sphere packing that hadn't moved in 48 years. Total compute cost: about $2,000. Everyone is talking about the math. I want to talk about the zero.
The proofs weren't just generated. They were formalized as Lean 4 certificates and published on GitHub, and the repository's "sorry" count sits at zero. In Lean, "sorry" is the placeholder you drop in when you're asserting a step without actually proving it. Zero sorries means every single line, across all ten proofs, was mechanically checked with no human judgment in the loop. That's the part that should matter to anyone building with these models, not the fact that a language model did graduate-level algebra.
I spend most of my working life on the gap between a model that sounds right and a model that is right. One of my open-access papers is specifically about formal verification for language model outputs, and the short version of what I learned is this: fluency is cheap and correctness is expensive, and almost every production incident I've cleaned up traces back to someone treating the first as a proxy for the second. A model that writes a beautiful-looking proof with a hidden gap is worse than useless, because it looks finished. A model that writes a proof a theorem prover can check line by line either produces the zero, or it doesn't ship.
That's what Astra actually demonstrates. Not "AI can do math now" — plenty of systems have chipped away at open problems before. What's new is coupling generation with a verifier that has no opinion, no benefit of the doubt, and no way to be talked into accepting a hand-wave. Lean doesn't care how confident the model sounded. It either checks or it doesn't.
I've built systems where this exact pattern is the difference between a demo and a deployment. Threat detection models that flag targets don't get to be "mostly right" with a persuasive explanation attached — the output has to clear a hard verification gate before a human ever sees it, because the cost of a fluent false positive is not the same as the cost of a boring true negative. Multilingual systems handling Indic languages have the same problem in miniature: a hallucinated answer in Hindi or Kannada often reads more confidently than the correct one, because the model has less real signal to hedge on and fills the gap with fluency instead. In every one of these cases, the fix was never "get a smarter model." It was building a verifier the model couldn't sweet-talk.
That's also why I'm not that impressed by the $2,000 number, even though it's the headline everyone else grabbed. Cheap generation was never the bottleneck. Cheap generation has been here for two years and it's why LinkedIn is full of demos that don't survive contact with a real dataset. The bottleneck was always trust — knowing which of the model's confident-sounding outputs to believe. Astra's real contribution is showing that for domains with a formal verifier available, that trust problem has a clean answer: don't trust the model's confidence, trust the checker.
The catch, and it's a big one, is that most of what we ask production AI systems to do doesn't have a Lean 4 waiting to check it. There's no proof assistant for "is this customer support answer correct," no compiler for "did this RAG system retrieve the right clause from a 200-page contract." For math and formally specified code, the verification loop closes cleanly. For almost everything else, we're still building the equivalent by hand — test suites, guardrail models, human review gates, hallucination detectors — because the checker doesn't exist off the shelf.
So take the Astra result as a preview, not a template. It tells you what production AI looks like in the narrow slice of the world where a mechanical verifier already exists: generation gets cheap, correctness gets checked, and the zero is the only number that matters. For everyone else, the job right now is building that verifier yourself, for your domain, before you ship. The model that writes the answer was never the hard part. The system that catches it when it's wrong always was.