All posts
// / Blog

OpenAI cut GPT-6 prices in half. That's not the number I care about.

Controversial take: a 50% price cut on a cloud API doesn't remove a single constraint from the systems I actually build.

This week OpenAI cut prices on GPT-6 Sol and Luna, its two workhorse-tier models, by half. Sol now runs $2 per million input tokens and $10 per million output. Luna is $0.10 and $0.50. OpenAI says these are permanent rates, not a promotion, and on their internal factuality evals Sol makes roughly half as many mistakes as its predecessor while getting close to flagship-level reliability. Cheaper and more accurate at the same time. That's a genuine jump, and I don't want to undersell it.

But read that paragraph again and notice what it's actually describing: a change to a bill. Not a change to where the model runs, what it can see, or how it fails.

I've shipped 250+ AI systems into production across automotive, defence, and education, and in almost none of them was the per-token API price the thing standing between a demo and a working deployment. An offline multilingual AI avatar I built for Mercedes-Benz Germany has to run inside a vehicle or a showroom kiosk with no guaranteed uplink — a 50% API discount is irrelevant if there's no API to call. Real-time drone threat detection I worked on for the Indian Army and Police has a latency budget measured against a moving target, not against a round trip to somebody else's data center, and the data involved doesn't leave the perimeter, full stop, at any price. Research platforms I've run at IIT Bombay and IISc serving 3,000+ users have to work through the connectivity a public university actually has, not the connectivity a Bay Area office has.

None of that changes because Luna got 50% cheaper. Connectivity, data residency, and latency are architecture requirements. They don't show up on an invoice, and no invoice fixes them.

Where the price cut does matter is real, so give it its due: the high-volume, non-critical, already-connected inner loop of an agent pipeline. Summarizing logs. Triaging tickets. Drafting the tenth pass of a report nobody will read twice. At $0.10/$0.50, Luna is cheaper than most teams' own inference bill for a comparable open-weights model once you count the GPU time and the on-call engineer babysitting it. If that's your workload and your data is already fine living in someone else's cloud, take the discount and move on. I would.

But watch what the discount quietly does to planning conversations. It's easy to let "which frontier model is cheapest this quarter" become the whole infrastructure decision, because it's the number everyone can see and compare. The harder, more consequential question doesn't show up on a pricing page: which layer of your stack is even allowed to talk to a frontier API at all, given where the data has to live and what happens if the network drops mid-request. That question gets asked once, in a design review, months before anyone requests a quote — and it's the one that determines whether you're building on a model that runs where your data is, or one that just got cheaper to rent.

I keep coming back to the same bias after a decade of shipping this stuff: build for on-device and offline first, and treat cloud API access as an optimization you're allowed to add later, not a foundation you build on top of. Foundations don't get bought back at a discount.

Price the model. Then ask whether it's even allowed on the network you're deploying to. That second question doesn't get a 50% discount.

#LLMOps#on-device-ai#AI-economics#production-ai#edge-ai