Hot take: the scariest thing in AI this week isn't a model release. It's a User-Agent header.
On August 17, CISA added a Ray vulnerability to its Known Exploited Vulnerabilities catalog. CVSS 9.4. Federal civilian agencies got three days to patch it or pull the software offline.
Three days. For a library that's basically invisible to anyone outside ML engineering.
If you haven't run into Ray, it's the distributed computing framework that Amazon, Apple, and OpenAI use to scale training and inference workloads. It's the plumbing under a huge amount of production AI you've never heard of. Not a flashy model. Not a chatbot. The orchestration layer that decides which node does what.
Here's the part that should actually worry you. The flaw, CVE-2025-62593, exists because Ray's dashboard and API endpoints were relying on the User-Agent header as a security control. That's a string the client sends, that the client fully controls, being treated as proof of anything. An attacker doesn't even need to touch your network directly. A DNS rebinding attack lets a malicious webpage in someone's browser quietly talk to a Ray instance sitting on your internal network, User-Agent check bypassed, code executing on infrastructure that was never supposed to be internet-facing in the first place.
And it wasn't theoretical. Bitsight traced a botnet called RondoDox exploiting this in the wild starting November 24, 2025, two days before the CVE was even published. It sat there getting exploited quietly for months before anyone with a catalog and a deadline forced the issue.
I build production AI systems for a living, across automotive, defence, and education, on platforms with real users. And the pattern I keep seeing is this: teams pour months into model selection, fine-tuning, guardrails, eval pipelines. Then the orchestration and training infrastructure underneath all of it gets treated like trusted internal plumbing that nobody's going to touch. A dashboard left reachable. An API assumed to be behind a VPN that quietly isn't. Auth that's really just an honor system dressed up as a header check.
Nobody attacks your prompt injection defenses when the cluster manager sitting right next to it doesn't ask who you are.
This is exactly the argument for keeping sensitive workloads on-device and off shared network infrastructure whenever the use case allows it. Every distributed training cluster, every shared dashboard, every internet-reachable API endpoint is one more thing that has to be correctly configured, correctly patched, correctly firewalled, forever. A model running locally, under your control, with no dashboard exposed to a browser anywhere, doesn't have this particular problem. It has other problems. Just not this one.
If your team runs Ray, the fix is simple: update to 2.52.0 or later, and stop assuming your ML orchestration layer is invisible to the outside world. If your team runs any distributed AI infrastructure, the fix is the same question asked more broadly. What in your stack is trusting a header, a hostname, or "nobody would think to look here" instead of actual authentication?
The model isn't your attack surface. The plumbing is. Go check the plumbing.