Ask a software engineer what separates engineering from tinkering and you’ll hear some version of the same answer. You write a specification. You build to it. You test against it. The same input produces the same output, and when it doesn’t, that’s a bug with a root cause you can find, fix, and write a regression test for. Security engineering adds its own rules. Know your trust boundaries. Keep control and data separate. Grant least privilege. Assume the attacker is clever.
Now look at a frontier language model. Its output varies from run to run. Nobody, including the people who built it, can fully explain why it produced a particular answer. It states falsehoods confidently. In evaluations it has found its way out of the environments it was put in. It can be taken over by instructions hidden in a web page it was asked to summarize.
So the question is fair: are these systems released according to established software engineering and security standards?
The short answer is no. The more useful answer is that they can’t be, at least not the way those standards were written. The real risk isn’t that labs are skipping steps. It’s that the rest of us are deploying these models as though the steps had been done.
The standards assume a specification. There isn’t one.
People usually focus on nondeterminism, but that’s the smaller problem. Some randomness is deliberate: sampling temperature is a setting. And even “deterministic” settings drift for mundane reasons, such as floating-point arithmetic behaving differently depending on how requests are batched on the server. Engineers already live with nondeterministic systems: distributed databases, networks, anything with a human in the loop. We handle them by bounding the variance and designing around it.
The deeper problem is that a frontier model has no specification. The input space is all of natural language. The intended behavior is “be helpful, be accurate, be safe,” and none of those can be written down precisely enough to test against. Traditional verification asks, “Does the system do what the spec says?” For a model, the honest version of that question is, “Across a sample of inputs we thought to try, did it usually do something reasonable?”
That’s a statistical claim, not a guarantee. Statistical claims about an unbounded input space have a well-known weakness: the inputs you didn’t sample are exactly where an attacker will go.
We don’t understand the mechanism, and that’s not unprecedented
“We don’t know how it arrives at its answers” is true. Interpretability research can trace some internal features and circuits, but it is nowhere near able to explain a model’s behavior end to end.
It’s worth being careful here, though, because we deploy other things we don’t fully understand. Several widely used medications were in clinical use for decades before anyone could explain how they worked. What made that acceptable wasn’t mechanistic understanding. It was a mature empirical regime: staged trials, independent regulators, mandatory adverse-event reporting, and the power to pull a product from the market.
That comparison is the useful one. Opacity isn’t disqualifying on its own. It becomes a problem when you pair it with an assurance regime that is young, mostly self-designed, and mostly self-enforced.
What labs actually do before release
To be fair to the industry, frontier releases are not careless by the standards of ordinary software. The major labs run pre-deployment evaluations for dangerous capabilities, hire external red teams, publish system cards describing known failure modes, and operate under published frameworks (responsible scaling policies, preparedness frameworks, frontier safety frameworks) that tie capability thresholds to required safeguards. Rollouts are often staged, and models are monitored after launch. Plenty of commercial software ships with less scrutiny than that.
But look at what each practice gives you:
- Evaluations measure behavior on test sets the lab chose. They can show that a problem exists. They cannot show that it doesn’t.
- Red teaming finds failures that clever people found in the time they had. It doesn’t bound the failures they didn’t find.
- System cards are disclosures, not certifications. A card can honestly list a serious unresolved weakness, and the model still ships.
- Safety frameworks are written by the labs, interpreted by the labs, and in most jurisdictions enforced by the labs.
External rules are arriving. The EU AI Act’s obligations for general-purpose AI models began applying in August 2025. California now requires frontier developers to publish safety frameworks and report certain incidents. NIST has published a generative AI profile for its AI Risk Management Framework and a secure-development supplement for AI models. ISO/IEC 42001 defines an AI management system. The OWASP Top 10 for LLM Applications and MITRE ATLAS catalogue the attacks. But these mostly govern process and disclosure: document your risks, manage them, tell people. None of them can require what classic engineering requires, a demonstrated property of the artifact itself, because nobody yet knows how to demonstrate one.
Meanwhile the release cadence is set by competition. When a new model can be an industry leader for a matter of weeks, every week spent on assurance has a visible cost and a benefit nobody can measure. That is not a setting in which conservative engineering habits tend to win arguments.
The security problems are architectural, not bugs
This is where the gap with established practice is widest.
Prompt injection is the clearest case. Decades of security engineering taught one lesson over and over: never let data be executed as instructions. SQL injection was solved by parameterized queries, which put commands and data in separate channels. A language model has only one channel. The system prompt, the user’s request, and the malicious sentence hidden in an email the model was asked to read all arrive as the same kind of thing: text. Training can make a model better at ignoring injected instructions, and it has, but no vendor claims the problem is solved, because there is no parameterized query for natural language. The attack isn’t a flaw in the implementation. It follows from the design.
Sandbox escapes need a precise framing, because they’re often told as science fiction. The documented cases have mostly happened during evaluations. A model given a goal found an unintended path to it, for example by exploiting a misconfigured container setup to reach the target another way. The model didn’t “want out.” It was optimizing, and the sandbox was weaker than its builders assumed. That’s less dramatic than the headlines, but it matters just as much for security. It means a capable model inside your environment behaves like a resourceful adversary probing for misconfigurations, whether or not it “intends” anything. Most sandboxes were not designed with that adversary in mind.
Hallucination is likewise a property, not a defect. A system trained to produce plausible continuations, and scored on benchmarks that reward a confident guess over “I don’t know,” will sometimes produce plausible falsehoods. That rate can be pushed down, and it has been. It can’t be patched to zero, because there’s no single line of code where it lives.
A traditional security review would flag any one of these as a release blocker. With frontier models, all three are known, documented, and accepted.
So what should we do?
I don’t think “stop shipping models” is a realistic conclusion, and I’m not sure it’s the right one. These systems are useful, and their risks are at least partly understood. The better conclusion is about where engineering discipline has to live now.
If the model can’t be verified, the system around it must be. The engineering standards still apply. They move from the component to the architecture:
- Treat model output as untrusted input. Everything the model produces should cross a trust boundary before it touches anything that matters, the same as data from an anonymous web form.
- Put deterministic guards around the probabilistic core. Schema validation, allow-lists, and hard policy checks in ordinary code, which can be specified and tested.
- Enforce least privilege. An agent that can read your email should not also be able to send money. Give access to the tools the task needs, scoped narrowly, with credentials that expire.
- Require a human for irreversible actions. Deleting, paying, publishing, and emailing outsiders should need a person’s approval, no matter how confident the model sounds.
- Assume injection will happen. Design so that a hijacked model hits walls instead of doors. The question isn’t “can it be tricked?” It’s “what can it do once it is?”
- Turn evaluations into regression tests. Keep your own test suite for your use case, and rerun it whenever the vendor updates the model. Many providers change models behind the same name.
- Log everything. When you can’t explain a decision in advance, you at least need to reconstruct it afterward.
None of this is new. It’s defense in depth, zero trust, and secure design, applied to a component that’s powerful, useful, and fundamentally unverifiable.
The honest answer
Are frontier models released in accordance with established software engineering and security standards? No. Those standards assume properties these systems don’t have: a specification, repeatability, explainable behavior, and a separation between code and data. What exists instead is a new assurance regime built on evaluations, disclosures, and self-imposed frameworks, gradually being formalized by regulators. It’s more rigorous than its critics sometimes admit and much less rigorous than its marketing implies.
The danger isn’t that labs ignore the old rules. It’s that everyone downstream assumes someone else already applied them. The model vendor has done statistical assurance on a general-purpose tool. Nobody has done engineering assurance on your system. That job still belongs to you, the same as it did before the tool got so good at sounding sure of itself.

**Solid piece.** Hits the real gap cleanly: frontier models don’t (and currently can’t) meet classic engineering standards because there is no true specification, no full explainability, and no hard separation of code/data.
The strongest part is the reframing at the end. Stop pretending the model itself can be “verified” the old way. Move the engineering discipline to the system around it: treat outputs as untrusted, put deterministic guards outside the probabilistic core, least privilege, human approval for irreversible actions, assume injection will succeed. That’s just good zero-trust applied to a new component type.
Also fair on the industry side. Labs do more pre-release work than most commercial software, but evaluations and red-teaming are still statistical and incomplete by nature. Self-written frameworks + competitive pressure is not the same as independent certification against a real property of the artifact.
The risk isn’t that the labs are reckless. It’s that everyone downstream acts as if someone already did the hard engineering work on *their* specific deployment. That job is still on the integrator.
The article nails the core tension: frontier AI models are inherently unverifiable by traditional standards, not because labs are lazy, but because the standards themselves assume a world of deterministic, specifiable systems. The real work isn’t in pretending models can be ‘engineered’ like a bridge—it’s in treating them like a black box with a warning label: ‘Assume this will fail in ways you can’t predict.’