The Monoculture Problem: Why AI-Generated Code Fails in the Same Places

the challenge of ai security moculture

AI-generated code is not uniquely buggy. It is uniformly buggy, and that is the part attackers should find exciting.

The usual debate asks whether a model writes worse code than a human. That is the wrong question. Humans have been shipping SQL injection since before most of today’s models’ training data was written. The more interesting question is what happens when millions of developers ask a handful of models for the same login endpoint, and get back the same code with the same mistake in the same place.

In agriculture, that is called a monoculture. It is very efficient right up until the blight arrives.

Human bugs are artisanal. Model bugs are mass-produced.

When a human writes a vulnerability, it usually reflects that person’s habits, deadline, and caffeine level. Two developers building the same feature tend to fail in different ways. That variety is accidental, but it is protective: a flaw found in one codebase tells an attacker little about the next.

Models don’t work that way. They learned from the same public code, they converge on the same idioms, and they tend to make the same choices when a prompt leaves security unspecified. The failures are correlated.

Veracode’s 2026 GenAI Code Security Report shows how uneven, and therefore how predictable, those failures are. Across more than 100 models, the average security pass rate was 56%, barely moved from 55% in the first report. Syntax is now essentially solved. Security is not.

Vulnerability classAverage security pass rate
Cryptographic algorithms87%
SQL injection83%
Cross-site scripting (CWE-80)15%
Log injection (CWE-117)12%

That spread is the tell. Models have learned the famous lessons, like parameterized queries. They still fail at flaws that depend on following untrusted data through an application. If you were choosing where to look first in a codebase you suspected was generated, this table is a decent treasure map.

One finding, many targets

Correlated bugs change the attacker’s economics. Vulnerability research is expensive; reusing it is cheap. When the same model produces the same flawed pattern across thousands of projects, a single discovery stops being a finding and becomes a search query.

The workflow writes itself:

  1. Prompt popular models for common features (file upload, password reset, webhook handler).
  2. Note which insecure constructs come back reliably.
  3. Turn those constructs into code-search patterns or scanner signatures.
  4. Sweep public repositories, exposed apps, and bug bounty scopes for matches.

None of this requires knowing that a given codebase was AI-generated. It only requires that a lot of codebases were. According to the same Veracode report, AI now authors roughly half of committed code in organizations that have adopted these tools. The haystack is increasingly made of identical needles.

Attackers also have the same AI tools defenders do. The step from “this pattern is common” to “here is a scanner for it” is now an afternoon, not a quarter.

Slopsquatting: the vulnerability that didn’t exist until the models invented it

The purest example of monoculture risk is the hallucinated dependency. Models sometimes recommend packages that do not exist. An attacker registers that name on npm or PyPI, fills it with something unpleasant, and waits for the next developer to paste the install command. The practice is called slopsquatting, a term coined by PSF Developer-in-Residence Seth Larson.

The research behind it, We Have a Package for You! (Spracklen et al., USENIX Security 2025), tested 16 code-generation models across 576,000 Python and JavaScript samples. The findings, as summarized by Socket:

  • 19.7% of recommended packages did not exist, adding up to more than 205,000 unique invented names.
  • Open-source models hallucinated at 21.7% on average; commercial models at 5.2%.
  • When prompts that triggered a hallucination were rerun 10 times, 43% of the fake names came back every single time, and 58% came back more than once.

That last number is the monoculture argument in one statistic. A random mistake is useless to an attacker. A mistake that repeats on demand is a registration opportunity. You don’t have to guess which names developers will be told to install; you can simply ask.

The names are also convincing. Only 13% were simple typos of real packages, so the old typosquatting defenses, which look for names suspiciously close to popular ones, mostly don’t apply.

Clean code, dirty secrets

Monoculture would matter less if every generated line got a careful review. It doesn’t, for two reasons.

Volume. AI makes code cheap to produce, and review capacity did not scale with it. Veracode’s framing is blunt: the failure rate stayed flat while the amount of code it applies to surged. A 44% flaw rate on a trickle is a backlog. On a flood, it is a strategy.

Confidence. Generated code usually compiles, runs, follows conventions, and arrives neatly formatted. It looks like the work of a careful senior engineer, and people review it accordingly, which is to say lightly. Polish is not correctness, but it is very good at impersonating it.

There is also the matter of where models learned to code. They reproduce what was common in their training data, and common is not the same as current. Deprecated APIs, outdated crypto defaults, and tutorial-grade shortcuts all survive in the output long after the humans who wrote them have learned better.

In fairness to the machines

The monoculture argument has real limits, and it is worth stating them before anyone starts a movement.

Human code was never that diverse. A generation of developers copied the same Stack Overflow answers and the same framework tutorials. Monoculture predates LLMs; models mostly industrialized it.

Predictability cuts both ways. If a model’s failure modes are consistent, defenders can write targeted static analysis rules and catch them at scale. A known weakness is a fixable weakness. The same Socket-summarized study found some models could flag their own hallucinated packages more than 75% of the time when asked, which suggests cheap self-checks help.

Model choice matters. Veracode’s leading model passed 68% of security tasks, while more than half of models sat at 50–53%. Reasoning models averaged 56% against 51% for non-reasoning ones. Security is not improving on its own as models get bigger, but it does vary enough that choosing a model is now partly a security decision.

None of this makes the risk go away. It does mean the outcome depends less on the model than on what happens after it hands you the code.

Breaking up the monoculture

You cannot make every model write different code. You can make sure its predictable mistakes don’t survive contact with your pipeline.

  • Treat generated code as untrusted input. Review it like a pull request from a stranger who types very fast and is extremely confident.
  • Verify every dependency exists and is the one you meant. Check the registry, the maintainer, the publish date, and the download history before installing. A package created last Tuesday with a perfect name deserves suspicion.
  • Scan in the workflow, not after it. Static analysis and software composition analysis belong on every pull request, so flaws are caught while the author still remembers writing them.
  • Put your review effort where models are weakest. Input that flows into HTML, logs, file paths, and shell commands deserves the closest look.
  • Ask for security explicitly. Prompts that specify parameterized queries, output encoding, and authorization checks get better results than prompts that assume them.
  • Pin and lock dependencies. Lockfiles and allowlists turn a hallucinated package from a silent install into a failed build.

The bottom line

AI-generated code is not dangerous because it is worse than human code. In some areas it is measurably better. It is dangerous because it is homogeneous, produced in enormous volume, and reviewed less carefully than it deserves. Those three properties together turn individual bugs into shared infrastructure, and shared infrastructure is exactly what attackers like to find.

The fix is not to stop using the tools. It is to stop trusting that clean-looking code is safe code, and to build pipelines that assume every model has habits, because every model does.

Sources