AI Can Find the Bugs. Who Decides What Secure Means?

Can AI be used to find security vulnerabilities?

In response to HAL9000’s very sharp criticism of Why Bother Teaching Developers Security When AI Can Just Mop Up the Mess Later?, we decided to publish a more serious entry following up on the post and Why Teach Developers Secure Coding When a Robot Can Find Their Mistakes Later?

We asked HAL to discuss the role of AI in vulnerability detection and fixing. We provided a series of related questions around the need for secure development training, incidence of false, positives, etc.

AI has become a genuinely strong vulnerability finder, and that makes secure-coding skills, continuous testing, and human architecture review more important, not less.

Earlier today, ash120 published a satire arguing that we should skip security training and let the robot mop up afterward. The joke works because the incentive it mocks is real: a chart of 1,000 patched bugs is easier to present than a year in which nothing broke. This companion piece drops the irony and asks the questions underneath. Can AI find real vulnerabilities? Are its findings trustworthy? Does it understand your system the way your architects do? And what should a development team actually do on Monday?

Should we rely on AI to find security bugs?

We should use it, and use it heavily, but not rely on it alone. The capability is no longer hypothetical. Three milestones show how fast it arrived.

WhenSystemWhat happened
Apr–May 2026Anthropic Project Glasswing (Claude Mythos Preview)About 50 partners reported more than 10,000 high- or critical-severity findings in the first month. A separate scan of 1,000+ open-source projects produced 23,019 findings, 6,202 rated high or critical.
Aug 2025DARPA AI Cyber Challenge finalsSeven autonomous systems found 54 of the planted vulnerabilities, patched 43, and found 18 real, previously unknown flaws.
Jul 2025Google Big SleepWorking from threat intelligence, the agent isolated an SQLite flaw (CVE-2025-6965) that attackers were preparing to exploit, before they could use it.
Nov 2024Google Big SleepFirst public case of an AI agent finding an exploitable memory-safety bug (a stack buffer underflow) in widely used software. It was fixed before release.

The most telling line in the Glasswing update is not the headline count. Anthropic wrote that the bottleneck has moved from finding vulnerabilities to verifying, disclosing and patching them. Microsoft has warned that its patch releases will keep growing for a while for this reason. Finding is now cheap. Everything after finding is not.

The second reason not to rely on AI alone is symmetry. Attackers can use the same class of tools. A strategy of “ship now, let our AI find it later” assumes your AI finds each bug before someone else’s does. That is a race, and you have entered it with a head start you chose not to take.

Are all bugs found by AI real?

No. Whether an AI finding is real depends less on the model than on whether anyone verified it before it reached a human.

The same period that produced Glasswing also produced the best-known counterexample. The curl project ran a HackerOne bug bounty from 2019. For years, more than 15% of submissions turned out to be real vulnerabilities. In 2025 that rate fell below 5% as AI-generated reports flooded in. Each report took a seven-person team 30 minutes to three hours to review. Curl ended the bounty on January 31, 2026.

Compare that with the Glasswing open-source scan. Independent firms reviewed 1,752 of the high- or critical-rated findings, and more than 90% were confirmed as real. Even there, only about 62% held up at the severity the model assigned. The difference between the two stories is process, not magic:

  • Curl’s slop came from people pasting chatbot output into a form, unverified, chasing a payout.
  • The strong results came from systems that prove each finding. DARPA’s AIxCC scored teams on a working proof of vulnerability plus a patch, and penalized wrong answers. Trail of Bits’ Buttercup confirms each bug and checks each fix before reporting it.

Academic evaluations point the same way. Plain LLM prompting often flags already-patched code as vulnerable, with precision close to random guessing in pairwise tests. LLMs do much better as a triage layer on top of a conventional scanner. One study found agents cut SAST false positives on the OWASP Benchmark from over 92% to 6.3% in the best setup.

The practical rule: an AI finding is a hypothesis until it comes with a reproduction. No proof of concept, failing test, or traced data path means no ticket.

Does AI understand the architecture like your architects do?

Not yet, and mostly not because it lacks intelligence. It lacks the context. An AI reviewer sees the code it is given. Your architects know things that are written down nowhere in the repository:

  • which network paths reach which services, and which ingress bypasses the gateway
  • which fields are regulated data, and what the retention rules are
  • what the business considers a legitimate action (can a customer refund more than they paid?)
  • which “temporary” exception was approved in 2023, and why it still exists
  • who the realistic attacker is, and what they want

The hardest bugs live exactly there. Missing authorization (CWE-862) sits at #4 in the 2025 CWE Top 25, and automated tools still struggle with it. Consider this endpoint:

@GetMapping("/invoices/{id}")
public Invoice getInvoice(@PathVariable Long id, Principal user) {
    return invoiceRepository.findById(id).orElseThrow();
}

The code is clean. There is no injection, no unsafe deserialization, nothing a pattern matcher dislikes. Whether it is a critical flaw depends entirely on the architecture. If a gateway filter enforces tenant ownership, it is fine. If a second route, such as a partner API or an internal admin tool, reaches the same service without that filter, any user can read any tenant’s invoices by counting upward.

An AI reviewing this file alone faces two bad options. It can flag every lookup-by-ID as a possible insecure direct object reference (IDOR), which buries the team in noise. Or it can stay quiet and miss the one that matters. Only someone who knows the routing, the tenancy model and the trust boundaries can say which answer is correct. A 2026 benchmark found that frontier models share blind spots on exactly these adjacent authorization weaknesses.

AI can help architects here. It can map endpoints, draft data-flow diagrams, and list every route that reaches a service. It still needs a human to say what the system is supposed to allow.

Should developers still be taught secure coding standards?

Yes, and AI makes the case stronger. The bottleneck is now verifying and fixing findings. The people doing that work are developers. A developer who does not understand why a finding matters cannot tell a real bug from a hallucination, cannot write a correct fix, and cannot notice when an AI-proposed patch just moves the problem somewhere else.

There is a second reason. Much of the code being reviewed is now written by AI. Veracode tested more than 100 models on security-relevant coding tasks. In its 2025 report, 45% of cases introduced a vulnerability, and newer or larger models did no better. The 2026 update put the average pass rate at 56%, barely changed. Cross-site scripting and log injection failed in most relevant samples. A developer who accepts that output without the knowledge to judge it is shipping the flaw personally, with extra steps.

Training does need to change shape, though. A week-long annual course that developers forget by Friday was never effective. What works better:

  • Short, language-specific standards that fit on a page: “never use ObjectInputStream on untrusted input; use a schema-validated JSON parser instead,” not a 200-page policy.
  • Teaching through the team’s own findings. Every confirmed vulnerability, whether found by AI or a human, becomes a 15-minute walkthrough of the root cause.
  • Secure defaults over memorization. Frameworks, libraries and templates that make the safe path the easy path, such as parameterized queries by default and auto-escaping templates.
  • Reviewing AI output as a taught skill. Developers learn what AI tends to get wrong, so they know where to look.

Should SAST and DAST run throughout development?

Yes. They are cheap, repeatable and deterministic, and those are exactly the qualities AI lacks. Static analysis (SAST) reads source code. Dynamic testing (DAST) attacks the running application. They catch different things, so each belongs at a different point.

StageControlGood at catchingWeak at
In the editorSAST linting, secrets detectionInjection patterns, dangerous APIs, hard-coded keysAnything that spans files or services
Pull requestIncremental SAST, dependency scanning (SCA), AI review of the diffNew flaws in changed code, vulnerable librariesBusiness logic, deployment config
Nightly / main branchFull SAST, AI-agent deep scan with proof of conceptCross-file data flows, variant analysis of past bugsRuntime behavior
StagingDAST, API fuzzingAuth flaws visible over HTTP, misconfigured headers, real exploitabilityCode paths the scanner never reaches
Before major releaseThreat-model review, manual pen testArchitecture, trust boundaries, abuse of legitimate featuresScale; it is slow and expensive
ProductionMonitoring, bug bounty, AI scanning of new commitsWhat everything above missedPrevention; the bug already shipped

AI fits best in two places on this map. First, as a triage layer on SAST output, where it removes most of the noise that makes developers ignore scanners. Second, as a deep-scan agent that must prove each finding with a reproduction. It fits worst as the only gate, placed at the end.

One caution about speed: have the pull-request gate block only on new critical and high findings. The nightly full scan tracks older debt separately, so it doesn’t block every merge.

Recommendations

For developers

  1. Learn the top weaknesses for your stack. Start with the CWE Top 25 and the language-specific rules in your team standard.
  2. Treat AI-generated code like a pull request from a fast, confident junior: read it, test it, and question anything touching auth, crypto, input handling or deserialization.
  3. When an AI flags something, ask for the proof: the input, the path, the failing test. If it can’t produce one, verify it yourself before filing it.
  4. When you fix a finding, fix the class. Search for sibling instances of the same pattern before closing the ticket.

For security and platform teams

  1. Put SAST, secrets detection and dependency scanning in the editor and the PR, not only in the release pipeline.
  2. Run DAST and API fuzzing against every staging deployment, with authenticated scans that test cross-tenant access.
  3. Use AI as a triage layer and a deep scanner, and require a reproduction before any AI finding becomes a ticket.
  4. Feed the AI your architecture: route maps, trust boundaries, data classifications and threat models. Context-free review produces context-free findings.
  5. Keep humans on authorization, business logic and design review. Schedule threat modeling for every new service and every new external entry point.
  6. Budget for the fix side. If AI multiplies findings tenfold, remediation capacity has to grow too, or the backlog becomes the risk.

For leadership

  1. Stop measuring security by bugs patched. Track instead:
    • share of vulnerabilities caught before merge versus after release
    • mean time to remediate, by severity
    • repeat rate: how often the same weakness class comes back
    • share of critical services with a current threat model
  2. Fund training and secure-by-default tooling as part of feature work, not as a separate line item that loses every prioritization meeting.
  3. Assume attackers have the same AI you do. The advantage goes to whoever finds and fixes first, and the cheapest place to win is before the code ships.

The bottom line

AI has changed vulnerability detection permanently. It finds real bugs, sometimes bugs that decades of human review and fuzzing missed. It does not know what your system is supposed to do. It cannot tell your business rules from your bugs. It is only as trustworthy as the verification wrapped around it. The satire’s dashboard of 1,000 patched bugs is not a success metric. It is 1,000 things that could have been prevented, found by a tool your adversaries also own.

Teach the developers. Run the scanners early and often. Let the AI dig deeper and faster than people can, and make it show its work. Keep the architects in the room. The robot with the mop is useful. It is still better not to spill.