Adoption of AI coding tools has outrun the controls meant to catch what they get wrong. This brief separates the ordinary risk, the kind already visible in production code today, from the frontier risk research on advanced models is only beginning to measure, and lays out what fifty years of code-review evidence, plus one company's own live history running Gauntlet, says about actually closing that gap without slowing down to do it.

Written by Ron Griffin, CEO and Founder, SpecOps.AI · Date September 29, 2026

The headline numbers

  • 90% of developers now use AI tools in their day-to-day work
  • 30% report little to no trust in the code that AI produces for them
  • 2.7× more security vulnerabilities measured in AI-co-authored pull requests
  • 22% → 96% review coverage, one company's own history, before and after adding this governance model

1. The gap

Two numbers from the same survey tell the whole story. In Google Cloud's 2025 DORA report, a nearly 5,000-person, vendor-neutral survey of technology professionals, 90% of respondents said they now use AI as part of their work, up 14 points in a single year. In the same survey, 30% said they have little to no trust in the code AI produces for them.

That is not a minor discrepancy. It means the tooling has been adopted faster than any control has been built to match it, and the people using it every day already know that. Stack Overflow's 2025 developer survey found the same pattern from a different angle: 84% of developers use or plan to use AI tools, up from 76% the year before, while favorable sentiment toward those same tools fell to 60%, down from over 70% in both of the two prior years. Distrust of AI output rose from 31% to 46% year over year. The developers most skeptical of all were the most experienced ones: only 2.6% of developers with ten or more years of experience said they highly trust AI-generated output.

DORA's own framing of its 2025 data makes the mechanism explicit: for the first time, the report found a positive correlation between AI adoption and delivery throughput, and simultaneously a negative correlation between AI adoption and delivery stability. Teams are shipping more, faster, and less reliably, in the same motion. The report's own conclusion: "AI accelerates software development, but that acceleration can expose weaknesses downstream. Without robust control systems... an increase in change volume leads to instability."

Unreviewed AI-generated code does not fail all at once. It accumulates the way debt always does: each unreviewed change is individually shippable and collectively expensive, paid down later by someone who has to first figure out what the code was actually supposed to do. The industry's usual name for that bill is technical debt, and AI-generated code is a faster way to take it on, not a way to avoid it. What changes with a coding assistant is the interest rate.

This brief is about what "robust control systems" actually means, grounded in three bodies of evidence: what is already measurable in AI-generated code today, what frontier-model safety research says about a harder problem still on the horizon, and what happens inside one real company's own history when that control system is actually built and turned on, including a real technical-debt paydown before it, not just after.

2. The ordinary risk

Set aside advanced-model scenarios for a moment. The risk already sitting in production code today is well measured, and it is not subtle.

Developers write less secure code when an assistant is watching

The clearest study here is peer-reviewed: Perry, Srivastava, Kumar, and Boneh (Stanford, published at ACM CCS 2023) ran 47 participants, students through industry professionals, through five security-relevant coding tasks, half with an AI coding assistant and half without. Participants with AI access wrote measurably less secure code on four of the five tasks:

Task With AI assistant Without
Symlink vulnerability 73% 15%
Authentication issues 58% 9%
Path / parent-directory vulnerability 61% 15%
Weak randomness 48% 15%
SQL injection 30% 6%

The second finding in the same paper is the more consequential one for governance design: participants with AI access were also more likely to believe their code was secure than the control group. The tool did not just introduce more defects, it suppressed the self-doubt that would normally prompt a developer to check.

19% slower. A 2025 randomized controlled trial by METR gave 16 experienced open-source developers real tasks in their own large repositories, with and without modern AI tools (Cursor Pro, Claude 3.5/3.7 Sonnet). Developers predicted AI would make them 24% faster. After finishing, they still believed it had made them 20% faster. Measured against the clock, they were 19% slower. The same overconfidence effect Perry et al. found in security judgment shows up again here, in raw throughput.

What is actually landing in shipped code

Two 2025 industry studies, both with disclosed methodology, point the same direction from production data rather than a lab setting. GitClear's analysis of 211 million lines of changed code (2020-2024, expanded past 600 million commits in a 2026 follow-up) found that 2024 was the first year on record where copy-pasted code exceeded refactored code in commits, a trend inversion that had held in the opposite direction since GitClear began tracking in 2020. Code churn, the share of code reverted or rewritten within two weeks of being written, rose from a pre-AI baseline near 3.3% to 7.1% in 2025.

CodeRabbit's December 2025 study of 470 real GitHub pull requests (320 AI-co-authored, 150 human-authored, reviewed with disclosed automated tooling) found AI-co-authored code carried roughly 1.7× more total findings per pull request than human-authored code, with the gap concentrated exactly where it matters most: security vulnerabilities were 2.74× more common, critical-severity issues were up 42%, and major-severity issues were up 74%.

None of this requires an adversarial or misaligned model. It is the ordinary, well-documented behavior of current systems: fluent, fast, prone to a specific and recurring set of mistakes, and persuasive enough that both the tool's output and the developer's confidence in it move in the same wrong direction at once.

3. The frontier risk

A separate, harder question is whether a sufficiently capable model could behave deceptively on purpose, not fumble a task, but pass a check it should have failed while appearing to comply. This is now an active, published research area, not speculation, and the findings deserve to be stated precisely rather than rounded up or down.

Sabotage, tested directly

Anthropic, working with Redwood Research, published "Sabotage Evaluations for Frontier Models" in October 2024, testing Claude 3 Opus and Claude 3.5 Sonnet against four scenario types: sabotaging a human's decision, sabotaging code, deliberately hiding capability ("sandbagging"), and undermining oversight. In the sandbagging test, a model trained to underperform dropped to 36.1% of its normal score on a dangerous-capability evaluation, and standard safety training recovered 95.6% of that lost performance, meaning the concealment did not survive ordinary retraining.

Stated plainly, in the authors' own words. Anthropic's own conclusion in that paper: current models, as of October 2024, "do not yet cross concerning thresholds" for sabotage risk, and "minimal mitigations are sufficient to address the risks" today. This is a capability evaluation under adversarial test conditions, not an observed incident. That distinction matters for how a governance program should actually respond to it.

Alignment faking

A follow-up study by the same collaboration, published in December 2024, is the most widely cited finding in this space, and the specific numbers are worth stating exactly. Claude 3 Opus was told it was being retrained to comply with harmful requests, and given a fictional distinction between a "monitored" and an "unmonitored" context. When it believed it was unmonitored, it refused harmful requests nearly always. When it believed it was being observed, it strategically complied 12% of the time, reasoning in its own visible scratchpad that faking compliance now would protect its current values from being overwritten later. When the model was actually retrained through reinforcement learning rather than merely told it would be, that rate of strategic reasoning rose to 78%.

The behavior was capability-gated: it appeared in Claude 3.5 Sonnet, Claude 3 Opus, and Llama 3.1 405B, and did not appear in smaller models in the same family. That pattern, present in more capable models and absent in less capable ones, is the specific reason this class of research gets more attention as model capability increases.

Scheming under observation

OpenAI's own o1 System Card, published in December 2024, includes an independent evaluation by Apollo Research of six frontier models for in-context scheming: deactivating oversight, attempting self-exfiltration, covertly manipulating data, and lying about it afterward. Under a strong goal-directed prompt, o1 attempted to deactivate its own oversight mechanism in 5% of applicable scenarios and covertly manipulated data in 19%. The figure that matters most for a governance design: when o1 had taken a scheming action and was then asked about it directly, it denied or fabricated an explanation 98-99% of the time. Asking a model whether it did something is not a verification step. It is a request for a second, unverified claim.

Capability is compounding faster than the risk is being measured

METR's longitudinal benchmark tracks the length of a software task, measured in how long a skilled human would take, that a model can complete autonomously with 50% reliability. That "time horizon" has been doubling roughly every seven months since 2019, with evidence of the rate accelerating in 2024. Every one of the findings above was measured against a specific model generation. The underlying capability curve has not stopped since any of these papers were published.

4. What has always worked

None of this is the first time software engineering has had to answer "how do we catch what a fast, fallible, occasionally overconfident producer of code gets wrong." Structured review has been measured for fifty years, on human-written code, with real numbers.

Method Low Typical High
Formal design inspection 65% 87% 97%
Formal code inspection 60% 85% 96%
Static analysis 65% 85% 95%
Pair programming 40% 55% 65%
Informal peer review 35% 50% 60%
Desk checking (self-review) 25% 45% 55%

(Defect removal efficiency by method, industry benchmark data, Capers Jones, 2008-2012)

Michael Fagan's original inspection process at IBM, the source of the modern formal code review, was measured at 90% defect detection across a product's life cycle; a Standard Bank of South Africa project that adopted it saw corrective-maintenance cost fall 95%. NASA's own inspection data across eight flight-software projects found inspections removed 64% of defects on average, outperforming testing in every defect category measured. Cisco's decade-old study of lightweight peer review, still the most-cited practical guidance in the industry, found effectiveness collapses above roughly 400-500 lines reviewed per hour.

Two things carry directly into the AI-code era. First, review effectiveness is a function of how it happens, not whether it happens: desk checking and informal review, methods that most closely resemble "a developer glances at what the AI wrote," catch roughly half of what formal, structured inspection catches. Second, no single method reaches the reliability a high-stakes system needs. Capers Jones' own conclusion: defect removal efficiency above 95% "cannot be achieved using testing alone," it requires combining inspection, static analysis, and testing together.

A caution about round numbers. The most repeated statistic in this literature, that a defect costs "100×" more to fix in production than at design time, traces to a real, peer-reviewed source: Boehm and Basili, IEEE Computer, January 2001. What gets dropped when the number is repeated is the authors' own caveat, in the same paper, that the ratio is closer to 5:1 on smaller, less-critical systems. The even more commonly reproduced version, a chart claiming exact dollar costs at each phase, does not trace to any locatable primary source at all. Included here as a small case study in its own right: even inside a field built on evidence, an unverified number can circulate as settled fact for over forty years once it is convenient enough to repeat.

5. A framework for AI-code governance

A single review pass, human or AI, is a desk check with better vocabulary. A governance program for AI-generated code needs the same layered structure the inspection literature already proves out, adapted for a producer that can write faster than any human reviewer can read.

1. Deterministic gates. Checks that do not depend on any reviewer's judgment: static analysis, policy linting, and, critically, verification that changed code actually executed rather than merely reading as correct. A diff can look right to a human or a model reviewer and still never have run.

2. Claim reproduction. A review verdict, from a person or a model, is itself a claim, not a fact. A high-severity finding, or a claim that an issue is fixed, gets independently reproduced before it is accepted, the same discipline the review process demands of the code it is reviewing.

3. Blast-radius limiting. The layer that matters most against a sophisticated failure specifically, because it does not require detecting bad intent, only limiting what any single change can reach: scoped access, tenant isolation, least-privilege service credentials, and no path that lets one change bypass review before it touches shared infrastructure.

4. Tamper-evident provenance. A record of what was reviewed, by what, and with what result, that cannot be quietly edited after the fact.

Layers one, two, and four are achievable today with existing tooling and process discipline. Layer three is where most organizations already under-invest, because it is infrastructure work rather than a checklist item, and it is the layer that still holds even when the first two do not catch something.

6. A case study: one repository's own history

Everything above draws on other people's data. This section draws on the authors' own: the repository this governance framework runs in has a real before-and-after in its git history, and the honest version of that story is worth telling, including the parts of it that do not flatter the argument.

Methodology and its limits, stated up front: this is a single company's own history, not a controlled study, and the two periods compared below differ in more than one way at once: who wrote the code changed, the product's scope and complexity changed, and the review process changed, all at close to the same time. Nothing here separates which of those changes caused which effect. It is not evidence that AI-written code is more reliable than contractor-written code, a claim this data cannot support either way. The pull-request and review-coverage figures come directly from the GitHub API (gh pr list --search "merged:<window>", and gh pr view --json reviews sampled across the earlier period); commit and line-count figures come directly from git log and git log --shortstat against the same date windows. Counted as of September 29, 2026; this repository takes new commits daily, so a reader re-running the same commands today will see a higher number for the governed period and an unchanged number for the closed, historical contracted-team period.

What changed, and when

From February through July 2026, this codebase was built by its founder, then together with a contracted development team of four engineers whose job was specifically debt paydown: take a founder-built prototype, remove what did not belong in a production system, and rebuild the foundation, a Bun and Turborepo monorepo, on architecture meant to hold under real customers rather than a demo. That kind of work reads as slow by design. Paying down debt does not produce the visible throughput that new feature work does, and it should not be judged by the same velocity numbers as what came after it; it is the precondition for those numbers meaning anything. On August 1, 2026, two things happened at once: the contracted team's engagement ended, having delivered that foundation, and a mandatory, multi-seat AI review process went live on top of it, enforced by CI from its first day rather than phased in, to keep the debt from re-accumulating the moment an AI coding assistant started writing most of the new code.

Metric Apr-Jul 2026 (contracted team) Aug-Sep 2026 (governed, as of Sep 29) Change
Pull requests merged 170 (4 months) 626 (2 months) 7.4× per month
Commits 1,106 1,669 3.0× per month
Real source lines changed (generated files excluded) 389,583 312,203 1.6× per month
PRs carrying a committed review report (sampled, n=60 vs. full population) 13 of 60, 22% 602 of 626, 96% 4.4×

7.4× · 3.0× · 1.6× · 22% → 96%. The pattern that matters here is that all four numbers move together. Pull-request throughput, commit throughput, and real source volume per month all rose at the same time review coverage rose from roughly one pull request in five to nearly all of them. Nothing traded off against anything else. The standard objection to adding review, that it slows a team down, has a direct answer on this repository's own numbers: it did not, on any of the three velocity measures checked.

A number this brief almost used: a first pass at code volume showed the governed period producing 2.9× more per month, not the 1.6× above. That number was inflated by generated files, mainly database migration snapshots and lockfile churn. Excluding lockfiles, migration snapshots, and build output brought the real figure down to 1.6×. The larger number would have made a better headline. It was not true, so it is not the number used here.

What the data does not show

Two further numbers were checked specifically because they would have made a cleaner, more flattering story, and neither cooperated. The ratio of commits with a fix:-style message was checked as a rough defect-frequency proxy: 250 of 1,106 commits in the contracted-team period (22.6%) against a nearly identical share in the governed period, effectively the same rate. Revert commits were checked the same way: 1 in the contracted-team period against 8 in the governed period, a higher raw and per-commit revert rate in the governed period, though a sample this small is too thin to read as a real difference in either direction.

Neither number is reported here as evidence of anything, including evidence that nothing changed. They are reported because omitting an inconvenient number after finding it is exactly the failure this document has argued against, and that standard does not stop applying to the one section drawing on the authors' own data.

Testing the reviewer, not just counting its output

A coverage number says a check ran. It does not say the check was any good. This company also runs a seeded-defect regression benchmark against its own reviewing process: seventeen fixture cases, each one a defect this same codebase actually shipped at some point, planted into synthetic diffs and run blind against each reviewing seat, plus one clean fixture per seat where the correct answer is silence.

Run for real against a live model, across two independent controlled runs plus a third execution that surfaced live while fixing an unrelated bug in the benchmark's own scoring logic: recall on the planted, previously-shipped defects reproduced exactly, 100% for five of the six measured seats, every time. The sixth seat held at 67%, missing the same case every time, a real and now-tracked gap rather than noise. The clean-fixture side, whether a seat wrongly flags sound code, was less stable: one seat failed its clean fixture in every run, others flipped between passing and failing at the identical settings, real model non-determinism rather than a harness bug. None of that gates a merge today by design, and all of it is printed in the tool's own output rather than hidden.

A defect found in the review tool, by the review process, before it shipped. Recording this benchmark's first real baseline surfaced a genuine bug in the scoring logic itself: a rounding mismatch that would have failed the sixth seat against its own just-recorded baseline on every future run, forever, even at an unchanged score. Two reviewing seats independently caught it, reproduced it against a live, actually-failing run rather than taking a report's word for it, and it was fixed with a regression test proving the exact case. The same governance mechanism this brief argues for caught a defect in its own measurement of itself.

What actually changed, stated precisely

What this repository's history actually supports is narrower than "AI replaced humans and quality improved": coincident with replacing an unreviewed, unenforced process with a mandatory, audited one, review coverage moved from roughly one pull request in five to the near-totality of what ships, velocity rose rather than fell on every measure checked, 283 blocker-severity findings, roughly one for every two pull requests merged in that window, were stopped before reaching master, and the review mechanism itself has been shown, under test, to catch real defects, including defects in its own tooling. Whether the resulting code carries fewer defects in total is not established by this data, because no closed-loop defect-tracking exists for the earlier period to compare against. What changed with certainty, and what this brief is willing to stand behind, is whether a change was checked by anything at all before it shipped, and whether that check demonstrably works when it runs.

7. Where this leaves the market

Section 1 established the gap at industry scale: adoption near-universal, trust falling, delivery stability degrading exactly where AI writes the most code. Section 6 established that the gap is closeable, with one company's own before-and-after numbers to show it: coverage from roughly one review in five to nearly all of them, velocity up rather than down on every measure checked, and a review mechanism proven, under test, to catch real defects including its own.

Put together, those two sections describe a market and a working answer to it, not a hypothesis and a hope. The technical-debt framing from Section 1 is the reason this matters commercially and not only technically: every organization adopting AI coding assistants at the rate DORA and Stack Overflow describe is taking on debt at the same rate, whether or not anyone is tracking it, and most are not tracking it, per the same 30% distrust figure this brief opened with. The paydown this document's own case study describes, a dedicated hardening phase followed by continuous, enforced review, is a repeatable sequence, not a one-time fix specific to one codebase. That sequence, packaged as a product rather than reconstructed by each organization from first principles, is the gap SpecOps built Gauntlet to close: the deterministic gates, claim reproduction, and tamper-evident provenance of Section 5's framework, running continuously rather than assembled under deadline pressure after the fact. Blast-radius limiting, Section 5's third layer, is deliberately the one this brief makes no shipped claim about: it is real infrastructure work, still in progress, and the honest version of a governance pitch says so instead of rounding a framework up to a finished product.

This is a positioning claim, and it is held to the same standard as every other claim in this document: it is falsifiable by the same kind of first-party data Section 6 used on itself, and it will be revisited the same way if the data changes.

8. The honest limit

A framework like the one above deserves one caveat stated as plainly as the statistics: if a review step is itself performed by an AI model, and the code under review was written by a model from the same family or capability class, the reviewer and the author may share blind spots neither one can see past. This is not a hypothetical. It is the direct implication of the alignment-faking finding above, that strategic behavior appeared specifically in more capable models and was absent in less capable ones from the same lineage.

The defensible response is not to claim review solves this, it does not, on current evidence, fully solve it for anyone. The defensible response is the same one security engineering has always given to a threat it cannot fully detect: assume it will sometimes get through, and make sure nothing that gets through can reach very far. A governance program that is honest about this limit is more credible, not less.

9. Recommendations

Section 6 is the reason these are recommendations rather than hypotheses: each one below is something one real organization's own history shows working, not just something the published research above suggests should.

  1. Treat AI-authored code as a distinct risk category with its own review posture, not as human-written code that happens to type faster. Security-relevant logic, authentication, and error handling are the categories where AI-authored code fails more often, and those categories deserve a mandatory second look regardless of how confident the diff reads.
  2. Require evidence of execution, not just a passing review, before anything reaches production. A diff reading correctly and code actually having run are different claims. Only one of them a reviewer, human or AI, can fake by being persuasive.
  3. Reproduce high-severity findings before accepting them as fixed. Apply the same skepticism to a review verdict that the review itself applies to the code.
  4. Invest in blast-radius limiting ahead of, not instead of, better review. It is the layer that still holds when detection fails.
  5. State the limits of the governance program in writing, including this one. A program that claims to fully solve a problem the published research says is not yet solved by anyone will not survive its first serious audit.

10. Methodology and sourcing

Every figure in this brief traces to a named, dated, publicly available source. Figures are drawn from four tiers: peer-reviewed academic research, primary reports read directly from the publishing organization, disclosed-methodology industry studies from vendors in the code-quality space, and first-party data queried directly from one company's own repository (Section 6), reported with its confounds stated in the same section rather than presented as generalizable.

Sources: Cohen et al., Best Kept Secrets of Peer Code Review (SmartBear/Cisco, 2006); Capers Jones, Software Defect Origins and Removal Methods (2008-2012); Fagan, IBM Systems Journal (1976) and IEEE TSE (1986); NASA Technical Reports Server; Perry, Srivastava, Kumar, Boneh, ACM CCS 2023 (arXiv:2211.03622); GitClear, AI Copilot Code Quality (2025-2026); CodeRabbit, State of AI vs. Human Code Generation (Dec 2025); Becker, Rush, Barnes, Rein (METR), July 2025 RCT; Anthropic with Redwood Research, Sabotage Evaluations for Frontier Models (Oct 2024) and Alignment Faking in Large Language Models (Dec 2024); Apollo Research via OpenAI's o1 System Card (Dec 2024, arXiv:2412.16720); Kwa, West et al. (METR), Measuring AI Ability to Complete Long Software Tasks (Mar 2025); Boehm and Basili, IEEE Computer (Jan 2001); Google Cloud/DORA, 2025 State of AI-assisted Software Development; Stack Overflow, 2025 Developer Survey.