blog
DEVOPS

How Arnica Governs AI-Generated Code Repo-Wide October 2026

Posted 
October 9, 2026
|
0
 min
Illustration of AI-generated code streaming from a neural network toward connected repositories protected by a security shield

If your security tooling triggers on code pushes and relies on developers having plugins installed, you're already working with a model that doesn't account for cloud-running agents. Those agents commit under bot identities, generate complete PRs in one pass, and often self-validate before any human sees the result. There's a better way to get coverage across every repo, regardless of how or where the code was written.

TLDR:

  • AI-assisted developers introduce security findings at 10x the rate of peers, and 45% of AI-generated code samples contain OWASP Top 10 vulnerabilities
  • IDE plugins and pipeline scanners can't govern cloud-running agents: they produce complete draft PRs with no workstation, no IDE session, and no incremental intercept point
  • Security policy must live inside agent config files (.cursor/rules/, CLAUDE.md, .github/copilot-instructions.md) so every agent run reads it regardless of where execution happens
  • Route findings by identity graph, not git author: 82% of findings trace back to developers no longer at the company, and cloud-agent commits land under bot identities with no human owner
  • Arnica connects through your SCM to scan 100% of repos from day one, writing security rules directly into agent config files and tracing multi-file vulnerabilities that single-file SAST misses

The Scale of the AI Code Security Problem

AI coding tools have delivered a genuine productivity jump. The problem is what comes with it.

Empirical research across Fortune 50 enterprises found that AI-assisted developers produce commits at three to four times the rate of their peers but introduce security findings at 10 times the rate. Remediation backlogs grow faster than security teams can close them, and the volume only increases as agent adoption scales.

The code quality picture is equally concerning. Veracode tested over 100 LLMs on security-sensitive coding tasks and found that 45% of AI-generated code samples introduce OWASP Top 10 vulnerabilities, a rate the CSA AI-Generated Code Vulnerability Surge research note reports as unchanged across multiple testing cycles spanning 2025 through early 2026. A separate CSA survey of over 1,500 security leaders found that 92% are concerned about AI agent security implications, yet most organizations report deep gaps in governance coverage.

Ad-hoc tooling does not answer the dangers of vibe coding. Point scanners catch issues after code is written. IDE plugins reach maybe 20% of developers. Neither was built for a world where agents generate full draft PRs on remote infrastructure with no workstation in the loop. The math demands repo-wide governance, not individual developer controls.

Why AI-Generated Code Introduces Vulnerabilities

AI models generate code by predicting the next plausible token sequence, not by reasoning about trust boundaries or attack surfaces. A model trained on millions of public repositories inherits every insecure pattern those repositories contain, and reproduces them confidently.

A dark, dramatic digital illustration showing a glowing neural network or AI brain at the center, with streams of code flowing outward from it. Hidden within the code streams are subtle visual cues of danger — faint red warning glows, cracks in the code lines, shadowy lock icons breaking apart. The overall aesthetic is deep navy blue and dark teal with red and orange accent highlights suggesting embedded vulnerabilities. Abstract, cinematic, no text or letters anywhere.

Three root causes drive most AI-generated vulnerabilities:

  • Pattern completion over security reasoning: when a model fills in an authentication flow, it matches patterns from training data. If insecure patterns were common in that data, they surface in the output. The model has no security intent.
  • Hallucinated dependencies: approximately 20% of AI-generated code samples reference packages that do not exist. Attackers register those names, a technique called slopsquatting, and anyone who installs the suggested dependency pulls in malicious code.
  • Insecure-by-default patterns: hardcoded credentials, missing input validation, client-side authorization checks, and SSRF-prone API calls appear repeatedly because they were common in training corpora.

The agent model compounds each of these. Traditional development involved incremental commits, giving security tools multiple intercept points. AI coding agents deliver complete draft PRs in a single pass, often self-validating in a sandbox before any human sees the result. That collapses the review window considerably, and tools built around the incremental commit model miss it entirely.

How AI Coding Agents Changed the Development Model

AppSec for AI coding agents starts with understanding the old model: human developers commit code in steps. A function here, a fix there, a comment updated before lunch. Each push is an intercept point, and traditional AppSec tools were designed around exactly that rhythm.

AI agents work on a different model entirely. Cursor, GitHub Copilot, and Claude Code generate a complete feature or bug fix in one pass, self-validate it in a sandbox, and deliver a full draft PR before any human has looked at a single line. The scanning trigger that legacy tools depend on arrives all at once or never arrives in the traditional sense at all.

Cloud-running agents widen this gap further. When an agent executes remotely, there is no workstation in the loop, no IDE plugin to intercept output, no incremental save event to trigger a scan. The only surfaces left are the PR itself and whatever rule context the agent read during generation. Tools built assuming a human initiates code in steps face a structural mismatch here, not a configuration gap that a tuned policy can close.

The Most Dangerous Vulnerability Classes in AI-Generated Code

The table below maps the six vulnerability classes most commonly found in AI-generated code to the specific reason AI agents amplify each one.

Vulnerability ClassWhy AI Agents Make It Worse
Injection flaws (SQLi, command injection)Pattern-matched from training data; model has no concept of trust boundary
SSRFAI-generated API integrations frequently omit server-side request validation
Hardcoded secretsCredentials appear in training corpora; agents reproduce them without flagging risk
Client-side auth bypassAuthorization checks land on the client because that pattern was common in training data
Slopsquatting~20% of AI-generated code references packages that do not exist, creating exploit-ready name targets
Multi-file context-dependent risksLogic spread across files that single-file scanners cannot trace end-to-end

The OWASP Top 10 for Agentic Applications 2026 now formally catalogs ten agent-specific risk categories, including Agent Goal Hijack, Tool Misuse, and Memory and Context Poisoning. Security teams scoping their threat model to human-written code patterns are missing this surface entirely.

Multi-agent AppSec risks deserve particular attention. An agent generating a feature across several files may sanitize input in one file, re-fetch it from cache in another, and apply an authorization check gated by a feature flag in a third. No individual file looks dangerous. The vulnerability only exists in the combination, which is exactly what rule-based, single-file SAST misses.

Why Traditional Shift-Left Tools Cannot Govern AI Agents

IDE plugins were the dominant shift-left investment of the last decade. The problem is they depend on a developer sitting in an IDE, saving files, and pushing code incrementally. That assumption fails completely when a Copilot Agent or Cursor cloud agent generates and commits code on remote infrastructure. There is no workstation. There is no IDE session. The plugin has nothing to intercept.

Adoption was already a problem before agents entered the picture. Industry-average IDE plugin adoption sits around 20%, meaning most repositories were ungoverned even under the best conditions. Agents drop that number to zero for every commit they author.

Pipeline-dependent scanning has a related failure mode. These tools trigger on a code push event, run checks inside the CI build, and return findings. When an agent delivers a complete 40-file draft PR in a single event, the pipeline sees one large artifact instead of a series of checkpoints. Findings arrive after the full feature is already written, making remediation a rework cycle instead of a course correction.

The deeper issue is that code-push scanning as a reliable intercept point is eroding. Agents that self-validate in a sandbox before opening a PR can bypass push-triggered scans entirely if the validation environment does not include security tooling. The event the scanner was waiting for either never comes, or arrives too late to shape what the agent produced.

What Securing AI-Generated Code Across All Repos Actually Requires

Detection-first tools scan what already exists. Governance-first tools shape what gets written. That distinction is the starting point for any clear-eyed evaluation of AI code security.

A detection-first approach catches issues after an agent has produced a complete draft PR, and findings arrive as a remediation queue the team works backward through, a core theme in agentic AI security. A governance-first approach writes security policy into the agent's generation context before a single line is committed, so by the time the PR exists, the policy has already run.

Both layers are necessary. Governance without detection leaves gaps for novel risks no rule anticipated. Detection without governance means every AI-generated PR arrives as a full remediation problem.

Coverage is where most programs fail before either layer can work. IDE plugin-based controls reach roughly 20% of developers under ideal conditions, and zero percent of cloud-running agents. For a program to govern AI-generated code across all repos, coverage must equal SCM coverage, not workstation coverage, which requires connecting through the SCM event stream instead of relying on developer-installed tooling.

Three lifecycle phases require distinct controls:

  • Before code is written: security policy must live inside the agent's configuration files, in the repository itself, so every agent run reads and applies it regardless of where execution happens.
  • As code is generated: AI SAST must trigger on push and PR events, including draft PRs that agents open before any human review begins.
  • After code reaches PR stage: findings need triage, ownership attribution, and routing to the developer currently active in that codebase, not the original author who may have left.

How to Route Security Findings to the Right Developer

Routing a security finding to a developer who left six months ago means nobody fixes it. Per Arnica's internal benchmark, 82% of findings were authored by developers no longer at the company, which explains why most backlogs stagnate regardless of scanning quality.

A dark-themed digital network diagram showing an identity resolution system. At the center, a glowing node represents a unified developer identity. Connected around it via luminous lines are multiple identity sources: a bot/robot icon representing an AI agent, a corporate email badge, a messaging platform icon, a code repository symbol, and a faded ghosted silhouette representing a departed employee. From the center node, a bright arrow routes toward an active human figure highlighted in blue-green. The overall aesthetic is deep navy and dark teal with blue and amber accent glows, abstract and cinematic. No text, no letters, no words anywhere in the image.

Cloud-agent commits make this worse. When GitHub Copilot Agent or a Cursor cloud agent opens a PR, the git author is a bot identity. Any routing mechanism anchored on git author email hits a dead end.

Solving this requires an identity graph. CODEOWNERS names team aliases and individuals who frequently rotate off; neither survives attrition. An identity graph maps SCM identity, Slack or Teams identity, personal versus corporate email, and agent-bot identity back to one resolved human, kept current from actual repository activity.

Two cases need explicit handling:

  • Cloud-agent commits: the bot identity must be traced back to the human who dispatched the agent so agent-authored code still has a human owner.
  • Departed authors: routing falls to a product-level security champion, a developer currently active in that repo who can fix, review, and merge the change.

Delivery also needs to match finding type:

  • A hardcoded secret warrants a direct Slack or Teams message plus automated source-code-layer remediation, part of a broader effort to prevent exposed secrets in AI coding assistants.
  • A dependency finding should arrive as a version-upgrade suggestion.
  • A SAST or IaC finding lands as an in-chat explanation with a code suggestion on the PR itself.

Same routing logic, different output format based on what the developer actually needs to act.

How to Classify Pull Requests by Risk Level

Not every file in a pull request carries equal risk. A test fixture update and an authentication middleware change are both "code changes" in the same PR, but treating them identically is what creates reviewer overload.

Risk classification works by scoring each file against signals that actually predict downstream impact:

  • Security relevance: does the file touch authentication, authorization, cryptography, input handling, or secrets management?
  • Code complexity: how many branches, dependencies, and call sites does the changed logic reach?
  • Changed surface area: is this a two-line fix or a full rewrite of a core module?
  • Dependency modifications: are package versions changing, and if so, what is the reachability exposure?

Arnica Code Review uses Focus, Look, and Skim classifications to sort files into tiers before any human sees them. Focus files require full reviewer attention, Look files merit a check, and Skim files are routine changes safe to approve quickly. Based on Arnica's internal data, roughly 70% of files in a typical PR land in Skim, so reviewers concentrate effort on the fraction that actually ships risk.

The alternative is undifferentiated comment flooding, where a bot surfaces observations on every file regardless of risk level and the reviewer's job becomes triaging the tool instead of reviewing the code. That cognitive overhead compounds as AI agents increase PR volume. Staged classification inverts this: safe files clear automatically, and human judgment goes where it cannot be replaced.

Regulatory and Compliance Requirements for AI Code Security

The compliance picture for AI-generated code hardened in 2026. The EU AI Act introduced explicit human-in-the-loop requirements for high-risk AI systems. In Arnica's view, those requirements are likely to extend to AI coding agents operating in governed software environments as enforcement guidance matures. That means documenting that a human reviewed AI-generated code before it shipped, and not merely a checkbox that a scanner ran.

Existing frameworks apply unevenly. NIST AI RMF 1.0 and ISO/IEC 42001:2023 were architected before autonomous agents existed, so they cover AI risk at a product level but say little about agent-generated source code. NIST SP 800-218 and Executive Order 14028 create SBOM requirements that governed FinTech and HealthTech industries must satisfy, but static SBOMs generated at release time miss dependencies introduced by agents mid-sprint.

According to the CSA State of AI Cybersecurity 2026 survey of over 1,500 security leaders, 92% of organizations are concerned about AI agent security implications, yet most report material gaps in AI security governance.

What a compliance team needs demonstrable in an audit comes down to four artifacts:

  • A per-PR record of what AI generated and what a human reviewed, so auditors can verify oversight was real and traceable.
  • A continuous SBOM updated as dependencies change, not snapshotted at release, so agent-introduced packages don't fall outside the audit window.
  • Evidence that security policy was applied during agent code generation, not retrofitted after the fact.
  • A traceable chain from finding to remediation with timestamps and owner attribution for every issue flagged.

Building an AppSec Program for AI-Generated Code

Arnica approaches agentic development security, a central concern of AI code security, as a sequenced program, not a collection of independent controls. The order in which you deploy each layer determines whether the program actually holds.

  • Phase 1: Build repo-wide visibility. Connect to your SCM and confirm scanning covers 100% of repositories, including those without CI/CD pipelines configured. Every ungoverned repo is an active blind spot. That requires pipelineless coverage driven by SCM events, not scanner configs someone has to maintain per repo.
  • Phase 2: Add governance at the generation layer. Scanning after the fact is a second line of defense. Security policy needs to live inside each AI coding agent's configuration files, in the repository itself, so agents read it during generation regardless of where they run.
  • Phase 3: Integrate risk-based triage into PR review. Every PR should arrive pre-classified. Reviewers need to know which files require full attention and which are safe to approve quickly before they open the diff. Without this, increasing PR volume from agents translates directly into reviewer overload.
  • Phase 4: Build the feedback loop. Findings that developers consistently dismiss signal either a tuning problem or a real risk being ignored. The program needs a structured way to capture that signal and route it to security before any model behavior changes.

The minimum viable baseline is coverage plus scanning. A fully mature program adds governance at generation, identity-aware attribution, compliance evidence per PR, and a learning loop that keeps detection quality improving as agent tooling evolves.

How Arnica Governs and Secures AI-Generated Code Across All Repos

Arnica connects through your SCM and begins scanning 100% of repositories from day one, no CI/CD pipeline changes and no IDE plugins to deploy. Coverage equals repo coverage, not workstation coverage, which is the only model that holds when cloud agents are generating and committing code with no developer machine in the loop.

Here is how each capability maps to a specific failure mode.

The Agentic Rules Enforcer

Rules are written directly into each agent's configuration files: .cursor/rules/ for Cursor, CLAUDE.md for Claude Code, .github/copilot-instructions.md for GitHub Copilot, GEMINI.md for Gemini, and .augment/rules/ for Augment. Because the rules live in the repository itself, they govern every agent run regardless of where execution happens. If a developer removes the block, the next qualifying PR restores it automatically. The default pack maps to OWASP ASVS Level 2, and agents cite the rule ID inline so security teams can verify that governed controls actually landed in committed code.

AI SAST With Multi-File Reasoning

Arnica's AI SAST performs semantic, multi-file reasoning, not pattern matching against fixed signatures. A vulnerability that lives across three files, sanitized in one, cached and re-routed in a second, and conditionally authorized in a third, is the kind of risk that single-file rule packs cannot surface. Arnica traces data flow and logic across the full file graph to find it.

Identity Resolution and Routing

Per Arnica's internal benchmark, 82% of findings were authored by developers no longer at the company. Cloud-agent commits compound this by landing under bot identities with no human git author. Arnica's identity graph resolves the currently active human behind every commit, re-attributing agent-bot commits to the developer who dispatched the agent and routing to a product-level security champion when the original author has left.

Adaptive Backlog Management

When a CVE crosses into the CISA KEV catalog or an EPSS score moves materially, Arnica re-routes the finding through the identity graph to the current owner and restarts the SLA timer. A finding that scored Medium at discovery does not stay dormant once it becomes actively exploited.

The Developer Feedback Loop

When developers consistently dismiss a finding category, Arnie surfaces a proposed rule refinement for security-operator review. Nothing changes until a human approves it, so developer signal informs tuning without handing developers control over what the scanner learns.

Arnica was named a Representative Vendor in the Gartner Hype Cycle for Platform Engineering 2026 under Software Supply Chain Security, included in the Forrester Agentic Development Security Tools Market Overview Q2 2026, and designated "Best for Mid-Market" by Latio Tech in their 2026 AppSec Market Report, the only vendor among the six reviewed to carry that specific label.

"Your board is going to ask you where AI is being used in development and whether it is secure. Arnica gives you the inventory and the proof, without requiring your developers to change how they work." -- Arnica CISO messaging guide

Final thoughts on AI Code Security and Repo-Wide Governance

The productivity gains from AI coding agents are real, but so is the 10x increase in security findings that comes with them. Point tools and IDE plugins were not built for a world where agents open 40-file PRs on remote infrastructure with no human in the loop. Coverage has to equal SCM coverage, not workstation coverage, or the whole program has blind spots by design. Get started for free and connect your repos to find out what's currently ungoverned.

FAQ

How do I route security findings to the right developer when the original author has left the company?

Routing to a departed author means the finding sits unfixed. An identity graph solves this by mapping SCM identity, Slack or Teams identity, and agent-bot identity to one resolved human kept current from real repository activity. When the original author is gone, the finding routes to a product-level security champion (a developer currently active in that codebase) and not to a stale CODEOWNERS alias or an empty inbox.

How do I automatically re-attribute a security finding when the original code author was an AI agent, not a human?

Standard git author attribution breaks when a Copilot Agent or Cursor cloud agent opens a PR under a bot identity. Arnica's identity graph traces the bot identity back to the human who dispatched the agent, so agent-authored commits still carry a human owner for routing and SLA purposes. Without this re-attribution step, any routing mechanism anchored on git author email reaches a dead end on every cloud-agent commit.

How do I classify pull requests by risk level so reviewers only focus on files that actually need attention?

Arnica Code Review classifies each file in a PR into Focus, Look, or Skim tiers before any human opens the diff, using signals including security relevance, code complexity, changed surface area, and dependency modifications. Roughly 70% of files in a typical PR land in Skim, so reviewer attention goes to the fraction that actually carries risk. Without pre-classification, increasing AI-generated PR volume translates directly into reviewer overload as every file competes equally for attention.

What is agentic rules enforcement, and how does it differ from IDE plugin-based security controls?

Agentic rules enforcement writes security policy directly into each AI coding agent's configuration files in the repository itself (.cursor/rules/ for Cursor, CLAUDE.md for Claude Code, .github/copilot-instructions.md for GitHub Copilot), so the policy runs during code generation regardless of where the agent executes. IDE plugins require a developer at a workstation, which means they reach roughly 20% of developers under ideal conditions and zero percent of cloud-running agents. Rules that live in the repo govern every agent run by default.

How should a CISO assess AppSec platforms for agentic development governance when most PRs are AI-generated?

Three criteria separate platforms that hold in an agentic environment from those that do not: coverage through the SCM instead of through developer-installed tooling, so repo coverage equals governance coverage; a governance layer that operates before code is written, and also after a PR exists; and identity-aware routing that handles cloud-agent commits and departed authors without falling back to CODEOWNERS. Detection-first scanning alone covers the output but not the generation step, which is where most AI-introduced vulnerabilities are fixed most cheaply.

Reduce Risk and Accelerate Velocity

Integrate Arnica ChatOps with your development workflow to eliminate risks before they ever reach production.  

Try Arnica