AGENTCONN

Field report · · AgentConn Team

Four Horsemen of Agent PRs and How to Stop Them

Agent PRs fail four ways: context collapse, false confidence, review fatigue, and architectural drift. Data-backed survival strategies.

AI AgentsCode ReviewCoding AgentsPull RequestsDeveloper Workflow2026
Four horsemen silhouettes made of glowing code characters riding across a dark terminal diff view with red and green code changes

Four Horsemen of Agent PRs and How to Stop Them

Agent-generated pull requests grew from under 1% of GitHub PRs to 27.6% in just fourteen months. Anthropic’s 2026 Agentic Coding Trends Report puts the number higher: 41% of all new code is now AI-generated, with tools like Claude Code growing 6x in workplace adoption in under a year. The PR flood is here. The review infrastructure is not.

The problem is not that agents write bad code. It is that agent-authored PRs fail in ways human-authored PRs never did, and our entire review process was designed for a world where a human could explain their reasoning when asked. A forensic study of 33,000 agent-authored PRs found that agents achieve an 83.77% acceptance rate versus 91.01% for humans, but the gap hides the real story: the types of failures are fundamentally different. Agents do not make typos or forget semicolons. They produce syntactically correct, compilable code that violates contracts, breaks cross-service dependencies, and introduces architectural regressions that pass every test in the suite.

After months of tracking this data across academic research, industry reports, and practitioner experiences, we see four recurring failure modes that account for the overwhelming majority of agent PR disasters. We call them the Four Horsemen.


Horseman 1: Context Collapse

The agent knows the file. It does not know the system.

The first and most pervasive failure mode is context collapse: the agent generates code that is locally correct but globally wrong. It fixes the function but breaks the contract. It refactors the module but violates the architectural boundary. It adds the feature but duplicates logic that already exists three directories away.

FeatBit’s analysis of the 2026 productivity paradox identified four types of context loss in AI-generated pull requests: requirement context (what the business actually needs), codebase context (how the system fits together), review context (what previous reviewers flagged), and organizational context (team conventions, deployment constraints, implicit rules that live in people’s heads).

The MSR 2026 empirical study found that 23% of rejected agent PRs were duplicates, with agents frequently submitting PRs for issues already being addressed by another contributor. The agent had no idea someone else was working on the same problem. It had no concept of “someone else.”

Stack Overflow blog — Are Bugs and Incidents Inevitable with AI Coding Agents? Analysis of bug rates in AI-generated code

View original post on Stack Overflow Blog →

The Stack Overflow engineering blog quantified the damage: AI-generated code produces 1.7x as many bugs as human code, with logic and correctness errors running 1.75x higher and security findings 1.57x more common. But here is the insight most teams miss: these are not bugs of incompetence. They are bugs of context. The agent wrote correct code for the wrong problem because it could not see the full system.

The context collapse trap: Agents excel at small, well-defined changes. The MSR study found that about 28% of agent PRs merge almost instantly. The danger zone is anything that touches multiple services, crosses architectural boundaries, or requires understanding of business rules that live outside the codebase.

How teams are surviving it:

  • Contract tests at PR gates. If the agent’s change breaks an API contract, the PR fails CI before a human ever sees it. Tools like Pact and Specmatic catch cross-service breaks that unit tests miss.
  • Architecture fitness functions. ArchUnit-style assertions that codify your architectural boundaries. “No module in domain/ may import from infrastructure/” catches the boundary violations agents love to introduce.
  • Codebase-aware agents. Greptile showed that agents with deep codebase context (dependency graphs, not just file content) produce significantly fewer rework cycles. Codex sits at 5-6% rework versus the human baseline of 10%.

Horseman 2: The Confidence Paradox

It looks right. It reviews right. It breaks in production.

New Relic 2026 State of AI Coding Report — AI-generated code grades higher in review yet triggers rise in production incidents

View original report on New Relic →

This is the most insidious horseman, and the New Relic 2026 State of AI Coding report gave it a number: 94% of engineering leaders rate AI-generated code as higher quality than human code at the time of review. The code reads well. The variable names are descriptive. The comments are helpful. The structure is clean.

Then it hits production. 78% of those same respondents report more incidents once deployed. 82% experienced at least one production failure tied to AI-generated code in the past six months. 74% say at least a quarter of AI code needs significant rework within twelve months.

The paradox is structural: agents are optimized to produce code that looks correct to a reviewer scanning a diff. They have absorbed millions of examples of what “good code” looks like syntactically. What they have not absorbed is what “correct behavior” looks like under concurrent load, with stale caches, during a partial network partition, or when the third-party API returns a 429 instead of a 200.

As Addy Osmani wrote in what may be the most important essay on code review this year: “Code generation became cheap while understanding stayed expensive.” Agentic code review means reviewing code whose author cannot explain itself. Classic review validates a colleague’s reasoning. Agentic review must reconstruct reasoning that was never written down.

Addy Osmani blog — Agentic Code Review: reviewing code whose author cannot explain itself

View original post on addyosmani.com →

The New Relic number that should scare you: 62% of engineering teams now ship AI-generated code to production without line-by-line manual verification. They trust the confidence of the code itself. The confidence paradox means the code that looks most trustworthy is often the code that fails most surprisingly.

How teams are surviving it:

  • Property-based testing in CI. Instead of testing specific inputs and outputs (which agents are great at generating tests for), property-based tests define invariants the code must maintain. Hypothesis (Python) and fast-check (TypeScript) catch the edge cases agents miss.
  • Canary deployments as review. If the code looks right in review and passes tests, deploy it to 1% of traffic first. The production environment is a better reviewer than any human for the failure modes agents introduce.
  • Semantic diffing. Tools that show behavioral changes rather than textual changes. “This function now returns null for inputs > 1000” is more useful than “line 42 changed from >= to >.”

Horseman 3: Review Fatigue

The queue grows faster than the team can read.

The math is brutal. Faros AI’s telemetry data shows that AI adoption correlates with 98% more PRs that are 154% larger, while review times have grown 91%. Zero-review merges are up 31%. The review queue is now a conveyor belt moving faster than anyone can watch.

This is not a people problem. It is a structural problem. As Tian Pan wrote: “The dominant failure mode of code review in 2026 is that reviewer instincts that worked on human-authored PRs break down on agent PRs because the bugs cluster in different places and the artifacts the reviewer sees are no longer the artifacts that matter.”

Hacker News discussion — Agentic AI PRs sit in the review queue 5.3x longer than unassisted ones

View discussion on Hacker News →

When a human writes a PR, reviewers develop a sense for where bugs hide. They check the boundary conditions, the error paths, the off-by-one opportunities. When an agent writes a PR, the code is syntactically flawless. The bugs are in the assumptions, not the implementation. Reviewers have to develop entirely new instincts, and most have not had time to.

Theo’s response to ThePrimeagen’s skeptical takes on AI coding agents captures the tension well. ThePrimeagen’s core complaints, which Theo largely concedes, include vibe-coded slop PRs as a real and growing review burden, juniors who let the agent do everything never building intuition, and auto-merge tooling that ships without a human gate. Theo pushes back on one point: experienced engineers using agents as power tools genuinely hit a higher productivity ceiling. The question is whether the review infrastructure can keep up.

The comparative study of agentic PRs found another pattern that amplifies review fatigue: agents ghost when they receive subjective feedback. A human contributor adjusts their approach when a reviewer says “this doesn’t fit the project’s pattern.” An agent either ignores the feedback or generates an entirely new PR from scratch. 61.38% of agent-authored PRs carry no recorded review activity at all.

The review fatigue test: Count the unreviewed agent PRs in your org's last sprint. If it is over 20%, you have a structural problem, not a discipline problem. No amount of "review your PRs" Slack reminders will fix a conveyor belt moving faster than humans can watch.

How teams are surviving it:

  • Tiered review. Not all PRs deserve the same scrutiny. Automated checks handle formatting, dependency updates, and simple refactors. Human review focuses on business logic, API contracts, and cross-service changes. We covered this in depth in 10x PRs, 1x Reviewers.
  • AI-assisted review tools. CodeRabbit ($40M ARR, 700% YoY growth) and Greptile now handle the first pass across millions of PRs. The key insight: the AI that writes the code should not be the AI that reviews it.
  • Reviewability as a metric. If a PR cannot be understood in under 10 minutes, it is too big or too poorly structured, regardless of who authored it.

Hacker News discussion — Google Engineers Launch Sashiko for Agentic AI Code Review of the Linux Kernel, 111 points

View discussion on Hacker News →


Horseman 4: Architectural Drift

One PR is fine. A hundred PRs is a slow-motion rewrite.

The subtlest horseman does not show up in any single PR. It shows up over weeks and months as agent-authored changes silently shift the architecture. Each change is reasonable in isolation. Together, they constitute a drift that no one approved and no one noticed until the system became unmaintainable.

SoftwareSeni’s analysis documents how agents produce contract violations, cross-service dependency breaks, and architectural regressions that pass tests. The tests pass because each change is locally correct. The architecture degrades because no test asserts “the system as a whole still makes sense.”

This is the horseman that connects to the 48K files deletion incident we covered recently. Blast radius is the real infrastructure problem. When an agent can make hundreds of changes per day, the cumulative architectural impact dwarfs anything a human team would produce, because each individual change is too small to trigger alarm bells.

StarkRavingFinkle blog — Agentic Coding: Who Will Review All That Code? Analysis of zero-review merges increasing 31%

View original post on starkravingfinkle.org →

The Anthropic trends report notes that developers can “fully delegate” only 0-20% of tasks to agents, yet 41% of code is AI-generated. The gap is filled by supervision that is often cursory. As Andrej Karpathy warned, the coming “slopacolypse” is not code that fails immediately but code that is “almost right, but not quite,” degrading system quality gradually.

How teams are surviving it:

  • Architecture fitness functions. Automated assertions that run in CI: dependency rules, layering constraints, module coupling limits. If the agent’s PR increases coupling beyond the threshold, it fails.
  • Weekly architecture diff reviews. Not PR-by-PR review but periodic “zoom out” sessions where the team looks at how the system changed over the past week. Module dependency graphs and coupling metrics are your friends here.
  • Agent guardrails. As we covered in The Anti-Slop Linter for Your Coding Agent, quality gates that run on every agent build are the cheapest form of architectural protection.

Hacker News discussion — ProofShot: Give AI coding agents eyes to verify the UI they build, 161 points

View discussion on Hacker News →


The Contrarian Take: More Review Is the Wrong Instinct

Against the grain: The instinct to hire more reviewers or slow down the merge rate is backwards. Greptile's data shows agent-generated code already produces fewer rewrites than human code (Codex at 5-6% rework vs human baseline of 10%). The problem is not quality — it is the mismatch between what we review for and what actually breaks.

Here is the uncomfortable truth the data supports: agent code fails differently, not more frequently. The MSR study found that only 35.7% of rejected agent PRs reflected genuine agentic failures. 31.2% were rejected for workflow constraints (duplicate PRs, wrong branch, format issues) and 33.1% lacked observable decision rationale. More than half of “agent failures” are actually process failures.

As Simon Willison articulated, the key skill is not review but verification: “being able to confidently instruct agents on how to make changes and then confidently verify that those changes have been applied in the correct way.” We explored this distinction in Verify, Don’t Review.

The shift is from reviewing code (reading diffs) to verifying behavior (running assertions). Review asks “does this code look correct?” Verification asks “does this code do the correct thing?” The first question depends on human judgment that scales linearly. The second can be automated.

Hacker News discussion — HyperProbe: Agents that do read-only debugging in production, 69 points

View discussion on Hacker News →


The Survival Playbook

For teams shipping agent-authored PRs today, here is the minimum viable defense:

  1. Gate, don’t review. Move your quality enforcement from human review to automated gates. Contract tests, architecture fitness functions, property-based tests, and mutation testing catch the four horsemen’s failure modes at CI time.

  2. Separate the generators from the judges. The AI that writes the code must not be the AI that reviews it. CodeRabbit, Greptile, and Augment exist because this separation is fundamental, not optional.

  3. Tier your review investment. Agent-generated PRs that only touch files within one module and pass all automated gates need a 2-minute glance. PRs that cross service boundaries or change API contracts need a full human review. Allocate accordingly.

  4. Track architectural drift explicitly. Run coupling metrics, dependency analysis, and module boundary checks weekly, not just per-PR. The horseman you do not see is the one that kills you.

  5. Invest in blast-radius limits. Sandbox the agent’s scope per task. An agent working on a login flow should not be able to touch the billing module. We covered the infrastructure for this in It Fails on the Harness, Not the Model.

The four horsemen are not going away. Agent PRs will only increase in volume. But they are survivable, and the teams that build the right verification infrastructure now will have a structural advantage over those who keep trying to read every diff by hand.


Originally published at AgentConn

The AgentConn Weekly

Weekly digest of new AI agent releases, framework comparisons, and deployment guides. Built for builders.

Weekly. Unsubscribe anytime.

Explore AI Agents

Discover the best AI agents for your workflow in our directory.

Browse Directory