Field report · · AgentConn Team
Four Horsemen of Agent PRs and How to Stop Them
Agent PRs fail four ways: context collapse, false confidence, review fatigue, and architectural drift. Data-backed survival strategies.
Four Horsemen of Agent PRs and How to Stop Them
Agent-generated pull requests grew from under 1% of GitHub PRs to 27.6% in just fourteen months. Anthropic’s 2026 Agentic Coding Trends Report puts the number higher: 41% of all new code is now AI-generated, with tools like Claude Code growing 6x in workplace adoption in under a year. The PR flood is here. The review infrastructure is not.
The problem is not that agents write bad code. It is that agent-authored PRs fail in ways human-authored PRs never did, and our entire review process was designed for a world where a human could explain their reasoning when asked. A forensic study of 33,000 agent-authored PRs found that agents achieve an 83.77% acceptance rate versus 91.01% for humans, but the gap hides the real story: the types of failures are fundamentally different. Agents do not make typos or forget semicolons. They produce syntactically correct, compilable code that violates contracts, breaks cross-service dependencies, and introduces architectural regressions that pass every test in the suite.
After months of tracking this data across academic research, industry reports, and practitioner experiences, we see four recurring failure modes that account for the overwhelming majority of agent PR disasters. We call them the Four Horsemen.
Horseman 1: Context Collapse
The agent knows the file. It does not know the system.
The first and most pervasive failure mode is context collapse: the agent generates code that is locally correct but globally wrong. It fixes the function but breaks the contract. It refactors the module but violates the architectural boundary. It adds the feature but duplicates logic that already exists three directories away.
FeatBit’s analysis of the 2026 productivity paradox identified four types of context loss in AI-generated pull requests: requirement context (what the business actually needs), codebase context (how the system fits together), review context (what previous reviewers flagged), and organizational context (team conventions, deployment constraints, implicit rules that live in people’s heads).
The MSR 2026 empirical study found that 23% of rejected agent PRs were duplicates, with agents frequently submitting PRs for issues already being addressed by another contributor. The agent had no idea someone else was working on the same problem. It had no concept of “someone else.”
View original post on Stack Overflow Blog →
The Stack Overflow engineering blog quantified the damage: AI-generated code produces 1.7x as many bugs as human code, with logic and correctness errors running 1.75x higher and security findings 1.57x more common. But here is the insight most teams miss: these are not bugs of incompetence. They are bugs of context. The agent wrote correct code for the wrong problem because it could not see the full system.
How teams are surviving it:
- Contract tests at PR gates. If the agent’s change breaks an API contract, the PR fails CI before a human ever sees it. Tools like Pact and Specmatic catch cross-service breaks that unit tests miss.
- Architecture fitness functions. ArchUnit-style assertions that codify your architectural boundaries. “No module in
domain/may import frominfrastructure/” catches the boundary violations agents love to introduce. - Codebase-aware agents. Greptile showed that agents with deep codebase context (dependency graphs, not just file content) produce significantly fewer rework cycles. Codex sits at 5-6% rework versus the human baseline of 10%.
Horseman 2: The Confidence Paradox
It looks right. It reviews right. It breaks in production.
View original report on New Relic →
This is the most insidious horseman, and the New Relic 2026 State of AI Coding report gave it a number: 94% of engineering leaders rate AI-generated code as higher quality than human code at the time of review. The code reads well. The variable names are descriptive. The comments are helpful. The structure is clean.
Then it hits production. 78% of those same respondents report more incidents once deployed. 82% experienced at least one production failure tied to AI-generated code in the past six months. 74% say at least a quarter of AI code needs significant rework within twelve months.
The paradox is structural: agents are optimized to produce code that looks correct to a reviewer scanning a diff. They have absorbed millions of examples of what “good code” looks like syntactically. What they have not absorbed is what “correct behavior” looks like under concurrent load, with stale caches, during a partial network partition, or when the third-party API returns a 429 instead of a 200.
As Addy Osmani wrote in what may be the most important essay on code review this year: “Code generation became cheap while understanding stayed expensive.” Agentic code review means reviewing code whose author cannot explain itself. Classic review validates a colleague’s reasoning. Agentic review must reconstruct reasoning that was never written down.
View original post on addyosmani.com →
How teams are surviving it:
- Property-based testing in CI. Instead of testing specific inputs and outputs (which agents are great at generating tests for), property-based tests define invariants the code must maintain. Hypothesis (Python) and fast-check (TypeScript) catch the edge cases agents miss.
- Canary deployments as review. If the code looks right in review and passes tests, deploy it to 1% of traffic first. The production environment is a better reviewer than any human for the failure modes agents introduce.
- Semantic diffing. Tools that show behavioral changes rather than textual changes. “This function now returns null for inputs > 1000” is more useful than “line 42 changed from
>=to>.”
Horseman 3: Review Fatigue
The queue grows faster than the team can read.
The math is brutal. Faros AI’s telemetry data shows that AI adoption correlates with 98% more PRs that are 154% larger, while review times have grown 91%. Zero-review merges are up 31%. The review queue is now a conveyor belt moving faster than anyone can watch.
This is not a people problem. It is a structural problem. As Tian Pan wrote: “The dominant failure mode of code review in 2026 is that reviewer instincts that worked on human-authored PRs break down on agent PRs because the bugs cluster in different places and the artifacts the reviewer sees are no longer the artifacts that matter.”
View discussion on Hacker News →
When a human writes a PR, reviewers develop a sense for where bugs hide. They check the boundary conditions, the error paths, the off-by-one opportunities. When an agent writes a PR, the code is syntactically flawless. The bugs are in the assumptions, not the implementation. Reviewers have to develop entirely new instincts, and most have not had time to.
Theo’s response to ThePrimeagen’s skeptical takes on AI coding agents captures the tension well. ThePrimeagen’s core complaints, which Theo largely concedes, include vibe-coded slop PRs as a real and growing review burden, juniors who let the agent do everything never building intuition, and auto-merge tooling that ships without a human gate. Theo pushes back on one point: experienced engineers using agents as power tools genuinely hit a higher productivity ceiling. The question is whether the review infrastructure can keep up.
The comparative study of agentic PRs found another pattern that amplifies review fatigue: agents ghost when they receive subjective feedback. A human contributor adjusts their approach when a reviewer says “this doesn’t fit the project’s pattern.” An agent either ignores the feedback or generates an entirely new PR from scratch. 61.38% of agent-authored PRs carry no recorded review activity at all.
How teams are surviving it:
- Tiered review. Not all PRs deserve the same scrutiny. Automated checks handle formatting, dependency updates, and simple refactors. Human review focuses on business logic, API contracts, and cross-service changes. We covered this in depth in 10x PRs, 1x Reviewers.
- AI-assisted review tools. CodeRabbit ($40M ARR, 700% YoY growth) and Greptile now handle the first pass across millions of PRs. The key insight: the AI that writes the code should not be the AI that reviews it.
- Reviewability as a metric. If a PR cannot be understood in under 10 minutes, it is too big or too poorly structured, regardless of who authored it.
View discussion on Hacker News →
Horseman 4: Architectural Drift
One PR is fine. A hundred PRs is a slow-motion rewrite.
The subtlest horseman does not show up in any single PR. It shows up over weeks and months as agent-authored changes silently shift the architecture. Each change is reasonable in isolation. Together, they constitute a drift that no one approved and no one noticed until the system became unmaintainable.
SoftwareSeni’s analysis documents how agents produce contract violations, cross-service dependency breaks, and architectural regressions that pass tests. The tests pass because each change is locally correct. The architecture degrades because no test asserts “the system as a whole still makes sense.”
This is the horseman that connects to the 48K files deletion incident we covered recently. Blast radius is the real infrastructure problem. When an agent can make hundreds of changes per day, the cumulative architectural impact dwarfs anything a human team would produce, because each individual change is too small to trigger alarm bells.
View original post on starkravingfinkle.org →
The Anthropic trends report notes that developers can “fully delegate” only 0-20% of tasks to agents, yet 41% of code is AI-generated. The gap is filled by supervision that is often cursory. As Andrej Karpathy warned, the coming “slopacolypse” is not code that fails immediately but code that is “almost right, but not quite,” degrading system quality gradually.
How teams are surviving it:
- Architecture fitness functions. Automated assertions that run in CI: dependency rules, layering constraints, module coupling limits. If the agent’s PR increases coupling beyond the threshold, it fails.
- Weekly architecture diff reviews. Not PR-by-PR review but periodic “zoom out” sessions where the team looks at how the system changed over the past week. Module dependency graphs and coupling metrics are your friends here.
- Agent guardrails. As we covered in The Anti-Slop Linter for Your Coding Agent, quality gates that run on every agent build are the cheapest form of architectural protection.
View discussion on Hacker News →
The Contrarian Take: More Review Is the Wrong Instinct
Here is the uncomfortable truth the data supports: agent code fails differently, not more frequently. The MSR study found that only 35.7% of rejected agent PRs reflected genuine agentic failures. 31.2% were rejected for workflow constraints (duplicate PRs, wrong branch, format issues) and 33.1% lacked observable decision rationale. More than half of “agent failures” are actually process failures.
As Simon Willison articulated, the key skill is not review but verification: “being able to confidently instruct agents on how to make changes and then confidently verify that those changes have been applied in the correct way.” We explored this distinction in Verify, Don’t Review.
The shift is from reviewing code (reading diffs) to verifying behavior (running assertions). Review asks “does this code look correct?” Verification asks “does this code do the correct thing?” The first question depends on human judgment that scales linearly. The second can be automated.
View discussion on Hacker News →
The Survival Playbook
For teams shipping agent-authored PRs today, here is the minimum viable defense:
-
Gate, don’t review. Move your quality enforcement from human review to automated gates. Contract tests, architecture fitness functions, property-based tests, and mutation testing catch the four horsemen’s failure modes at CI time.
-
Separate the generators from the judges. The AI that writes the code must not be the AI that reviews it. CodeRabbit, Greptile, and Augment exist because this separation is fundamental, not optional.
-
Tier your review investment. Agent-generated PRs that only touch files within one module and pass all automated gates need a 2-minute glance. PRs that cross service boundaries or change API contracts need a full human review. Allocate accordingly.
-
Track architectural drift explicitly. Run coupling metrics, dependency analysis, and module boundary checks weekly, not just per-PR. The horseman you do not see is the one that kills you.
-
Invest in blast-radius limits. Sandbox the agent’s scope per task. An agent working on a login flow should not be able to touch the billing module. We covered the infrastructure for this in It Fails on the Harness, Not the Model.
The four horsemen are not going away. Agent PRs will only increase in volume. But they are survivable, and the teams that build the right verification infrastructure now will have a structural advantage over those who keep trying to read every diff by hand.
Originally published at AgentConn







