An engineer opens a pull request at 4pm. It is 900 lines, it touches eleven files, the tests pass, and the description says "implement subscription pausing as specified in TICKET-4412". The reviewer knows, without being told, that the author did not write most of it. And the reviewer has a decision to make that their team has never discussed: what is their job here?
Under the old implicit contract, review was a conversation between two people who had both thought about the problem. One had thought hard enough to write it; the other checked their reasoning. That contract is broken now, and most teams have not replaced it with anything. They have the same checklist, the same approval button, and quietly more code arriving than before.
The process needs rewriting, not the tooling. Here is what actually changes.
The author's signal is gone
Human-authored code carries information beyond its content. Every line cost the author effort, so lines were scarce and roughly deliberate. If a function was 200 lines, someone had decided it needed to be. If an abstraction existed, someone had wanted it enough to build it. Reviewers read that signal without noticing they were reading it.
Generated code carries none of it. Lines are free, so there are more of them. Abstractions appear because they are conventional rather than because anyone needed them. Error handling is thorough in places that cannot fail and absent where it matters. Nothing in the diff distinguishes a considered decision from a default.
Worse, the code is fluent. It reads as though written by someone competent and confident, with consistent naming and tidy structure. Fluency suppresses scrutiny — it is much easier to be sceptical of awkward code than of code that looks like it knows what it is doing. Reviewers are being asked to apply more scepticism to material that invites less.
The reviewer's question has changed from "did you think about this correctly?" to "did anyone think about this at all?" Those require different reading, and most review checklists still assume the first.
Authorship is not optional
The most important norm to establish is also the simplest: the person who opens the pull request owns the code, regardless of what produced it.
That means being able to explain any line in it, having read all of it, and accepting the consequences when it breaks. "The agent wrote that part" is not an acceptable answer in a review thread or in an incident review. If you would not be comfortable defending a line you typed yourself, do not submit it.
This sounds obvious written down. It is not the default behaviour, and it erodes silently — an engineer under deadline pressure skims a generated diff, sees nothing alarming, and submits. Making the expectation explicit and repeating it is genuinely most of the work here. Teams that state it clearly behave differently from teams that assume it.
The practical corollary is that submission implies self-review first. Read your own diff before anyone else does. This one habit catches a large share of the leftover debris — the speculative abstraction, the unused parameter, the defensive check for a condition that cannot occur — and it is exactly what reviewers most resent spending their time on.
Diff size is now a process control, not a preference
Review quality falls off sharply with diff size. This was known long before agents; it is why "keep pull requests small" appears in every engineering handbook and why it was routinely ignored. What has changed is that the cost of producing a large diff has collapsed, so the natural size of a pull request has drifted upward with nothing pushing back.
A reviewer can genuinely assess maybe 200–400 lines with full attention. Beyond that, behaviour changes in a predictable way: they read the first files carefully, skim the middle, and check that the tests pass. That is not review. It is a ceremony with an approval attached.
So a size ceiling stops being a style guideline and becomes a control you enforce. Not as a lint rule that people learn to bypass, but as a norm with a stated reason: above this size we cannot review it properly, so we will not pretend to. The pushback is that decomposition costs the author time. It does. That is the trade — author time is now the cheap resource and reviewer attention is the scarce one, so spending the former to protect the latter is correct.
Two adjustments help. Separate mechanical changes from behavioural ones into different pull requests, because a rename across sixty files reviewed alongside a logic change means the logic change gets lost. And have generated work declare its intended surface up front, so a diff that wandered outside it is visible immediately — something we build into the work-item shape described in our note on how agents change the development lifecycle.
Review what the machine structurally cannot
Automated review handles the mechanical layer well — missing await, unhandled rejections, resource leaks, convention drift. There is no reason for a human to spend their first ten minutes there. But the boundary is sharp, and human review should be deliberately aimed at the far side of it, as we set out in more detail in our note on what automated pull request review catches and misses.
The questions worth a human's attention, in rough order of value:
- Is this the right change? The code may be correct and solve a problem nobody had. Generated code is never wrong about syntax and frequently wrong about intent, because it optimises for satisfying the stated requirement rather than for the requirement being right.
- Does it hold the invariants the system depends on? The rules enforced three modules away, in a reconciliation job, or only in someone's head. This is where the genuinely expensive bugs live, and evidence for them is never inside the diff.
- Is it safe to deploy, in this order? Migration sequencing, backwards compatibility with clients already running, whether old code will encounter the new schema during rollout.
- Is the complexity necessary? The most reliable smell in generated code is unearned abstraction — an interface with one implementation, a configuration option nobody asked for, three layers where one would do. Every one of those is permanent maintenance cost incurred for nothing.
- Would the tests have caught the bug? Not "are there tests". Generated suites reliably assert what the implementation already does. Pick the most important new branch and ask whether any test would fail if it were wrong.
- Can the team maintain this? If it uses a pattern nobody else on the team uses, it is a liability at 3am regardless of its elegance.
Question four is the one reviewers most often let through, because rejecting working code for being more complicated than necessary feels pedantic. It is the correct call. Complexity that nobody chose deliberately is the easiest kind to remove and the most expensive kind to keep.
Tier by blast radius, not by size
The volume problem does not get solved by asking reviewers to read faster. It gets solved by spending attention unevenly and on purpose.
Not all changes carry equal risk, and treating them identically means either over-reviewing trivia or under-reviewing the dangerous parts. Usually both. A copy change in a presentational component and a change to token validation should not pass through the same process.
| Tier | Examples | Process |
|---|---|---|
| Critical | Authentication, authorisation, payments, data deletion, migrations | Two human approvals, one senior; size ceiling enforced hard; no exceptions for urgency |
| Standard | Business logic, APIs, data access | One human approval focused on intent and invariants; automated pass first |
| Low risk | Presentational components, copy, configuration with tests | Automated checks plus lightweight human sign-off |
| Mechanical | Formatting, renames, dependency bumps with green CI | Automated verification; separated from behavioural changes |
Building this taxonomy for your own codebase is a half-day exercise and it is the highest-leverage process change available. It is also the one that makes the volume increase survivable — you are not reviewing less, you are reviewing the right things more.
Route it mechanically where you can. A change touching your authentication paths should require the critical process automatically rather than depending on someone noticing. Encoding the taxonomy as configuration rather than as a wiki page is what makes it hold:
# review-policy.yml — tier is derived from what the diff touches,
# never from who opened it or how urgent they say it is.
tiers:
critical:
paths:
- 'src/auth/**'
- 'src/billing/**'
- 'db/migrations/**'
- 'src/**/permissions.ts'
approvals: 2
require_codeowner: true
max_diff_lines: 400 # hard stop, no urgency override
checklist: [intent, invariants, deploy-order, rollback]
standard:
paths: ['src/**']
approvals: 1
max_diff_lines: 600
checklist: [intent, invariants, test-discriminates]
low_risk:
paths: ['src/components/**', 'content/**']
approvals: 1
checklist: [intent]
mechanical:
# Must not be mixed with behavioural change in the same pull request.
labels: ['formatting', 'rename', 'dependency-bump']
approvals: 0
requires: [ci_green, no_behavioural_diff]
# Applies to every tier: authorship is not delegable.
assertions:
- author_confirms_self_reviewed
- generated_code_declared_scope_respected
The max_diff_lines entry on the critical tier with no override is the line that matters, and it is the one teams are tempted to soften. The whole point is that urgency is precisely when review discipline is most valuable and least likely to be applied voluntarily.
This is also how the PR Review agent in our agent platform is configured: it clears the mechanical layer and flags which tier a change falls into, so human attention arrives already pointed at the right questions.
The team-level risks nobody puts on the roadmap
Two slower problems deserve naming, because neither shows up in a metric until it is well advanced.
Reviewer fatigue. Review is cognitively expensive and it has no visible output. When volume rises, review quality degrades before anyone reports a problem, because the approvals keep arriving on time. Watch for the signs: approvals within two minutes of opening, comment counts falling while diff sizes rise, the same one or two people reviewing everything. Review load needs to be a planned, distributed cost, not something absorbed between other work.
Learning loss. Struggling through an implementation is how engineers build models of a system. Accepting a generated one and reading it does not produce the same understanding, and the gap shows up eighteen months later in people who can ship features but cannot debug the system or reason about a design trade-off. This is a real cost and it lands on the most junior half of your team hardest. Some work should be done the slow way on purpose — and reviewing thoughtfully is itself one of the best remaining ways to learn a codebase, which is another reason not to let review become a rubber stamp.
What this means in practice
Write down the authorship norm and say it out loud in a team meeting: you own what you submit, you have read all of it, you can explain any line. Then make self-review before submission an expectation. Those two things cost nothing and change behaviour more than any tool you could install.
Build the risk taxonomy and tier your review process against it, so senior attention concentrates on authentication, money, migrations and data rather than being spread evenly across everything. Enforce a diff-size ceiling on the critical tier with a stated reason rather than as a style rule. Let automation own the mechanical layer entirely, and rewrite your human checklist to cover only what it cannot see — intent, invariants, deploy safety, unnecessary complexity, whether the tests discriminate, and whether the team can maintain it.
Review has become the most important stage in the pipeline rather than the last chore before merge, and it is worth staffing and scheduling accordingly. If you are working out how to restructure yours around a much higher volume of code, we are happy to compare approaches — it is the question we get asked most often once teams start shipping agent-written work.