Two defects, same pull request. The first is a missing await on an audit-log write inside a request handler — the function returns, the response ships, and the promise resolves into nothing. An automated reviewer catches that essentially every time. The second is a change to an order-cancellation handler that sets status = 'cancelled' without also clearing the reserved-inventory row, breaking an invariant that is enforced by a reconciliation job in a different service. The automated reviewer approves it without comment.
Both defects are real. Both ship bugs. The difference between them is not difficulty, and it is not model capability. It is where the evidence lives. The first defect is fully visible inside the diff. The second requires knowing something the diff does not contain.
That distinction explains almost every result you will get from an automated reviewer, and it is the right frame for deciding what to trust it with.
Local correctness versus systemic correctness
A language model reviewing a pull request sees a few hundred lines of changed code, some surrounding context if your tooling is good, and whatever instructions you gave it. It is reasoning over that window. It is not reasoning over your system.
Within the window, it is genuinely strong. Pattern recognition over code is what these models are best at, and a surprising share of production defects are local pattern violations — the kind a very well-rested reviewer would spot on the first read and a tired one would miss on the third.
Outside the window, it has no basis for judgement and, critically, it usually does not know that. The failure mode is not "I cannot assess this." It is silence, or worse, confident approval.
What it catches reliably
In our experience the dependable categories share a property: a competent engineer could identify the defect with no knowledge of the system beyond the diff itself.
- Unhandled promise rejections and missing
await— including the subtle version where the value is awaited but the error path is not. - Null and undefined paths the type system was talked out of, particularly after an
asassertion or a non-null!. - Resource leaks — a file handle, connection, subscription or interval acquired on a path that can throw before release.
- N+1 queries visible in the hunk — an
awaitinside a loop over a collection, with a database call on the inside. - Missing branch coverage — a new conditional with no corresponding test, which is a mechanical observation over the changed files.
- Error handling that swallows — a
catchthat logs and continues where the caller needed to know. - Unsafe type assertions and validation gaps at boundaries where external data enters typed code.
- Forgotten cleanup — a feature flag referenced but never read, a debug log at info level, a commented-out block, a hardcoded value that belongs in config.
- Convention drift — naming, error-construction patterns, module layout that deviates from what the rest of the file does. Consistency is a pure pattern-matching task and models are excellent at it.
Two of these are worth more than the rest combined, for an unglamorous reason: missing await and swallowed errors are both silent in testing and loud in production, and both are exactly the kind of thing human reviewers stop seeing after twenty minutes of reading.
What it misses reliably
The misses are not random. They cluster around a single cause — the information required to make the judgement is not in the diff.
| Defect class | Why the reviewer cannot see it |
|---|---|
| Cross-module invariant violation | The rule is enforced somewhere else in the system, or nowhere explicit at all |
| Migration ordering and backfill safety | Requires knowing deploy sequencing and whether old code will run against the new schema |
| Backwards incompatibility for in-flight clients | Requires knowing who calls this and which versions are still live |
| Concurrency under real load | The race is between two executions; the diff shows one. Lock ordering, idempotency and retry interactions are invisible |
| Cost and latency regressions | A correct-looking call added to a hot path is only wrong if you know the path is hot |
| Authorisation at a trust boundary | Depends on where the boundary sits in your architecture, not on the shape of the code |
| Product intent | The code may be flawless and solve the wrong problem. Nothing in the diff says so |
| Deletions of load-bearing code | Absence of evidence reads as absence of risk |
Concurrency deserves its own warning. Automated reviewers do sometimes produce a comment about locking or race conditions, which creates an impression of competence in this area. Read those comments carefully. They are typically pattern-triggered — the word transaction appeared, or a shared mutable structure is visible — rather than the product of reasoning about interleaved executions. The correct expectation is that concurrency correctness is not covered, and to design your review process accordingly.
The false positive problem decides everything
This is the part teams underestimate, and it is the difference between a tool people rely on and a tool people mute.
Consider a reviewer bot that comments fourteen times on a 200-line pull request. Four comments are useful. Ten are stylistic noise, restatements of what the code plainly does, or speculative concerns that do not apply. The precision is 29%. The engineer reading it has to evaluate all fourteen to find the four, which costs more attention than reading the diff unaided would have.
Within a week, the team learns to scroll past the bot. At that point your recall is irrelevant. A finding nobody reads has the same value as a finding you never produced.
An automated reviewer is a classifier, and it should be evaluated like one. Precision is the metric that determines whether it survives contact with your team; recall only starts to matter once precision is high enough that people still read the output.
The practical consequences are unpopular but straightforward. Comment far less than you can. Require both high severity and high confidence before posting. Aim for two or three comments on a typical pull request, not fourteen. Default to advisory rather than blocking, and promote a rule to blocking only after you have evidence of its precision on your own codebase.
Making that tractable means the reviewer should emit structured findings, not prose, so you can threshold and measure them:
interface ReviewFinding {
file: string;
line: number;
severity: 'blocking' | 'high' | 'medium' | 'nit';
/** Model-reported, calibrated against your own labelled history. */
confidence: number;
category:
| 'correctness' | 'resource-leak' | 'error-handling'
| 'performance' | 'security' | 'test-gap' | 'convention';
rationale: string;
/** Required for correctness findings: the concrete input that breaks it. */
failingScenario?: string;
}
const POST_THRESHOLD: Record<ReviewFinding['severity'], number> = {
blocking: 0.9,
high: 0.8,
medium: 0.9,
nit: 1.1, // effectively never: let the linter own style
};
export function selectComments(findings: ReviewFinding[]): ReviewFinding[] {
return findings
.filter((f) => f.confidence >= POST_THRESHOLD[f.severity])
.filter((f) => f.category !== 'correctness' || Boolean(f.failingScenario))
.sort((a, b) => b.confidence - a.confidence)
.slice(0, 5);
}
Two details in there carry most of the weight. Setting the nit threshold above 1.0 disables style commentary entirely — a formatter and a linter do that job deterministically and for free, and every style comment spends credibility you need elsewhere. And requiring a failingScenario for correctness claims is a cheap, effective filter: a model that cannot name the input that triggers the bug is usually pattern-matching on the shape of risky code rather than finding a real defect.
Context is the lever, not the model
When teams are unhappy with automated review quality, the instinct is to change models. The larger gains are almost always in what you put in the window.
- Send whole functions, not hunks. A diff hunk with three lines of context above and below is not enough to judge whether an early return skips necessary cleanup. Expand to enclosing function bodies.
- Include the repository conventions file. Your error-handling pattern, your logging rules, your "never do X" list. This converts vague style opinions into checks against a stated standard, which raises precision sharply.
- Include the related test files. The reviewer cannot flag a missing test for a new branch if it never saw the test file.
- Include the pull request description and the linked work item. Without stated intent, the reviewer can only assess whether the code is internally consistent, never whether it does the right thing.
- Include definitions of the symbols the diff calls. Resolving the signature of the function being called is often the difference between catching an argument-order bug and missing it.
Adding the conventions file and the enclosing function bodies is usually a bigger quality jump than any model upgrade, and it costs a few thousand extra tokens per review. That is the trade you want.
The composition that actually works
Treat automated review as the first of several layers, each covering what the previous one structurally cannot.
- Deterministic tooling first. Types, linting, formatting, dependency audit. If a rule can be expressed deterministically, never spend model attention on it.
- Automated review second, for local correctness, error handling and test gaps — the categories where evidence lives in the diff.
- Human review third, explicitly scoped to what the machine cannot see: is this the right change, does it hold the system's invariants, is it safe to deploy in this order, who else depends on this.
Naming that third scope is the most valuable thing you can do to your review checklist. Reviewers who know the mechanical layer is already covered stop spending their first ten minutes on null checks and start spending it on architecture. That reallocation is where the value is — not in removing humans from review.
This is how the PR Review agent in our agent platform is built: it runs after the deterministic checks, posts a small number of high-confidence findings, and is deliberately quiet about anything it cannot ground in the diff. It works alongside SDE and QA agents on the same backlog, which matters, because an implementation agent without a reviewer behind it just increases the load on the humans downstream.
What this means in practice
Start in advisory mode and measure. For the first month, label every comment the bot posts as useful or not useful. You will have a precision number for your own codebase within a few hundred pull requests, and it will be lower than the vendor benchmark. Tune thresholds against that number rather than against a marketing claim.
Be honest with your team about the boundary. Tell them explicitly that the bot covers local correctness and does not cover invariants, migration safety, concurrency or intent. A reviewer who believes the machine has already checked everything is a worse reviewer than one who had no machine at all — the failure mode of automated review is not bad comments, it is unearned confidence.
And keep the human review checklist short and pointed at the gaps. Three questions a machine cannot answer beat twenty it already covered. If you are wiring this into your own pipeline, or thinking about the broader LLM and agent systems around it, tell us how your review process looks today — the process usually needs more work than the model does. For the human side of the same problem, our note on code review in the age of AI-generated code picks up where this one stops.