Take a well-specified ticket: add a rate limit to an internal API, 429 on breach, per-API-key, sliding window, configuration in the existing settings module. Eighteen months ago that was roughly two hours of an engineer's day — forty minutes reading the surrounding code, fifty minutes writing it, thirty minutes on tests and self-review. Hand the same ticket to an agent today and a complete branch with tests exists in under ten minutes.
Here is the part nobody puts on the slide. The ticket still takes about two hours to land. The time moved. It now sits in the queue before a human opens the diff, in the twenty minutes that human spends deciding whether the sliding window is actually correct at the boundary, and in the follow-up conversation about whether per-API-key was the right axis in the first place — a question the ticket asserted and never justified.
That is the real story of agents in the software development lifecycle. Not acceleration. Redistribution. And if you do not know where the work moved to, you will staff the wrong side of it.
Writing code was never the bottleneck
This is uncomfortable for a profession that identifies with typing, but the data has been consistent for decades: the time between a work item being ready and the change being in production is dominated by waiting, not by authoring. Waiting for clarification. Waiting for review. Waiting for a release window. Waiting for the flaky test suite to go green on the third retry.
In a typical team, the implementation step is a minority of cycle time. So making implementation ten times faster does not make delivery ten times faster. It makes implementation stop being the constraint and hands the crown to whatever was second in line.
Queueing theory is unkind here. If you increase the arrival rate at a workstation without increasing its service rate, the queue does not grow a little. It grows non-linearly as utilisation approaches one. A review function that was comfortably absorbing twelve pull requests a day at 70% utilisation does not calmly absorb twenty. It saturates, and then wait times explode.
Agents do not remove work from the lifecycle. They move work from a stage that scales with headcount to stages that scale with attention — and attention is the resource you have least of.
Stage by stage: where it actually helps
It is worth being specific rather than generically enthusiastic, because the effect differs sharply by stage.
Specification and backlog refinement: gets harder
An agent will satisfy your acceptance criteria with unsettling literalism. A human engineer reading "rate limit per API key" pauses and thinks: what about unauthenticated traffic? What about our own internal service accounts, which share a key and will now throttle each other? They raise it in standup. An agent implements exactly what you wrote and the problem surfaces in production three weeks later.
This is not a model failure. It is a specification failure that the model faithfully reproduced. Underspecified tickets have always carried debt; agents make you pay it sooner and in public.
Implementation: genuinely transformed
For bounded, well-described changes inside an established codebase, first-draft quality is high and the marginal cost is close to zero. The categories that work best are the ones where the pattern already exists in the repository: a new endpoint alongside nine similar endpoints, a new migration, a new adapter behind an interface that already has three implementations, a mechanical refactor across forty files.
Testing: high volume, variable value
Agents produce enormous quantities of tests quickly. Whether those tests would have caught anything is a separate question, and coverage percentage will not answer it for you — generated suites are extremely good at asserting the behaviour the implementation already has. We wrote about the trap and the ways out of it in our note on making AI-generated tests actually useful.
Review: the new constraint
This is where the pressure lands, and it is worse than a straight volume increase. Human-authored diffs carry an implicit signal: the author suffered for every line, so lines are scarce and roughly intentional. Agent-authored diffs lose that signal. Code is cheap to produce, so diffs get bigger, more defensive, more speculative. A 400-line diff takes a careful reviewer somewhere between forty and ninety minutes to review properly, and review quality is known to fall off a cliff past a few hundred lines of change.
Worse, agent output is fluent. It reads as if it were written by someone competent and confident, which suppresses exactly the scepticism reviewers should be applying. Reviewers start skimming. Skimmed review is theatre.
Integration and release: mostly unchanged, now more contended
Merge conflicts, migration ordering and environment drift do not care who wrote the code. What changes is that more branches are in flight at once, so conflicts are more frequent and the integration window is more contended.
On-call: quietly riskier
At three in the morning, the relevant question is whether anyone on the team understands the change that broke. Code that no human ever fully read is code with no owner, and ownership is what makes incidents short.
What a good work item looks like when an agent is the consumer
The single highest-leverage change most teams can make is rewriting how they specify work. A ticket written for a human is a conversation starter. A ticket written for an agent has to be a contract, because there will be no conversation.
Four things matter more than everything else:
- Machine-checkable acceptance criteria. Not "handle errors gracefully" but "returns 429 with a
Retry-Afterheader; existing 200-path latency unchanged; new behaviour covered by a test that fails against the current main branch." - Explicit non-goals. Agents pattern-match toward completeness and will refactor adjacent code you did not ask about. Say what is out of scope.
- File and module boundaries. Naming the surface the change may touch is the cheapest way to keep the diff reviewable.
- The invariant behind the requirement. Say why, in one line. It is the only defence against a locally correct change that violates a system-level rule.
In practice we express this as structured metadata alongside the prose, so the definition of done is executable rather than aspirational:
interface AgentWorkItem {
id: string;
intent: string; // one line: the invariant this protects
scope: {
allow: string[]; // glob paths the change may touch
deny: string[]; // hard boundaries, e.g. 'db/migrations/**'
};
acceptance: AcceptanceCheck[]; // must be executable, not prose
nonGoals: string[];
maxDiffLines: number; // reject and re-plan above this
}
type AcceptanceCheck =
| { kind: 'test'; command: string; mustFailBefore: boolean }
| { kind: 'invariant'; assertion: string; verifiedBy: string };
const item: AgentWorkItem = {
id: 'API-2841',
intent: 'No single API key can exhaust shared request capacity.',
scope: {
allow: ['src/middleware/**', 'src/config/limits.ts', 'test/middleware/**'],
deny: ['db/migrations/**', 'src/auth/**'],
},
acceptance: [
{ kind: 'test', command: 'pnpm test middleware/rate-limit', mustFailBefore: true },
{ kind: 'invariant', assertion: 'p99 latency on the 200 path unchanged', verifiedBy: 'bench/api.bench.ts' },
],
nonGoals: ['Quota billing', 'Per-endpoint overrides', 'Changing the auth middleware'],
maxDiffLines: 400,
};
The mustFailBefore flag does a lot of quiet work. It forces a test that actually discriminates between the old behaviour and the new one, which rules out the most common class of worthless generated test.
Redesigning the pipeline around the new constraint
If review is the constraint, optimise review. That is the whole strategy, and it mostly means reducing the volume of human attention each change requires rather than asking humans to read faster.
| Lever | What it does | Cost |
|---|---|---|
| Hard diff-size ceiling | Forces decomposition; keeps changes inside the range where review quality holds | More work items, more planning overhead |
| Automated first-pass review | Clears mechanical defects before a human looks | Useless if noisy; see the limits below |
| Machine-verifiable acceptance | Moves "does it work" from opinion to CI | Requires real investment in test infrastructure |
| Tiered review by blast radius | Concentrates senior attention on auth, money, migrations and data | Needs an honest risk taxonomy of your own codebase |
| Agent writes the change, human writes the test | Keeps a human in the loop on intent, not syntax | Slower; worth it on high-risk paths |
Automated first-pass review deserves a caveat rather than a cheer. It is genuinely good at defects whose evidence is inside the diff and genuinely blind to violations of invariants enforced elsewhere in the system — the boundary is sharp enough that we wrote a whole note on what automated review catches and what it misses.
Tiered review is the lever most teams underuse. Not every change carries the same risk. A copy change in a marketing surface and a change to token validation should not receive the same process. Once you accept that, you can route mechanically: a change touching src/auth/** requires two human approvals regardless of size, a change confined to a presentational component with passing visual tests may need none.
Where this leaves the team
The composition of engineering work shifts rather than shrinking. Less time producing first drafts. More time on system design, on interface and invariant definition, on building the verification infrastructure that makes autonomous work safe to accept, and on review judgement.
Notice that every one of those is a senior activity. The uncomfortable implication is that agents raise the floor on output while raising the bar on the judgement needed to use that output safely. Teams that are strong on specification, testing and architecture get compounding returns. Teams that were already shipping underspecified work into a thin review process get the same problems, faster.
This is the design principle behind our own agent platform: the SDE, QA and PR-Review agents work the same backlog because implementation, verification and review are one loop. An implementation agent with no verification agent behind it just moves the queue.
What this means in practice
Instrument before you scale. Measure where cycle time actually goes today — time in backlog, time in implementation, time waiting for review, time in review, time waiting to deploy. If review is already your longest stage, adding implementation throughput will make delivery slower, not faster, and you will have the metrics to prove it before you spend the budget.
Then fix specification quality, because it is the cheapest change with the largest effect and it improves human-authored work too. Then invest in verification: the amount of autonomy you can safely grant is capped by the quality of your automated checks, and nothing else. Start agents on the reversible, well-patterned, low-blast-radius end of your backlog and expand the envelope as your evidence improves — the same way you would extend trust to a capable new engineer.
The teams getting real value here are not the ones that adopted the best model. They are the ones that treated this as a delivery-pipeline redesign rather than a tooling upgrade. If you are working through where the constraint actually sits in your own pipeline, we are happy to compare notes — and if the problem is broader than code, our AI automation work starts from the same question: which stage is really the bottleneck?