Point a capable model at a module sitting on 12% line coverage and you can be at 90% before lunch. The suite is green. The coverage badge is a pleasant colour. And the suite will not catch a single regression you care about, because almost every assertion in it was derived from the implementation rather than from what the code is supposed to do.
Here is the shape of it, taken from a discount calculator with an off-by-one in its tier boundary:
// The implementation, containing a bug: the tier boundary should be >=
export function discountFor(orderTotal: number): number {
if (orderTotal > 500) return 0.1;
return 0;
}
// The generated test. 100% line coverage. Passes. Useless.
it('returns 0.1 for orders over 500', () => {
expect(discountFor(501)).toBe(0.1);
});
it('returns 0 for orders of 500 or less', () => {
expect(discountFor(500)).toBe(0);
});
The second test does not merely fail to catch the bug. It codifies the bug. When someone later fixes the boundary to match the specification, this test goes red and a well-meaning engineer will "fix the test" back to the broken behaviour. You have taken a defect and given it institutional protection.
Coverage is a proxy, and generation games proxies perfectly
Line coverage measures whether a line executed during the test run. It says nothing about whether anything was asserted, whether the assertion was meaningful, or whether the expected value was derived from a specification or read off the current output.
That was always a weak proxy. It was tolerable when writing tests was expensive, because the cost of authoring imposed a rough discipline — a human writing an assertion by hand generally has some opinion about what the answer ought to be. Remove the cost and you remove the discipline. Generation optimises precisely for the measured thing, which is execution, not verification.
The question worth asking is not "what is our coverage". It is: would this suite have gone red before the last three incidents shipped? That question is answerable, and the answer for a freshly generated suite is usually no.
Four ways generated suites fail
Tautological tests
The category above. The model reads the implementation, computes what it returns, and asserts that. This is a circular argument rendered in code. It pins behaviour rather than verifying it, which has some narrow value in refactoring — but it is not testing, and calling it testing is how teams end up with false confidence.
Mock-shaped tests
Ask for a unit test of a service with four collaborators and you will often get a test that mocks all four. What remains under test is the orchestration glue. The assertions become "was repo.save called once with this object", which passes forever regardless of whether save actually persists anything or whether the object shape is right.
These tests are worse than no tests on two counts: they are coupled to implementation structure, so they break on every legitimate refactor, and they verify the test double rather than the system.
Snapshot sprawl
Snapshots are the path of least resistance for a generator, because no judgement about correct output is required — whatever comes out becomes the expectation. A few snapshots of stable, meaningful output are fine. Two hundred auto-generated snapshots produce a workflow where every change turns thirty snapshots red and the team resolves it by running the update flag. Now the suite asserts nothing at all, and it takes four minutes of CI to do so.
Tests generated from buggy code
The general case of the first failure. If the source of truth for generation is the implementation, then every existing defect becomes an expected behaviour. You are not writing a safety net. You are taking a photograph of the current state and framing it.
Mutation testing is the honest scoreboard
If coverage cannot tell you whether a suite has value, something else has to. Mutation testing does it directly: it introduces small changes to your source — flip a comparison operator, change a boundary, remove a statement, negate a condition — and reruns the suite. Each mutant that the suite fails to catch is a defect of that exact shape that could ship today.
Run it against our discount calculator. Mutate > to >= and the suite goes red, so that mutant is killed — but only because the test encoded the wrong boundary in the first place, which mutation score alone will not tell you. Mutate the return value of 0.1 to 0.11 and it also dies. Now mutate a branch that the generated tests executed without asserting on, and it survives. Survivors are where the real information is.
Stryker covers JavaScript and TypeScript; mutmut and cosmic-ray cover Python. The results are uncomfortable the first time. It is common for a suite with 85% line coverage to sit somewhere around 45–55% mutation score, and that gap is an exact measurement of how much of your suite is decoration.
Coverage tells you which lines ran. Mutation score tells you which bugs your suite would notice. Only one of those is a test result.
Mutation testing is expensive — you are running the suite once per mutant, so runtime scales with mutant count. Do not put it on every pull request. Run it nightly, or scoped to changed files, and treat the score as a gate on new code rather than a quest to fix the whole repository.
Where generation is genuinely excellent
None of this is an argument against generating tests. It is an argument about what you generate them from. Point the generator at the specification, the bug report or the input space rather than at the implementation, and it becomes one of the highest-value uses of a model in the entire delivery pipeline.
Regression tests from a failing trace
The single best use, with no close competitor. You have a bug report, a stack trace, maybe a request payload. Generating a minimal failing test from that artefact is fast, the correctness criterion is unambiguous (it must fail now and pass after the fix), and the resulting test is permanently valuable. Every incident should produce one, and nobody has ever enjoyed writing them by hand.
Boundary and edge-case enumeration
Humans are lazy about the unhappy path. Models are relentless about it. Empty collection, single element, duplicate keys, maximum integer, negative zero, unicode surrogate pairs, leap day, daylight-saving transition, timezone at the date line, string that looks like a number, deeply nested null. Ask for the edge cases and evaluate the list yourself — the enumeration is the value, the assertions still need your judgement.
Property-based tests
This is where generation and good testing genuinely align, because a property is a statement about intent rather than about implementation, so there is nothing to be tautological about. The framework then searches the input space far more thoroughly than any handwritten example set.
import fc from 'fast-check';
// A property is derived from the spec, not from the code.
// Any implementation satisfying the spec passes; the bug above does not.
it('discount never decreases as order total increases', () => {
fc.assert(
fc.property(
fc.double({ min: 0, max: 100_000, noNaN: true }),
fc.double({ min: 0, max: 100_000, noNaN: true }),
(a, b) => {
const [lo, hi] = a <= b ? [a, b] : [b, a];
return discountFor(lo) <= discountFor(hi);
},
),
);
});
it('applied discount never exceeds the order total', () => {
fc.assert(
fc.property(fc.double({ min: 0, max: 100_000, noNaN: true }), (total) => {
const off = total * discountFor(total);
return off >= 0 && off <= total;
}),
);
});
Monotonicity and bounded-output are properties a model can propose well when you describe what the function is for. The generator's job is proposing candidate invariants; yours is deciding which ones are actually true of your domain.
Parametrised table tests and fixture data
For pure functions with a specified input-output relation, table-driven tests are mechanical to expand and genuinely useful. Similarly, generating realistic fixture data — a few hundred plausible records with correct referential integrity and nasty-but-valid values — removes a real chore and improves test realism.
Generate from the spec, and close the loop
The operational change that matters is source of truth. If your work items carry machine-checkable acceptance criteria, those criteria are the correct generation input, and the resulting test is a genuine check on the implementation rather than a mirror of it. We described that work-item shape in our note on how agents change the development lifecycle; test generation is the stage where the discipline pays off most visibly.
Two rules make this concrete and are worth enforcing in CI:
- A generated test for new behaviour must fail against the unmodified branch. Run it before the implementation lands. If it passes, it is not testing the new behaviour, and it should be rejected automatically. This one check eliminates the tautological category outright.
- Generation should not see the implementation when the specification is available. Give it the acceptance criteria and the function signature. Withholding the body is the difference between verification and transcription.
The second operational change is closing the loop. A generator that emits files into a pull request has done the easy half. A QA agent that runs the suite against a real environment, observes the failures, distinguishes a genuine defect from its own bad assumption, and iterates is doing the job. That feedback loop is why the QA agent in our agent platform shares a backlog with the SDE and PR-Review agents rather than running as a standalone generator — tests written in isolation from execution are guesses.
The maintenance bill nobody costs
A test suite is code, and generated tests are code nobody has read. Three thousand generated tests that take eleven minutes of CI on every push impose a real tax: on build time, on the patience of engineers waiting for feedback, and on every future refactor, which now breaks four hundred tests that were coupled to structure rather than behaviour.
When those four hundred go red, the team faces a question they cannot answer: which of these failures indicate a real regression? Nobody knows, because nobody wrote them. The rational response is to delete or bulk-update them, at which point the entire exercise was negative value.
Volume is not the goal. A hundred tests that would each have caught a distinct real defect beat three thousand that assert the code does what the code does.
What this means in practice
Stop reporting coverage as a quality metric and start reporting mutation score on changed code. It is a harder number and a much more honest one, and it will immediately tell you whether your generated tests are worth their runtime. Set the gate on new code only — retrofitting the whole repository is a project nobody will finish.
Change what you generate from. Specifications and failing traces produce tests with real discriminating power; implementations produce mirrors. Add the "must fail before the fix" check to CI, because it is a handful of lines of pipeline configuration and it removes the single most common defect in generated suites. Review generated tests with the same seriousness as production code, and delete aggressively — a test that cannot fail is a liability with a maintenance cost.
Used well, this is genuinely one of the better applications of models in engineering: the unhappy paths finally get written, incidents reliably produce regression tests, and property-based testing stops being something teams mean to get around to. If you want help wiring that into a delivery pipeline rather than just generating files, our automation work is largely this shape of problem — tell us what your suite looks like today. For the review side of the same loop, see our note on what automated pull request review catches and misses.