The system runs payroll for 4,000 people. It was written in 2009, the last engineer who understood the tax calculation module left in 2019, and there are 11,000 lines in a single file called process.php with no tests. Everyone agrees it must be modernised. Every attempt has been abandoned, because the first question — what does it currently do? — has no answer anyone will commit to.
This is where AI assistance is genuinely transformative, and it is not where most people expect. The headline promise is automated translation: point a model at the old code and get the new code. That is the least valuable thing on offer and the most dangerous. The real value is in comprehension and in building the safety net that makes any change survivable.
Because the thing that blocks legacy modernisation is almost never the typing. It is that nobody knows what the system does, so nobody can tell whether a change broke it.
Why rewrites fail, briefly
The instinct is a clean-slate rewrite. It fails for a reason that has nothing to do with engineering skill.
The old system encodes a decade of accumulated correctness. Not in its architecture — the architecture is usually bad — but in its details. The special case for employees who transferred mid-quarter. The rounding rule that matters for one jurisdiction. The retry with the specific back-off that stops a downstream system from falling over. None of it is documented. Much of it looks like a bug until you find out why it is there.
A rewrite reproduces the parts that are obvious and loses the parts that are subtle, and you discover the difference in production, one angry edge case at a time. Meanwhile the old system is still live, still accumulating changes, and you are maintaining two of everything.
Incremental strangulation — routing traffic progressively from the old system to the new, module by module, with both running — is the approach that works. It has always been correct and has always been slow, because each increment needs you to understand one region of the old system well enough to replace it safely. That understanding step is where models change the economics.
Comprehension is the real win
Handing a model 800 lines of undocumented procedural code and asking what it does produces, in about a minute, something a human would take a day to write: a description of the control flow, the data it touches, the branches and what appears to trigger them, the side effects, and a list of things that look anomalous.
It will be roughly 85% right. That number sounds disqualifying and is not, because the failure mode is favourable — you are not accepting the output as truth. You are using it as a map to direct your verification, and reading code with a hypothesis is enormously faster than reading it cold. The 15% that is wrong tends to be wrong in ways that a targeted check exposes quickly.
Three uses are worth building into the work:
- Behavioural summaries per module, reviewed and corrected by an engineer, committed to the repository. You are creating the documentation that should have existed, as a by-product of needing it.
- Dependency and coupling maps. Which modules touch which tables, which globals are written where, what the implicit contracts are between regions of the code. This is the input to deciding where the seams are, and it is tedious to assemble by hand.
- Business-rule extraction. Pulling the conditional logic out as a list of stated rules — "employees with status T and a start date after the 15th are prorated using calendar days, not working days" — and taking that list to the people who own the process. Frequently half the rules turn out to be obsolete, and deleting a rule is cheaper than porting it.
That last one changes scope more than anything else. Modernisation projects are usually sized on the assumption that everything must be reproduced. It usually must not.
Characterisation tests: the highest-value use
You cannot safely change code you cannot test, and legacy code is untestable for structural reasons — global state, no dependency injection, side effects everywhere.
The way through is characterisation tests. Not tests of correct behaviour, which nobody can define. Tests of current behaviour: capture what the system does now, for a wide range of inputs, and pin it. When you then refactor, any deviation is visible immediately. If a pinned behaviour turns out to be a bug, you fix it deliberately, as its own change, with the test updated to say so.
Ordinarily this is dull, enormous work, which is why teams skip it. Generating it is exactly the kind of mechanical breadth models are good at — and unlike generating tests for new code, the usual objection does not apply. A characterisation test is supposed to assert what the implementation currently does. Tautology is the goal.
# Capture real behaviour from production inputs, then pin it.
# The expected values are recorded, not reasoned about — that is the point.
import json, pytest
from legacy_bridge import call_legacy_payroll
with open("fixtures/payroll_cases.json") as fh:
CASES = json.load(fh) # inputs sampled from a year of production traffic
@pytest.mark.parametrize("case", CASES, ids=lambda c: c["id"])
def test_payroll_matches_recorded_behaviour(case):
result = call_legacy_payroll(case["input"])
assert result == case["recorded_output"], (
f"Behaviour changed for {case['id']}. If this change is intended, "
f"update the fixture in the same commit and say why."
)
Two details matter. Sample the inputs from real production traffic rather than inventing them, because the distribution is the value — the odd cases are the ones you need pinned, and you will not think of them. And where the legacy system is not callable in isolation, run the old and new implementations side by side against the same input and diff the outputs. That comparison harness is often the single most useful artefact of the whole project, and it is worth building before any code is replaced.
Where AI-assisted refactoring genuinely helps
With comprehension and a safety net in place, mechanical transformation becomes low-risk and fast. The categories that work well are the ones where the change is structural and the outcome is verifiable:
| Task | Suitability | Why |
|---|---|---|
| Mechanical syntax and API migration | Strong | Pattern-uniform, compiler-verifiable, tedious at scale |
| Extracting pure functions from procedural blocks | Strong | Clear criterion; characterisation tests confirm equivalence |
| Adding types to untyped code | Strong | Inference from usage is exactly the model's strength |
| Introducing seams for dependency injection | Good | Repetitive and well-understood, but touches call sites broadly |
| Splitting a large file by responsibility | Good | Needs human judgement on the boundaries |
| Translating one language to another wholesale | Poor | Produces idiomatically wrong code carrying the old design |
| Redesigning the data model | Poor | Requires domain and business context the code does not contain |
| Deciding the target architecture | Not applicable | This is the actual engineering work |
The wholesale-translation row is the one to internalise. A model asked to convert 11,000 lines of PHP to TypeScript will produce 11,000 lines of TypeScript that looks like PHP — the same global state, the same procedural shape, the same undocumented special cases, now in a language where none of it is idiomatic. You have changed the syntax and kept every structural problem, and you have thrown away the one advantage the old code had, which is that it is battle-tested. This is the modernisation equivalent of a rewrite with extra steps.
Transformation without comprehension is just a rewrite you did not admit to. The order is always: understand, pin the behaviour, then change — and the model helps most with the first two.
Sequencing the work
The order that holds up in practice, on a system nobody understands:
- Map before touching anything. Generate module-level summaries and a dependency map, have an engineer verify them, commit them. You now have documentation and a basis for planning.
- Build the comparison harness. The ability to run old and new against identical input and diff the results. Everything downstream depends on this.
- Pin current behaviour with characterisation tests generated from production-sampled inputs, prioritising the highest-risk modules.
- Extract the business rules and interrogate them. Take the list to the process owners. Delete what is obsolete before you port it.
- Choose seams and strangle incrementally. Start with a module that is low-risk, well-bounded and has a clear interface — not the most painful one. The first increment is where you prove the process works.
- Route traffic progressively with a flag, shadowing where possible: run both, serve the old result, compare the new one, and promote when the diff is clean.
- Delete the old path deliberately. Unretired legacy code is the most expensive thing in this list, because you are now maintaining both.
Step five gets argued about. There is always pressure to start with the module causing the most pain. Resist it once — the first increment is as much about validating your comparison harness, your flagging and your rollback as it is about the code. Prove the machinery on something forgiving.
Steps one through three are where agent assistance scales well, because they are broad, mechanical and independently verifiable. That is the kind of work we put through our agent platform: generating characterisation suites across many modules in parallel, with the diff against recorded behaviour as an unambiguous pass criterion. Generation of that breadth needs executable verification behind it, which is the same argument we make in our note on AI-generated tests.
The honest limits
Three things will not be solved by any amount of model assistance, and pretending otherwise is how these projects go wrong.
Missing domain knowledge stays missing. If the code contains a rounding rule and no human alive knows which regulation it implements, a model can tell you what the code does but not whether it should. Those decisions need someone with authority over the business process, and finding that person is a project task, not a technical one.
Data migration remains the hardest part, and it is where these projects actually fail. Twelve years of accumulated data with schema changes, partial backfills, encoding inconsistencies and rows that violate constraints added after they were written. No model reconciles that for you. Budget for it as its own workstream from the start.
Finally, models generate code with confidence that is unrelated to their understanding of your system, and legacy code is dense with non-obvious load-bearing details. A plausible refactor that drops one special case is the characteristic failure, and it is why the characterisation suite comes before the refactor rather than after.
What this means in practice
Spend the first phase on understanding, not on code. Generate the module map and the behavioural summaries, verify them, and commit them to the repository. The documentation is worth having on its own, and it is what turns an unsizable project into a plannable one.
Then build the safety net before the first refactor. Characterisation tests from production-sampled inputs plus an old-versus-new comparison harness are what make incremental replacement safe, and they are the part most teams defer and then regret. Use models for the mechanical breadth — comprehension, test generation, type inference, uniform migrations — and keep architecture, data modelling and sequencing with your engineers.
Done this way, modernisation stops being a multi-year gamble and becomes a series of small, reversible, verified steps. It is slower than the rewrite you were hoping for and it finishes, which the rewrite generally does not. If you are sitting on a system nobody wants to touch, tell us what it does and what you know about it — mapping the unknown is where our software engineering work usually starts.