Two candidates arrive in the same planning session. Route inbound support tickets to the right team. Approve supplier invoices under a threshold. On a slide they are the same shape: read an input, make a decision, take an action. Both are described as "AI automation" and both get the same estimate.
They are not remotely the same problem. A misrouted ticket costs a few minutes and is corrected by the person who receives it — the error is cheap, visible and reversible. A wrongly approved invoice moves money to an external party, and recovering it involves finance, the supplier and possibly a lawyer. One of these can run autonomously at 92% accuracy from week one. The other should not run autonomously at 99.5% accuracy, because the 0.5% is unrecoverable.
Most automation programmes do not fail during implementation. They fail at selection, months earlier, in a meeting where nobody asked what happens when the system is wrong.
Six dimensions that actually predict success
Here is the assessment we run before writing any code. It is deliberately boring and it kills a lot of proposals, which is the point — the cheapest automation project is the one you correctly decline.
1. Volume multiplied by handling time
This is the only source of value, so compute it first and compute it honestly.
Take a process running 200 times a day at 6 minutes of human handling time. That is 1,200 minutes, or roughly 20 hours of human time per day — three people. Automate 70% of it end to end and you have recovered around 14 hours a day. That is a real project with a real return.
Now take a process running 3 times a day at 6 minutes. That is 18 minutes a day. Automate it perfectly and you have saved an hour and a half a week, against an engineering build, an integration surface, a monitoring burden and a permanent maintenance obligation. The arithmetic says do nothing, and the arithmetic is right.
People systematically overestimate the volume of processes that annoy them and underestimate the volume of processes that are merely tedious. Get the numbers from a system, not from a conversation.
2. Error cost and reversibility
Volume tells you the upside. This tells you what you can actually ship.
Reversibility is the single best predictor of whether a process can run autonomously. Not accuracy — reversibility. A system that misroutes a ticket has created a correctable inconvenience. A system that sends the wrong email to a customer has created an impression you cannot retract. A system that approves a payment, deletes a record or files a regulatory submission has done something you may not be able to undo at any price.
The useful way to think about it: expected cost is error rate multiplied by cost per error, and if cost per error is effectively unbounded, no achievable error rate makes autonomy acceptable. That is not a modelling problem to be solved with a better prompt. It is a structural property of the process, and the correct response is a human approval gate, permanently.
3. Determinism of the decision
This is where we disappoint people, so we do it early.
If the decision rule is stable and expressible — route to the billing queue when the account has an open invoice and the message matches a known set of billing intents — then you want code. Not a model. Code is deterministic, testable, debuggable at 3am, free to run, and behaves identically on the millionth execution as on the first. A model will do the same job with a latency budget, a per-call cost, non-determinism and an evaluation harness you now have to maintain.
A large share of requests that arrive labelled "AI automation" are ordinary integration work. Two systems that do not talk to each other, a form that should write to a database, a report that someone assembles by hand from three exports every Monday. There is no inference problem anywhere in it. The honest answer is to build it as a piece of software, deliver it faster and cheaper than the AI version, and spend the model budget where variability actually demands it.
We say this to clients before we quote, and it costs us work occasionally. It is still the right call — an unnecessary model in a workflow is a permanent tax on reliability.
4. Input variability
So when is a model the right tool? When the input is unstructured and heterogeneous and no rule survives contact with it.
Free-text email threads where the request is buried in paragraph four. Invoices from 300 suppliers in 300 layouts. Support messages in multiple languages with typos and screenshots. Contract clauses that mean the same thing in nine different phrasings. This is the genuine use case: the variability is irreducible, a rules engine would need a thousand rules and would still miss, and a model handles the long tail gracefully.
The test is simple. If you can write the rules, write the rules. If every attempt to write the rules produces an ever-growing list of exceptions, you have found a real inference problem.
5. Availability of ground truth
Ask one question: after the system produces an output, can anyone determine whether it was correct, and how soon?
For ticket routing, yes — a reassignment is a labelled error, available within hours, essentially for free. For invoice coding, yes, at month-end close. For "assess whether this supplier poses a delivery risk", often no: there may be no event that confirms or refutes the judgement, and if there is, it arrives eighteen months later.
No ground truth means no evaluation, and no evaluation means no operation. You cannot detect drift, you cannot tell whether a prompt change helped, and you cannot justify raising the autonomy level. Processes without a feedback signal are the ones that quietly degrade for six months before somebody notices.
6. Process stability
A process that is restructured every quarter will consume your engineering capacity in maintenance. You are not automating a process; you are automating a snapshot of it. Ask when it last changed materially and who owns it. A process with no clear owner is a process that will change without telling you.
And the related trap: automating a broken process instead of fixing it. If the ticket queue needs routing because the intake form asks the wrong questions, the automation is an expensive workaround for a ten-minute form change. Map the process before you automate it and you will occasionally find the whole step should be deleted.
Scoring it
We make this explicit rather than intuitive, because an explicit score is arguable and an intuition is not. The gate matters more than the score — a candidate can look excellent on volume and still be disqualified on ground truth.
interface Candidate {
name: string;
runsPerDay: number;
minutesPerRun: number;
/** Unbounded when the action cannot be undone: payments, deletions, filings. */
reversibility: 'trivial' | 'costly' | 'irreversible';
/** 'rules' means build software, not a model. */
decisionType: 'rules' | 'judgement-over-structured' | 'judgement-over-unstructured';
groundTruth: 'immediate' | 'delayed' | 'none';
materialChangesPerYear: number;
}
const DAILY_HOURS_FLOOR = 4; // below this, the build rarely pays for itself
function assess(c: Candidate) {
const dailyHours = (c.runsPerDay * c.minutesPerRun) / 60;
const disqualifiers: string[] = [];
if (dailyHours < DAILY_HOURS_FLOOR) disqualifiers.push('insufficient volume');
if (c.groundTruth === 'none') disqualifiers.push('not evaluable — cannot be operated');
if (c.materialChangesPerYear > 3) disqualifiers.push('process too unstable');
const recommendation =
disqualifiers.length > 0 ? 'decline'
: c.decisionType === 'rules' ? 'build-as-software'
: c.reversibility === 'irreversible' ? 'draft-for-approval-only'
: 'automate-with-autonomy-ladder';
return { dailyHours, disqualifiers, recommendation };
}
Note that the highest-value processes are frequently not the ones with the most executive attention. Invoice line-item coding is nobody's strategic priority and is often the best candidate in the building: high volume, unstructured input, reversible before close, immediate ground truth, stable for years.
The autonomy ladder
Once a candidate passes, the question is how much authority to grant it. The answer is not a design decision made up front. It is a position you earn with measured accuracy.
| Rung | System behaviour | What it takes to get here |
|---|---|---|
| Shadow | Runs on live input, writes nothing, decisions logged and compared to the human's | Nothing. Always start here |
| Suggest | Proposes an option the human accepts or overrides in one click | Shadow accuracy that beats the current process |
| Draft for approval | Produces the complete action; a human approves before it takes effect | High acceptance rate on suggestions. Terminal rung for irreversible actions |
| Act with reversal window | Executes, notifies, holds a defined window in which it can be cleanly undone | A genuine undo path, plus measured accuracy at the target rate |
| Autonomous | Executes, escalates only low-confidence cases | Sustained accuracy, stable drift metrics, and reversible consequences |
Shadow mode is the rung teams skip and the one that does all the work. It costs almost nothing, it runs against real production traffic rather than a sanitised test set, and it produces the labelled dataset you need to argue for promotion. It also surfaces the input distribution you did not anticipate — the 8% of tickets that arrive as forwarded threads with four quoted replies — before those cases can cause damage.
Autonomy is not a setting you configure. It is a claim about measured accuracy and bounded consequences, and it should require evidence in the same way a production deployment does.
Within a rung, confidence thresholds do the routing. Below the threshold, escalate to a human with the reasoning attached. Above it, proceed. Two rules keep this honest: model-reported confidence must be calibrated against your own outcome data before you trust it as a threshold, and the escalation path needs a named owner and a service level. An escalation queue nobody watches is just a slower failure.
The unglamorous parts are what make it survive
The decision logic is a fraction of the work. What determines whether an automation is still running in a year is the operational scaffolding, and it is the first thing cut when a pilot is rushed.
- Idempotency. Every action needs a stable key so a retry cannot double-post an invoice or send a message twice. Assume every step will be retried, because eventually it will be.
- Decision audit log. For every execution: the input, the retrieved context, the decision, the confidence, the model and prompt version, and the action taken. When someone asks why the system did something in March, this is the only acceptable answer.
- Replay. The ability to rerun historical inputs through a new version and diff the decisions. This is how you ship a change without gambling.
- An off switch. One flag, no deploy required, that any operations person can flip. Every autonomous system needs one and it should be tested, not theoretical.
- Drift monitoring. Alert on the distribution, not just on errors: a sudden shift in confidence, in escalation rate, or in the mix of categories usually means the input changed and you are now operating outside your evaluation set.
When a workflow needs several specialised steps that hand work to each other — extract, validate, decide, act — the coordination becomes its own design problem, with its own failure modes around retries and partial completion. We work through those in our note on multi-agent orchestration patterns.
What this means in practice
Inventory before you build. List the candidate processes, get real volume and handling-time numbers from systems rather than from opinions, and score each one on error cost, reversibility, determinism, input variability, ground truth and stability. Expect most of the list to be disqualified, and expect a meaningful share of the survivors to be integration work with no inference problem in them at all. Both outcomes are wins — you have avoided spending a quarter on something that would not have paid back.
For whatever survives, start in shadow mode and climb the ladder on evidence. Build the audit log, the replay path and the off switch in the first version, not the second. And fix the process before you automate it, because automating a broken process just makes the brokenness faster and harder to see.
The pattern across the programmes that work is unremarkable: they picked fewer processes, picked them on reversibility and volume rather than on enthusiasm, and invested in the operational plumbing that makes an autonomous system safe to leave running. If you want a second opinion on your own shortlist — including which items we would tell you not to automate — send us the list. Sorting that out is the first thing we do in any AI automation engagement.