Every automation project we inherit has the same shape. Someone bought a tool, wired it to the CRM, and automated the most visible task in the department. Six months later the tool is still running, nobody trusts its output, and the team has quietly built a spreadsheet alongside it to do the real work.
The problem is almost never the model. It is the choice of what to hand over. After building automation into growth stacks across healthcare, aerospace, local services and SaaS, we've settled on a rule that has been right more often than any tooling decision we've made: automate the 60% of a workflow that is judgment-free, and leave the 40% that carries risk, taste, or relationship to a human.
This piece is the framework behind that number — how we find the 60%, how we protect the 40%, and the test we run before writing a line of code.
Why most automation lands on the wrong layer
Teams automate what is annoying. That instinct is understandable and almost always wrong, because annoying tasks are usually annoying precisely because they require judgment. Writing the follow-up email is annoying. Deciding who deserves a follow-up email is boring. The boring one is the automation candidate.
There's a second failure mode that costs more. Teams automate the output layer — the copy, the reply, the recommendation — because that's where the demo looks impressive. But output is where mistakes are visible to customers, and where a 5% error rate destroys the trust that made the workflow valuable. Meanwhile the input layer underneath it — enrichment, routing, deduplication, scoring, summarising, tagging, formatting — stays manual, invisible, and expensive.
The rule of thumb: automate toward the input, not toward the customer.
Move one layer down from wherever the excitement is. That's usually where the 60% lives.
Finding the 60%: the three-column audit
Before we scope any build, we map the workflow into three columns. It takes an afternoon and it has never failed to change the plan.
Column one — deterministic
Steps where the same input always produces the same correct output, and where being wrong is cheap and reversible. Data entry between systems. Deduplication. Lead routing by firm rules. Pulling firmographics. Formatting a report. Transcribing a call. Tagging a ticket by category.
This column is the 60%. It is unglamorous and it is where nearly all of the recoverable hours are.
Column two — probabilistic with a cheap check
Steps where a model is usually right, and where a human can verify in seconds rather than redo in minutes. Drafting a first-pass reply. Summarising a long thread. Suggesting a next action. Scoring intent.
Automate these, but never ship them straight through. The economics only work when review is genuinely faster than doing it manually — if your reviewer has to re-read the source to trust the summary, you have added work, not removed it.
Column three — consequential or relational
Steps where being wrong costs money, credibility, or a relationship. Pricing decisions. Anything a regulator reads. Clinical or legal content. The first human contact with a high-value account. Strategy. Negotiation.
This is the 40%. Protect it deliberately. The teams that pull ahead are not the ones who automate the most — they are the ones who buy back hours in column one and spend those hours in column three.
The test
For any step you're about to automate, ask: if this is wrong 5% of the time and nobody notices for a week, what does it cost? If the answer is "a little rework", automate it. If the answer is "a client", don't — instrument it instead.
The instrumentation nobody budgets for
Automation without measurement is just faster guessing. Every workflow we ship carries three things from day one, and they are non-negotiable in scoping:
- A ground-truth sample. A fixed set of real cases with known-correct answers, re-run on every change. Without it you cannot tell a model upgrade from a regression.
- An exception queue. Anything the system is unsure about goes to a human, and the reasons get counted. The pattern in that queue is the roadmap for the next iteration.
- A reversal path. Every automated action can be undone, and someone knows how. If it can't be undone, it belongs in column three.
This is the part clients push back on, because it feels like overhead on top of the thing they actually wanted. It is the difference between an automation that survives its first bad week and one that gets switched off. We build the measurement layer alongside the automation for the same reason we build analytics architecture alongside a funnel: an intervention you can't read is an intervention you can't defend.
What this looks like in practice
A services business came to us wanting an AI agent to write client proposals — the visible, exciting, column-three task. We ran the audit. Proposal writing stayed human. What we automated instead was everything feeding it: enquiry parsing, qualification scoring against historic close rates, service-line routing, pulling the right past-work examples, and assembling a pre-filled brief.
The proposal still gets written by a person. It now takes 20 minutes instead of two hours, because the thinking starts with the research already done. Nobody had to trust a model with the client relationship, and the recovered hours went into more proposals — which is what actually moved revenue.
That pattern repeats. In our local SEO work, the automation sits in data collection and reporting across dozens of locations, not in the pages themselves. In healthtech, it sits in operations and triage, well clear of anything clinical.
The sequencing that works
If you're starting from scratch, the order matters more than the tooling:
- Instrument first. You cannot find the 60% in a workflow you can't see. Two weeks of honest measurement beats two months of guessing.
- Automate one deterministic step end to end. Not five, partially. One, completely, with the exception queue running.
- Read the exceptions for a month. They will tell you what to build next, and it usually isn't what was on the roadmap.
- Only then move up a layer. Add probabilistic steps once the deterministic ones are boring and trusted.
Most teams try to skip to step four. That's why most automation projects have a spreadsheet running quietly alongside them.
Want the 60% mapped in your own stack?
We run the three-column audit as a fixed-scope engagement — you get the map, the sequencing, and an honest view of what's worth automating before anyone commits to a build. It's how every AI Automation project we take on starts.
The short version
The model you choose matters far less than the layer you point it at. Move down from the customer, not toward them. Take the deterministic 60%, instrument it properly, and spend the hours you recover on the 40% that a machine has no business touching. That split is the whole strategy — everything else is implementation.
If you want this thinking applied to your funnel rather than read about it, start a conversation, or see how it plays out across our client work.





