All postsStrategy
Why AI Should Make Your People More Capable, Not Fewer
A steering committee asked how many FTEs an order desk automation would save. The better question was how much demand the desk turned away. Why the headcount case caps growth, why context and judgement are worth more as answers get cheap, and how that changes what we automate, measure and decline.
The fourth slide of the steering committee deck had one number on it: 3.5 FTE. The project was an order desk automation for a distributor, and the finance lead had worked out that reading, checking and keying orders took three and a half people's time. The first question from the table was the one we hear in most first meetings. How many FTEs does this save?
We asked a different question back. How many quotes went out late last quarter, and how many of those customers ordered somewhere else? Nobody in the room knew. Everyone knew the two people sales called when a customer wanted something off-catalogue. Those two people were the real limit on the business, and the slide was quietly planning around them.
The headcount case is the easiest number to write
A headcount saving is certain, immediate and fits on one slide. Salary times people is a number finance already trusts. It is also a one-off. You bank it once, and the operation is then sized for the demand it had on the day you cut.
That works for a business that plans to stay the size it is. Most businesses paying for automation do not. They have more orders coming, a new market to open, a large customer asking for faster turnaround. For them the question is how much more demand the same operation can take on, and how fast it can respond. A project scoped to remove two people answers with a smaller team and a system that has stopped learning.
Cheap answers make context the scarce input
Reading an invoice, classifying an email or drafting a reply now costs very little, and every competitor can buy the same models on the same terms. The part of the work that separates one distributor from another sits elsewhere.
At the order desk it sat with two people. One knew which substitutions each large customer would accept without being asked. The other could tell from the wording of an email that a buyer was comparing quotes and needed an answer that afternoon, not on Thursday. Neither piece of knowledge was written down, and both decided whether orders were won.
A model does not know which customer is about to leave, or that a run of small complaints means a supplier is in trouble. It also cannot tell you where your operation is stuck. In our experience the value of AI in an operation rarely comes from a better model. It comes from finding the step that limits how much demand you can serve and removing it, which takes people who understand the business, the customers and the work. Cut them to fund the project and you lose the people who would have found the next constraint.
We start from the step that caps demand
A headcount case starts from the biggest block of effort. We start from the step that caps demand served: the queue that makes customers wait, the check only one person can do, the report that arrives too late to act on.
To find it, we pull timestamps from the order or ticketing system and look for where work waits, not where it is worked. Then we ask the team one question: when something unusual comes in, whose name gets said? Where the waiting and the same two names meet, that is usually the constraint. It is rarely the biggest block of effort.
At the order desk, the constraint was not order entry. It was the off-catalogue and special-terms cases, which waited for one of two people and set the pace for everything behind them. So the system reads every order, checks it against the catalogue, the price agreement and stock, and puts the routine ones through. The hard cases reach the two experts with the groundwork done: the customer's history, the likely substitutions, the margin at stake, and they decide. Their day moves from keying easy orders to the cases only they can settle, and the desk's capacity comes from that shift, not from anyone leaving.
The trap in this design is sending everything unclear to the experts. Then the bottleneck has moved, not gone. We cap what reaches them to cases where their knowledge changes the answer, and we track the age of their queue as a measure in its own right.
The system is built to spread expert judgement
Exceptions route by context, not by who is free: a substitution question goes to the person who knows that customer's product line, a terms question to the person who negotiated the agreement. Each decision is stored as a small record: what the system proposed, what the expert did, a reason code and one free-text line. Every week we sort those records by reason. A reason that repeats becomes a rule, with the expert's name on it. A case where the expert overruled the system becomes an evaluation case, and the system has to pass it before any change goes live.
Instinct gets the same treatment. When an expert holds an order because it does not look right, we ask for one line on what they saw. Often it is a pattern nobody had named: a new delivery address on a large first order, a price request that undercuts every earlier one. Once named, it is checked on every order, not only the ones that expert happens to see.
Over time, a newer colleague gets the same evidence, the same history and the reasoning behind earlier decisions on every case. Brynjolfsson, Li and Raymond studied a generative AI assistant in a customer support team ("Generative AI at Work", NBER, 2023) and found the largest gains among less experienced agents, because the tool carried the practices of the best agents to everyone else. We design for that effect, and it lasts only as long as the experts are there to learn from.
What we tell the CFO who asks for the headcount case
We do talk about cost, but we change what is measured. The value statement we agree before the build carries three or four of these:
- Demand served. Orders, quotes or cases cleared within the target time at the peak, with the same team.
- Response time. From customer request to useful answer, taken from system timestamps.
- Work turned away or delayed. Quotes past their promised time, orders lost with "no response" as the reason in the CRM, cases escalated because the team was full.
- Room for the next step. Whether a new market, customer or product line can be taken on without a new team to run it.
Then we say the rest plainly. If the business grows, the capacity goes into that growth. If it does not, the capacity shows up as hires you do not need as volume rises, and what to do with it is the business's decision. What we will not do is book the removal of the people the system depends on as the return. Once they are gone, the system stays as good as it was on the day it went live, and nobody is left who can make it better.
What we decline
We turn down work where the only stated goal is a smaller team, and we say so in the first conversation, before a Workflow Assessment.
We also decline a quieter version of the same thing: building a system on corrections from experts who already know they are being replaced. Their corrections dry up or turn defensive, the edge cases never get named, and the system learns only the easy cases. We would rather say that early than find it later in the evaluation set.
And we do not report FTE saved as the headline number, even when asked. Time saved can sit in the supporting detail. It is not why the work was done.
A saving is booked once
Two salaries saved is a number that stops moving the day it is booked. Two experts whose decisions now shape every case, and who spend their week on the customers and exceptions that need them, keep finding the next constraint. A system like this improves only while the people who understand the business are there to correct it, which is why we build it around them.