Skip to content

All postsAgentic AI

Autonomy Is Earned One Decision Class at a Time

Sponsors ask when the approvals can be switched off. Autonomy is not a setting on an agent. It belongs to each class of decision, moves up a five-rung ladder on evidence, and drops back down by rule.

August 12, 20255 min readWritten by Subash Natarajan

An order-handling agent makes about a dozen kinds of decision. It reads a delivery address. It maps a customer's product description to a catalogue item. It checks a price against an agreement. It accepts a substitution for an end-of-life product. Three months into running one, the sponsor asked the question every sponsor eventually asks: the agent is nearly always right, the approvals are slowing the team down, so when can we switch them off?

The question treats autonomy as a setting on the agent. Some of those decisions had been right on every reviewed case for weeks and could be undone in seconds. Others had been right just as often, but committed the company to something a customer would hold it to. One switch would have treated them the same.

We answered with a ladder. The sponsor now runs that conversation himself.

Autonomy belongs to a decision class

A decision class is a kind of decision defined narrowly enough that its cases share both their difficulty and their consequence. "Map a product description to a catalogue item" is one class. "Accept a substitution" is another. Classes are listed in the process map before the build, and every decision the agent takes is tagged with its class in the case record.

Autonomy is a property of each class, recorded with the date it was granted, the evidence behind it and the person who agreed it. The agent as a whole has no autonomy level. Asking whether the agent can run unsupervised is like asking whether the finance team can approve payments. Which payments?

The ladder has five rungs

  1. Suggest. The agent proposes; a person does the work. For new classes, and for decisions where the agent's job is to save reading time.
  2. Draft. The agent prepares the change in the target system as a draft; a person reviews and releases it.
  3. Act with approval. The agent executes after a person approves a decision card. The approval is a decision, not a rework, which is what makes it faster than a draft.
  4. Act and report. The agent executes; a named person reviews a daily report of its actions and can reverse any of them.
  5. Act silently. The agent executes and nobody looks, except through sampling and monitoring.

Anthropic's guidance on building effective agents (December 2024) argues for the simplest structure that works, with autonomy added only when it is needed. The ladder turns "only when needed" into a decision a business owner can take and defend.

The rung lives in the architecture, not the prompt

A rung that the agent reads from its instructions is a suggestion. A rung the architecture enforces is a control. So the rung of each class is a configuration record outside the agent, and it decides which tools the agent is offered for that class. On suggest, the agent gets read tools only. On draft, it gets a tool that writes drafts and nothing that releases them. On act with approval, the committing tool demands an approval token that only the decision card can issue, tied to that case and that version of the proposal. Only from act and report upwards is the committing tool callable directly.

Moving a class up the ladder is therefore a change to that record, made through the client's change process, and visible in the history. Nobody can promote a class by editing a prompt, and the agent cannot talk its way up a rung, because the tool it would need is not there.

Promotion needs four kinds of evidence

A class moves up one rung when four conditions hold and the process owner signs off in writing.

  • Accuracy on current cases. Measured on recently reviewed cases, not on the acceptance set. Necessary, and the least of the four.
  • Bounded consequence. The worst plausible error has a known cost and a known reversal. A class whose errors cannot be reversed, such as a payment that has left the bank, stops at act with approval.
  • A detection signal. Something after the fact would catch an error the reviewer no longer sees: a reconciliation, a customer confirmation, a daily report. Without one, promotion simply makes errors invisible.
  • An owner. A named person accepts responsibility for the class on its new rung.

A class can be right on every reviewed case and still be unsafe to promote, because the reviewed cases were the easy ones, or because the rare error would be expensive and undetectable. The evidence can be gathered automatically. The decision cannot, and the agent never proposes its own promotions.

Demotion is a rule, not a meeting

A class drops one rung automatically when a confirmed error appears after the reviewer stopped seeing it, when its inputs change shape (a new document layout, a new customer segment, a new market), or when its detection signal breaks.

A model upgrade, or a change to the rules, drops every class to act with approval until it re-qualifies. Sponsors like this rule least and value it most. It stops an improvement from quietly spending autonomy that was earned by a different system.

Evidence also ages when nothing changes at all. Every rung above draft carries a review date, and a class whose review lapses drops a rung. Autonomy that nobody has looked at for a year was granted to a business that no longer exists in quite that form.

Most classes settle in the middle

In the operations we work in, most classes end up on draft, act with approval, or act and report. Act silently is rare: we reserve it for classes that are reversible and continuously reconciled, such as matching a bank line to an open item, where any mistake surfaces in the next reconciliation run.

That is not timidity. An agent's value in an operation rarely comes from removing the last second of human attention. It comes from doing the reading, checking and preparation so the attention that remains is short and well spent. A class on act with approval, with a good decision card, costs a reviewer seconds per case. Moving it to act silently saves those seconds and removes the one place where someone would notice the world changing.

Two things we decline. A global autonomy switch, even for a pilot, because it ends with the most consequential decisions promoted alongside the easiest. And promotions during the close or a peak season, because an autonomy change is a process change and should go live when the people who would notice a problem have time to notice it.

The sponsor's question, answered

When can we switch off the approvals? For some decisions, now, and here is the evidence. For some, after another month of cases. For some, never, and here is why.

Discuss Your Operation With Our Engineers.

Describe one workflow and the systems it relies on. A senior engineer responds within two business days with an initial assessment: what we would build, what we would not, and why.

A senior engineer reads every request and replies within two business days.

Or book directly: discovery call calendar