Skip to content

All postsStrategy

Intelligence Is Getting Cheap. Accountability Is Not.

The price of a model answer is falling faster than any enterprise can plan around. The cost of standing behind that answer is not falling at all. Where durable value in enterprise AI will sit, and what that means for the business case, the architecture and the vendor contract.

November 5, 20247 min readWritten by Anubhai Mehta

A CFO sent us two proposals for automating supplier invoice handling last month and asked which was cheaper. Both priced the model in detail: tokens per invoice, invoices per month, a cost per thousand documents. Neither said who would answer to the auditor when the system posted an invoice to the wrong cost centre, what evidence that person would have, or how the posting would be reversed.

The model line was the smallest number on either page, and it will be smaller next year. The line that was missing is the one that will decide the cost of the project, and it will not be smaller next year. That asymmetry is, in our view, the most useful thing an executive can understand about enterprise AI at the end of 2024.

The price curve is real

When GPT-4 launched in March 2023, it cost $30 per million input tokens. In July this year OpenAI released GPT-4o mini at 15 cents per million input tokens. It is a smaller model, but for the reading, sorting and extracting that make up most operational work it is good enough, at a two-hundredth of the price. Anthropic, Google and the open-weight models are on the same curve.

David Cahn's essay "AI's $600B Question" (Sequoia, June 2024) asked where the revenue will come from to justify the infrastructure being built. For a company buying AI, the same numbers say something simpler: model capability is becoming a commodity input, priced like bandwidth or storage. Nobody builds a strategy on having cheaper bandwidth than a competitor. The model will not be where an enterprise wins or loses.

What does not get cheaper

In February, the Civil Resolution Tribunal of British Columbia decided Moffatt v. Air Canada. The airline's website chatbot had told a customer he could claim a bereavement fare after travelling. The airline's policy said otherwise, and the airline argued that the chatbot was responsible for its own statements. The tribunal disagreed: the company is responsible for everything on its website, whoever or whatever wrote it. The damages were a few hundred dollars. The principle was not.

That principle is the second cost line. Every decision a system takes on a company's behalf has to be answered for by someone in the company, to a customer, an auditor, a regulator or a board. Answering needs four things, and none of them is on the token price curve:

  • A named person. Someone whose job includes this class of decision and who agreed to the rules the system follows.
  • Evidence. What the system saw, which rule applied, who approved, what was written where. Kept long enough to survive an audit.
  • Reversal. A way to undo the decision in the system that holds it, by someone who knows how.
  • Review. People with the time and the information to check the cases that need checking, and a way to know which cases those are.

These costs scale with the consequence of a decision, not with its length in tokens. A wrong answer about a bereavement fare costs a few hundred dollars. A wrong change to a supplier's bank details costs whatever the next payment run sends to a fraudster. Making the model ten times cheaper changes neither.

Cheaper answers mean more to answer for

There is a second effect, and it runs the wrong way for anyone budgeting on the token price. When an answer costs a fraction of a cent, it becomes worth automating decisions that were never worth automating before: the small invoice, the routine address change, the low-value credit note. Each one is another decision someone must answer for. As the price of intelligence falls, the number of decisions a company hands to software rises, and the accountability bill rises with it.

A company that plans on the token price will see its AI costs go up while the price per answer goes down, and will not understand why. A company that plans on accountability will see where the money is going: into the review, evidence and reversal around each class of decision. That is also where the money can be saved, because accountability can be engineered to be cheaper. A decision that can be undone in seconds needs less review than one that cannot. Evidence captured as a side effect of the work costs nothing to collect at audit time. Review aimed at the cases that need it costs a fraction of review spread across everything.

Where the value will sit

If intelligence is cheap and accountability is expensive, value in enterprise AI moves to the places that make accountability cheap. We see three.

The systems of record. The ERP, the bank, the policy administration system and the ticketing system are where a company's authority, audit trail and money already live. Their permissions, approval workflows, period controls and change logs are decades of accumulated accountability. AI that acts inside those systems, through their own interfaces and under their own controls, inherits all of it. AI that works beside them, in a separate assistant that people copy answers out of, has to rebuild all of it or go without. We expect much of this year's chat-first work to be rebuilt for that reason.

The company's own rules and records. Every enterprise has an unwritten rulebook: which differences a clerk lets through, which customers get an exception, which supplier always sends the wrong reference. A system can only be answered for if its rules are written down and owned. The companies that write their rules down, and keep a record of every decision and every correction, will have something a competitor cannot buy from the same model vendor. It is less glamorous than a model and far harder to copy.

Whoever carries the outcome. Foundation Capital argued in April that AI will let companies sell finished work rather than software, and Sequoia's October essay on the reasoning era makes a similar case for services delivered as software. We think they are right about the direction. What follows for a buyer is less discussed: a vendor who sells the work is selling accountability, and that is exactly the part a buyer cannot fully hand over. The tribunal did not ask which vendor built the chatbot. The price of outcome-based AI will be set by how much of the error cost the vendor will carry, and the contract that matters will be about who pays, who fixes and who keeps the evidence when the work is wrong.

Agents make the gap wider

Until this year, most enterprise AI suggested. A person read a summary or a draft and decided. That kept accountability cheap, because a person was answerable for every action and the system was just a faster way to read.

That is changing quickly. In September Salesforce announced Agentforce, built to act on customers' behalf inside its platform. In October Anthropic released a model that can operate a computer's screen and keyboard. When AI moves from suggesting to acting, the cost of a mistake moves from a few minutes of someone's reading time to money that has moved and a customer who has been told something. The intelligence gets cheaper at the same moment the accountability gets more expensive.

What to change now

None of this argues for waiting. It argues for spending in different places.

  1. Rewrite the business case in three lines. The cost of intelligence, which will fall. The cost of integration into the systems of record. The cost of accountability: review time, evidence, reversal and the owner's time, which will not fall and grows with consequence. If a proposal cannot fill in the third line, it is not a proposal yet.
  2. Name the person before the system. For every AI initiative, write down who answers for its decisions and what evidence they would need to do it. If nobody will put their name to it, the initiative is a demonstration, whatever the slides say.
  3. Build inside the transaction. Prefer designs where AI prepares or takes an action through the system of record, under its existing permissions and logs, over designs where it sits in a separate window. Measure progress in transactions the system touches, not in users who log in.
  4. Write the rules down and keep the record. Start with one workflow. List the decisions in it, the rule for each, and who owns the rule. Log every decision the system takes and every correction a person makes. This record becomes the company's evidence, its training data and its bargaining position with any vendor.
  5. Negotiate the error, not the price. When a vendor offers to sell outcomes, ask who pays for a wrong outcome, how fast it is reversed, and who owns the rules and the decision log if the contract ends. The answers are worth more than any discount per document.
  6. Spend on making decisions cheap to stand behind. Pick a model that is good enough and design so it can be replaced. Put the engineering effort into reversible actions, evidence captured as the work happens, and review aimed at the cases that carry consequence.

The line that decides it

Within two years, the model will be the cheapest line in most enterprise AI budgets, and nobody will remember which one was chosen. What will separate companies is whether they can say, for every decision their software took, who stands behind it and how they know it was right. That capability is built slowly, one workflow at a time, and no model release will shorten it.

Discuss Your Operation With Our Engineers.

Describe one workflow and the systems it relies on. A senior engineer responds within two business days with an initial assessment: what we would build, what we would not, and why.

A senior engineer reads every request and replies within two business days.

Or book directly: discovery call calendar