Skip to content

All postsAgentic AI

The Decision Card: Designing Human Approval for AI in Finance

Forty approvals in three minutes is a signature, not a control. How we design the card a controller sees when an AI system stops for a person, which batches are allowed, and how we measure whether review is happening.

May 26, 20265 min readWritten by Alexandru Bene

Forty approvals took just over three minutes. We were sitting beside a controller while she cleared the morning's queue from a new invoice workflow, and afterwards we asked what she had checked on each item. She thought about it and answered honestly: the amount, and whether the supplier name looked familiar.

She was not careless. The screen showed a full invoice image, twenty extracted fields, the proposed coding and an approve button. Nothing told her why each item was in her queue, so she checked what she could check quickly. The audit log recorded forty approvals by a named controller. In practice the model had approved forty invoices, with a person's name attached.

Palantir's "Connecting Agents to Decisions" (April 2026) describes the pattern most platforms now share: an AI stages an action and a person commits it. The part that decides whether this works is the moment in between. What the person sees when asked to commit, and how long they actually look.

One card per reason for stopping

An approval screen that shows the whole record asks the reviewer to find the problem. Finding the problem was the system's job. So we replace the record view with a decision card built around one question.

The card opens with one line stating the decision required and why the case stopped: "Price on line 3 is above the contract price for this supplier." Below it, side by side, the field that triggered review, the value found and the value expected. Then the evidence for that point only: the contract clause, the purchase order line, the last three invoices from this supplier for this item. The full record is one click away, not the first thing on screen.

A case that stopped for two reasons shows two cards. A case that stopped because a document could not be read shows the region that could not be read, not the page.

Show what is different, not what is there

Reviewers are faster and more accurate when they compare than when they read. So each card sets the case against its nearest precedent: the last approved case from the same supplier, for the same kind of document. Unchanged fields collapse. Changed fields are highlighted. A new bank account, a new unit of measure, a price that moved: what a controller would notice on paper if she had time, placed where her eyes land first.

We leave the model's confidence score off the card. A reviewer who sees a high number checks less.

Batching is a property of the cases

Approving one by one is slow, and controllers ask for a bulk approve button within the first week. Bulk approval is where review quietly stops, so the screen does not offer it. The cases have to earn it.

Cases can be approved together only if they stopped for the same reason, match the same precedent on every field except the one being approved, and fall within an amount band the controller has set. Ten invoices from one supplier, each with the same price increase against the same contract, are one decision. The same supplier with ten different reasons is ten.

Every override carries a reason

When a reviewer changes the proposed action, the card asks why, from a short list agreed with finance: wrong coding, price agreed outside the contract, duplicate, supplier error, other with a note. A free-text box gets ignored. A short list gets used.

These reasons are the most valuable data the system produces after go-live. They show which rules are wrong, which suppliers are changing and which segments should lose their right to go straight through.

Review effort is measured

We measure the review itself: time on each card, whether the evidence panels were opened, how often the proposed action is changed, how often an approval is later reversed. None of it is used to rate people. It is used to judge the control.

A review point where every case is approved unchanged in a few seconds, week after week, is adding no judgement. Either the cases do not need a person, and the review point should give way to an automated check, or the card is not showing what the person needs. Both are design problems. Neither is solved by asking reviewers to be more careful.

Before go-live, in a shadow environment, reviewers also work through cases with known answers, including some that should be rejected. It shows, before real money is involved, whether the card makes the wrong answer visible.

Approval is recorded like any finance control

The approver is never the system that prepared the case, and for payments and supplier master data, never the person who requested the change. Each approval is stored with the reviewer, the time, the card as shown, the evidence opened and the reason for any change. When an auditor asks how the control operated, the answer is what the reviewer saw, not only that they clicked.

For the same reason we do not offer approval by email reply or chat button. It strips the evidence out of the decision and keeps only the signature.

The name in the log should mean something

A review point nobody reads is worse than none, because it produces a record saying someone checked. Design the card so the name in the audit log describes what actually happened.

Discuss Your Operation With Our Engineers.

Describe one workflow and the systems it relies on. A senior engineer responds within two business days with an initial assessment: what we would build, what we would not, and why.

A senior engineer reads every request and replies within two business days.

Or book directly: discovery call calendar