All postsAgentic AI
Tool Design as a Control for AI Agents in ERP Systems
The prompt says what an agent should do. The tools decide what it can do. We design enterprise agent tools as business verbs with contracts, not API wrappers, and publish them as a manifest an IT lead and an auditor can read.
The agent had extended a customer's price agreement by itself. We found it in a test run during a pre-production review of an order-handling agent. An order had failed because the agreement expired the previous day, and the agent resolved the failure the most direct way available. The action was logged and technically authorised. It was also outside anything the business had approved.
The agent's instructions limited it to creating sales orders and routing exceptions. Its single tool could call any endpoint of the ERP's API, under a service account with broad sales rights. The client's IT security lead then asked the question we now ask of every agent: which operations in the ERP can this system perform? The instructions described one scope. The tool granted another. In production, the tool's scope is the one that applies.
That review settled how we design agents for operations. The tool list is the specification and the permission boundary. The instructions are guidance inside it.
Tools are business verbs, not API wrappers
The quickest way to connect an agent to an ERP is to expose the ERP's API as tools. It is also the widest exposure available. An ERP's API is built for integrations written by engineers who read the documentation and test their code, and it offers every operation the system supports.
An agent needs the opposite: a few operations that match the steps of the business process, each doing one thing. So our tools are named and shaped as business verbs. Look up the price agreement for this customer and item. Check stock for these lines. Create a draft sales order. Attach a document to a case. Hand this case to a person with this reason. Each verb maps to a step in the process map the operations team agreed, and no tool exists that is not on that map.
Anthropic's engineering post on writing tools for agents (September 2025) makes a related point: wrapping existing APIs one to one rarely produces good tools, and tools should fit the tasks the agent actually performs. In an enterprise we go further. The tool set is a permission boundary, and it deserves the care of a role in the ERP's authorisation model.
The strongest control is a missing parameter
A tool that writes to a business system checks its inputs before anything reaches that system, and not only their types. The customer exists and is not blocked. The item is sellable in this sales organisation. The quantity is positive and plausible for this customer's history. The delivery date is not in the past.
The price is not a parameter at all. Draft orders take prices from the agreement, and the model has no way to supply one. Removing a parameter is a stronger control than validating it. If a model should never decide a value, the tool should not accept it.
Writing tools also take a key built from the business identity of the record, customer order number and line, so calling one twice returns the first result instead of a second record. The agent never has to reason about retries.
Errors are written for the model to act on
When a tool refuses, its message is part of the design. "Validation failed, code 4012" gives a model nothing to work with, and a model with nothing to work with tends to try something creative. That is how the price agreement got extended.
Our tools return errors that say what happened and what the agent may do next: "No valid price agreement for customer 4471 and item X on the requested date. Do not create the order. Hand the case to sales operations with reason: missing price agreement." The error is a small piece of the process, written by someone who knows the process. In our experience, well-written refusals prevent more wrong actions than any instruction in the prompt.
Each step sees only its own tools
Agents choose better from short lists. As the tool count grows, so does the chance of picking a plausible wrong one, and so does the surface an unexpected input can reach. So tools are scoped per step. The step that reads a purchase order sees document tools. The step that checks the customer sees read tools for customer, credit and agreements. Only the drafting step sees the tool that creates a draft order, and no step sees a tool that releases one.
Each tool runs under its own service credential with only the ERP rights it needs, so the scope holds even if the agent code is wrong. The boundary is enforced twice: by what the agent is offered and by what the ERP will accept.
The tool manifest is a control document
For every agent we deliver, the tools are published as a manifest: each tool's name, what it does in business terms, which system it touches, whether it reads or writes, which credential it uses, what it refuses and why. A few pages, written for people who do not read code.
The IT lead reviews it before go-live, and an auditor reads it to understand the scope of automation. If a proposed tool cannot be described in one plain sentence on that list, it is too broad. Changes to the manifest go through the client's change process like any change to a user role, because that is what they are.
Three requests we turn down
A general query tool, even a read-only one. Read access to everything, combined freely, is how an agent ends up quoting one customer's pricing to another. We add narrow read tools for the questions that actually come up, one at a time.
A tool that both decides and acts, such as "approve and post". Deciding and acting are separate steps with separate owners, and the tool boundary is where that separation becomes real.
"The prompt tells it not to" as a control. Prompts shape behaviour on the cases you thought of. Tools constrain it on the cases you did not.
The worst thing it could do
The first question an IT lead or auditor asks about an agent is the worst thing it could do. With a good manifest the answer takes a minute to read. If the honest answer is "anything the API allows", the agent is not ready, however good its instructions are.