All postsArchitecture
The Agent's Context Is Not the Record
An agent system holds three kinds of state: the system of record, the case record and the model's working context. Most production failures we are asked to review come from treating the third as one of the first two, and from plans that go stale while the agent works.
The ERP held the same supplier twice. A client's internal team had built an agent for supplier onboarding: it read the supplier's documents, checked the tax number, created the supplier, requested bank details and set up payment terms. In testing it was faultless. In its first week of real use, a routine deployment restarted the service while the agent was halfway through a case. When it came back, it created the supplier again.
The model had reasoned correctly both times. The fault was where the agent kept its knowledge of what it had already done: in a tool call and its response, inside a conversation history the restart discarded. The agent's memory of the world and the world had come apart, and nothing in the design noticed.
Most agent failures we review in production have this shape. They are state failures that look like reasoning failures, because the symptom shows up in the agent's output.
Three kinds of state, three owners
An agent inside an operation touches three kinds of state. Frameworks blur them, and the blur is where the trouble starts.
The system of record is the ERP, the CRM, the ticketing tool, the bank. It holds the business truth, applies its own validation and permissions, and belongs to the client's IT team. The agent changes it only through interfaces that already exist.
The case record is the durable account of the work: which case, which steps are done, what was decided at each, the evidence, who approved what. The workflow writes it, not the model. It lives beside the system of record, in a store the client owns, and never inside the ERP's custom tables or spare fields, where it would couple the agent to the ERP's upgrade cycle and mix working notes with business truth.
The working context is what the model sees for the current step: instructions, the relevant documents, facts pulled from the other two. It is built for one step and thrown away.
The rule follows. The context is a view, never a store. Anything that must survive a step is written to the case record or the system of record before the step counts as done. Kill the process at any instant and a fresh one rebuilds the context from those two sources and carries on. We test this literally, by killing the worker at random points during the acceptance run and checking that every case still finishes exactly once.
It is also a quick way to review any agent design, including a vendor's. Ask where each of the three kinds of state lives. If the answer for two of them is "in the conversation", the design has only one.
Intent is written before the side effect
Resumability comes from the order of writes. For every step that changes the outside world, the workflow records the intent (create supplier, with this business key), asks the system of record whether an earlier attempt under that key exists, makes the change only if it does not, then records the identifier the system returned. A crash between any two of those writes is harmless.
In the onboarding review, this sequence removed the duplicate without a single change to the agent's prompt. The team had spent a week tuning the prompt. The cost is a few extra writes and some latency per step, which in business workflows is always cheaper than one duplicate supplier, order or payment.
A plan is stale the moment it is written
Most agent frameworks treat the plan as the asset: reason once, then execute. In an operation that is backwards. A case that takes two days will see a buyer change the order, credit control block the customer and someone correct the delivery address, all while the agent holds a plan built on the old values.
So the plan is disposable and the preconditions are the asset. Each writing step carries the exact values it depended on when it was planned: the order version, the credit status, a hash of the address. Before executing, it compares them with the live record. Unchanged, it proceeds. Changed in a field the step does not use, it proceeds and logs the difference. Changed in a field it relies on, it stops and returns the case to planning, with both versions visible to a person.
In the systems we have reviewed, stale plans cause more wrong writes than wrong reasoning does. They are also invisible in testing, because test data never changes while the agent is thinking.
The same goes for master data. An agent that keeps its own copy for speed will act on a stale one sooner or later. It reads current values at the step that needs them.
Rebuilding context on every step has a real cost when a case carries long documents: a contract, a price list, a hundred-page tender. The answer is to cache what was derived from them, an extracted table or a summary of the terms, keyed by the version of the source it came from and never by the conversation. When the source changes, the key changes and the cache misses. Speed comes back without the context becoming a store.
Memory that changes decisions is a business rule
Long-running agents increasingly come with memory: summaries of past work, stored for later recall. Anthropic's engineering post on context engineering (September 2025) describes structured note-taking and compaction for long tasks, and for research and coding agents that is sound.
Business decisions need a harder line. If an agent has learned that a supplier always invoices per pallet while the purchase order counts cartons, that fact will change how future invoices are matched. Whatever the framework calls it, it is a business rule. So it is stored as an explicit record: the fact, the case it came from, who confirmed it, when it expires. A finance analyst can read it, correct it or delete it. Memory nobody can inspect, such as embeddings of old conversations that shift behaviour in ways no one can explain to an auditor, is not allowed to touch a business decision.
Where the business remembers
Separating state this way costs design time before the agent does anything impressive, and teams under pressure to demo skip it. The demo works anyway, because demos are short, single-user and never restarted. The return comes in production: the agent can be stopped at any point, two people looking at a case see the same history, and when the model is replaced next year no state changes, because none of it lived in the model.
An agent's context is where it thinks. The record is where the business remembers.