All postsArchitecture
Model Retirement: Pinned Versions, Qualification Runs and an Exit Plan
Every hosted model a workflow depends on will be retired, and its replacement will be better on average and different in the particular. How we pin versions, qualify a replacement on its error profile, and write the exit into the handover.
The finance team forwarded the email with one line on top: can we just switch? Below it, their model provider announced that the model version their document workflow depended on would be retired on a fixed date a few months away. The recommended replacement was newer, faster and better on every published benchmark.
We ran the replacement against their acceptance set that afternoon. Overall accuracy went up. It also read some European dates in a different order, rounded amounts differently on a handful of supplier layouts, and was more willing to fill a missing purchase order number with a plausible guess. Better on average, different in the particular. In a finance workflow, the particular is where the money is.
Every hosted model has a retirement date, announced or not. A workflow built on a model is built on a component with a known lifespan, and the architecture has to treat replacement as a certainty.
Pin the exact version, never an alias
Most providers offer aliases that point at the latest version of a model family. They are convenient in development and dangerous in production, because the model behind the alias can change with no change on your side. The workflow that passed its acceptance run last month may be running on a different model today.
Every production workflow we build references an exact, dated model version. The version is part of the release, recorded in the release evidence, and changed only by a release. If a provider does not offer pinned versions with notice before retirement, a consequential workflow does not run on that provider.
A thin interface makes the model replaceable
The workflow never talks to a model directly. It talks to a small interface per task: extract these fields from this document, classify this ticket into one of these queues. Each interface has a fixed input and output schema, and every response is validated against it. Behind the interface sit the model, its prompt and its settings.
The rest of the system does not know which model is in use. Swapping one is a change behind one interface, tested against that interface's cases, and two models can run side by side on the same inputs, which is how a replacement is qualified before it carries real work.
Prompts belong to the model they were written against. Their phrasing, examples and formatting instructions reflect how that model responds, and moved unchanged they often work and sometimes fail subtly. So a prompt is versioned with the model it was qualified on, and a model change is always a prompt review.
Qualification compares error profiles, not scores
A candidate replacement runs on the full acceptance set, then on a recent sample of production cases with known outcomes. The headline accuracy is the least interesting output. What matters is the difference in errors: which cases the old model got right and the new one gets wrong, and which fields and segments they fall in.
A replacement that is better overall but newly wrong on dates from one market, or on credit notes, does not go live until those cases are handled, by a prompt change, a validation rule or a narrower scope. The finance owner sees the comparison and signs off the switch, as they signed off the original go-live.
Every extracted value points to its source
The most dangerous difference that afternoon was not the dates or the rounding. Validation catches those, because a wrong date or amount usually breaks a check against the purchase order. It was the invented purchase order number. A plausible made-up value passes format checks by design.
So every value a model extracts must point to where it appears in the source document: the page and the region of text. The validation layer confirms the value is actually there. A value with no location is treated as missing, not found. This check does not care which model is in use, which is exactly why it matters at a model change. It turns a new model's creative gap-filling into an ordinary, visible exception.
Replacements run in shadow first
After qualification, the candidate runs in shadow on live traffic for a period that includes a month-end. It receives the same inputs as the production model, its outputs are compared field by field, and nothing it produces reaches the ERP. Differences are grouped by field and segment, so the finance owner reviews a short list of kinds of difference, not thousands of cases.
Shadow running also measures cost per completed case at real volume. A new model can move it either way: a different price per token, a different token count for the same prompt, a different share of cases needing a second pass.
Switching cost is set by the acceptance set
Executives usually hear vendor lock-in as a question about contracts and APIs. For a model, both are the cheap part. The interface above makes the swap a small code change. The cost of a model change is the qualification run, the review of differences, the prompt adjustments and the sign-off, and all of that depends on one asset: a set of real, labelled cases the client trusts. With a well-kept acceptance set, a switch is days of work. Without one it is a small project, and in practice the client stays on whatever it has, which is what lock-in means.
So the most useful thing a company can own to stay independent of any model provider is not a multi-vendor abstraction layer. It is its acceptance set, kept current, with named owners.
The exit is part of the handover
We budget re-qualification into the running cost of every workflow we hand over, expecting it at least once a year, and the client's own team runs it. A client who cannot qualify a replacement without calling us does not fully own the system.
Every handover names the models the workflow has been qualified on, not only the one in use. Where hosting rules allow, one of them is an open-weight model the client can run on its own infrastructure, so there is always a route that does not depend on a single provider's roadmap. The handover also records the provider's notice period, so someone knows when to start the next qualification. And the qualification is rehearsed once during handover, on a real candidate, while the builders are still there.
Two things we decline: upgrading because public benchmarks improved, since they contain none of the client's suppliers, languages or layouts; and fine-tuning a hosted model unless the client accepts that the fine-tune retires with its base model.
A scheduled piece of maintenance
Model choice feels like the central decision in an AI project. Over the life of a workflow it is one of the most temporary. The acceptance set is the part that lasts, and with it the next retirement email is a date in the maintenance calendar, not an incident.