Skip to content

All postsArchitecture

In Regulated Environments, the Model Is the Easy Part

Hosting, data minimisation, change control and evidence decide whether AI may run in a regulated operation. They change the design, not only the paperwork, and two of them pull in opposite directions.

September 30, 20255 min readWritten by Subash Natarajan

The servers would have no route to the internet, and that was not going to change. The head of security at a national digitalisation agency said it in the first architecture meeting, and it decided the rest of the project. The agency wanted AI ticket triage for its service desk: 100% on premises, air-gapped, with no dependency on a model vendor.

We built it that way, and it runs in production. The lesson applies well beyond air gaps. In a regulated operation, choosing and running the model is the easy part. The hard part is everything that decides whether the model is allowed to run: where data goes, what each step may see, how changes are approved and what evidence exists afterwards.

Hosting is a design input

Teams often design around a hosted model and ask about hosting at the end. In a regulated operation, the answer to "where may this data be processed" removes options before design starts. It decides whether a hosted model API is acceptable, in which region, under which contract, or whether models must run on the client's own hardware.

For financial entities in the EU, the Digital Operational Resilience Act has applied since January 2025. It treats reliance on external ICT providers as a risk to be managed, documented and exited if needed, and a hosted model API is that kind of dependency. So hosting, the approved list of models and providers, and the exit route are settled in writing before design, and they go into the contract. The architecture follows from them.

One consequence people underestimate: develop where you will run. Develop on a hosted model and deploy on a local one, and the two will behave differently in ways tests do not fully catch. The agency's system was built on the same kind of model, in the same kind of environment, as production.

Each step sees only the fields it needs

Data minimisation is usually written as a policy. We treat it as a property of each step. A triage step that decides which team should handle a ticket needs the ticket text and the list of teams. It does not need the requester's identity number, their ticket history or the attachments. Deterministic code assembles a view with only those fields before the model is called.

This reduces what could leak, and it makes the data flow explainable in one table: step, fields in, fields out, where it runs. That table is usually the first thing a data protection officer asks for, and producing it from the design is far easier than reconstructing it from logs.

We do not rely on scrubbing personal data out of free text before sending it elsewhere. Reliable redaction of names, addresses and identifiers in unstructured text is harder than it looks, and a redaction step that works on most tickets has still sent some personal data out. Free text is processed where the data is allowed to be.

Prompts, thresholds and models are code

Changes to production systems go through change control: request, test, approval, record. Prompts, thresholds and model versions are part of the production system, but they look like configuration, and anyone with access can edit them in a minute.

In a regulated environment they are code. A prompt change is tested against the acceptance set and approved like any release. Model versions are pinned and change only through a release. At the agency the environment enforced this on its own: code, libraries and model weights entered only through the approved transfer process, so nothing could change silently. Elsewhere we enforce it by making the configuration store part of the release.

One item is easy to miss: the acceptance set itself. Every change is approved because it passes the acceptance set, so whoever can edit the set can quietly lower the bar. Cases are added through the same change process as code, and removing one needs the process owner's sign-off and a written reason.

Every release carries an evidence pack

When a regulator, an internal auditor or a risk committee asks how the system works, the answer should already exist. Each release carries an evidence pack: the data flow table, the models and versions in use and where they run, the acceptance results compared with the previous release, the approval record and the list of changes. The release pipeline produces it. Nobody writes it for the audit.

The pack also makes change safe for the client's own team after handover, because it records what working looked like on the day of release.

Logging and minimisation pull in opposite directions

This tension is real and rarely stated. Auditability pushes towards logging everything: every input, output and intermediate step. Minimisation and retention rules push towards keeping as little personal data as possible for as short a time as possible. A system that logs full ticket text forever is very auditable and very hard to defend under data protection rules.

We resolve it by separating what a decision needs to be explained from what it needs to be replayed. The case record keeps the decision, references to the evidence, the model and rule versions, and a fingerprint of the input. The full input stays in the source system under its existing retention rules, and the record points to it. When the source record is deleted on schedule, the decision can still be explained, though no longer replayed with the original text. The client agrees that trade-off explicitly, per workflow, and it is written into the design.

The controls outlast the model

The air gap turned out to make the agency's system easier to own: pinned versions, no external calls, no subscription that could be repriced, nothing leaving the building. The model inside could be replaced tomorrow. The hosting decision, the data flow, the change control and the evidence would stay, because in a regulated operation those are the system.

Discuss Your Operation With Our Engineers.

Describe one workflow and the systems it relies on. A senior engineer responds within two business days with an initial assessment: what we would build, what we would not, and why.

A senior engineer reads every request and replies within two business days.

Or book directly: discovery call calendar