Your AI agent controls are aspirations until you can test them
AI agent controls only count if you can test them. A context layer is really a control layer, with every rule tied to a passing test and a line of defence.
Your AI agent controls are aspirations until you can test them
Call it enablement if you like. To a regulator, the thing you built is a control layer, and the only AI agent controls that count are the ones you can run on demand.
We have laid out the architecture. Now comes the reckoning a regulated enterprise cannot skip. Everything so far has been framed as enablement: helping agents act with the right context. Look at the same layer through a risk officer's eyes and it is something else: a control layer. Its business value, faster and more consistent agent delivery, is only bankable if every agent output is authorised, grounded, explainable, and auditable, per request. AI agent controls that cannot be demonstrated are not controls; they are intentions with good grammar. An excellent capability that cannot be governed, audited, or explained is, in a bank, a liability the moment it acts.
So the honest test of this whole enterprise is not 'does the agent give good answers'. It is 'can you prove, to someone whose job is to doubt you, that the controls operate'. That reframes the work, and it is where a great deal of AI governance falls over, because most of it is written in the aspirational voice: policies that should be followed, access that ought to be restricted. A risk officer has seen enough of those documents fail to hold up the moment something actually goes wrong, and reads the next one with the same doubt.
AI agent controls in the binding voice, tied to tests
The difference between a control and an aspiration is testability. A control is a must, it names who owns it, and it comes with an acceptance test someone independent can run. Treat the context layer that way and it resolves into a catalogue of specific obligations, not a set of aspirations.
A few, to make it concrete. The layer must expose exactly one governed entry point, with no bypass and no fail-open path, verified by showing that no code path returns context without passing policy, redaction, grounding, and quality. Every request must bind the accountable principal (a person, or a governed non-human identity with a named owner), the agent identity and version, declared purpose, channel, and a correlation ID, and the layer must reject any request missing them; a request that omits them fails validation. An unauthorised purpose must return a denial with empty context and empty evidence, verified by exactly that test. Sensitive fields must be stripped from real data in line, verified by the field being absent from the output and a redaction constraint present. High-impact actions must route to a named human, and the agent must not self-approve. The run ledger must be written by the layer, not the agent, and be hash-chained so tampering is detectable; the chain validates intact and fails after a single altered record.
Writing them this way earns its keep: each obligation traces back to one of the properties the design was derived from, and forward to a test that either passes or does not. When a third-line auditor asks whether redaction actually happens, the answer is not a paragraph of reassurance; it is 'run this test and read the output'. AI agent controls you can execute are controls you do not have to be believed on.
A control you can only describe is a promise. A control you can run is evidence. Auditors have learned to tell the difference, and so should the people buying AI platforms.
The residual risks worth stating honestly
The other half of a real control model is the risk register, and the tell of a serious one is that it does not claim everything is fine. Several risks here reduce nicely: over-broad context assembly, a policy engine outage tempting a fail-open path, ledger tampering: these fall to low, or near it, once the enforcement, default-deny, and hash-chaining controls operate as designed.
Others do not, and pretending otherwise would be the giveaway of a naive model. Prompt and indirect injection remains a medium risk, because it is adversarial and evolving, and no layer fully engineers it away. Model behaviour beyond its evidence, asserting past what it was grounded on, stays medium even with quality gating and human approval. Approval fatigue is a human risk: a control that routes to a person is only as strong as a person who genuinely reviews rather than rubber-stamps, so it needs sampling by the third line to stay honest. And concentration risk, dependence on a small number of external model or protocol providers, is a systemic exposure a single enterprise cannot fully mitigate at its own layer, and belongs on the enterprise risk register with a named owner rather than asserted as closed. Stating those plainly is not weakness; it is the thing that makes the rest of the register credible. A model that closes every risk to green has told you it was not looking hard.
Where the framing earns its keep
Map the controls to the frameworks the enterprise already answers to and the picture stops being abstract. Human approval and quality gating speak to the EU AI Act's human-oversight and accuracy obligations; the tamper-evident ledger speaks to its record-keeping requirement, due to move from August 2026 to December 2027 for standalone high-risk systems under the May deferral agreement. The no-bypass, default-deny, safe-degradation behaviour maps to DORA's operational-resilience expectations. Agent versioning, independent review, and ongoing monitoring line up with the model-risk discipline behind SR 11-7. Redaction and purpose-binding map to GDPR minimisation and purpose limitation. None of that is a claim of certification (the exact article mapping is a job for the compliance function), but it shows the layer was built against the obligations, not retrofitted to them.
Ownership has to be real, too. The controls distribute across three lines of defence: the delivery teams that operate the gateway and build context assets, the risk and compliance function that owns policy and approval thresholds and reviews changes before they go live, and internal audit that independently re-runs the ledger tests and samples the approvals. The layer does not replace that structure; it gives each line something concrete to hold.
What testable controls look like in the room
Take a composite from a tier-1 lender's annual model-risk review. In previous years, agent governance was a slide deck: here are our principles, here is our policy library, here is a screenshot of the approval workflow. This year the third line asks a different question: show me. The delivery team runs the control suite in front of them. The redaction test removes a restricted field from live data and prints the constraint. The unauthorised-purpose test returns an empty bundle. The ledger test validates the hash chain, then the reviewer alters one record and watches it fail. The review that used to take a fortnight of document exchange takes an afternoon, because the controls answered for themselves. That is the difference testability makes: governance that stops relying on trust.
The objection: isn't this compliance box-ticking that slows delivery?
The strongest counter from an engineering leader is that this sounds like process weight (controls, registers, three lines of defence) bolted onto teams that are trying to ship, and that it will slow agent delivery to the pace of the audit function. It is a real risk, and plenty of governance programmes have earned that reputation.
The answer is that testable controls cut the opposite way. The thing that actually slows agent delivery in a regulated business is not enforcement; it is the months a pilot spends in review because no one can prove it is safe. A control expressed as a passing test is faster than a control expressed as a document, because it turns a subjective argument into a re-runnable fact. The weight people fear comes from governance that is all prose and no execution: endless meetings to interpret ambiguous policy. Move the controls into tests and you spend less time arguing, not more. Box-ticking is the slide deck. The test suite is the thing that lets you skip it.
What to do differently on Monday
For any agent heading toward a regulated workflow, ask your teams to produce two artefacts, not one. The control catalogue is familiar. The second is the one that matters: for each control, the test that demonstrates it, runnable by someone who does not trust you. Where a control has no test, treat it as not yet real, whatever the policy says. Then insist the risk register states residuals honestly: if injection, model behaviour, approval fatigue, and provider concentration are all marked closed, send it back, because someone was optimising for green rather than truth. The goal is not a longer document. It is a shorter distance between 'we have a control' and 'watch it work'.
Where this leaves us
A context layer is a control layer wearing a friendlier name. Its AI agent controls earn their keep only when they are testable, owned across three lines of defence, honestly rated, and mapped to the obligations the enterprise already lives under. Build it that way and an impressive demo becomes something a bank can actually stand behind, because its controls are provable.
Enough principle. It is time to watch the whole thing run on a single case.
Next: one fraud investigation, three requests, and a layer that grants one, refuses one, and escalates one, then proves what it did.

Ankur Chrungoo
Principal · Enterprise AI and Agentic Systems
Principal Engineer and Architect at Bugni Labs with nearly 20 years across software engineering and architecture. Focused on Enterprise AI and Agentic Systems, spanning production architecture, agent orchestration, evaluation and governance, with extensive experience in regulated financial services. MSc Artificial Intelligence, Queen Mary University of London.
The Engineering Notebook
Once a month, a long read on what we're learning building governed AI for regulated enterprises. No hot takes, no roundups.
Related case studies
- Automating evidence extraction for regulatory narrativesReducing manual effort in regulatory narratives while improving traceability and consistency.
- Authorised payment fraud: designing for speed, signals and supervisionExperimenting with multi-agent fraud detection under tight sprint constraints.
- Building a cloud-native payment and data foundation for a new digital bankFrom concept to reference architecture, ISO20022 payments, data services and open banking adapters.
You might also enjoy
Agentic AI Is Not a Chatbot With Extra Steps
Unpack why agentic AI enterprise surpasses chatbots. Explore definitions, mechanisms, financial services examples, benefits like 3-5x velocity, and misconceptions for CIOs building governed AI systems.
PerspectiveBuild vs Buy for Enterprise AI
Compare building in-house AI solutions versus buying from vendors for enterprises. Review costs, timelines, pros, cons, stats, and top platforms to decide.
PerspectiveAI Vendor Lock-In Is a CIO Problem, Not Procurement's
AI vendor lock-in is usually fought as a pricing negotiation. In regulated institutions it is an architecture and concentration-risk decision the CIO owns.