AI Governance
Your AI Governance Framework Assumes the Model Only Answers
Almost every AI governance framework in circulation is built around a system that produces an output for a person to check. Agents do not do that. When the failure mode moves from a wrong answer to a wrong action, most of the controls you have stop reaching the risk.
July 31, 2026 · 13 min read

Open any AI governance framework written in the last three years, including the first version of ours, and trace where the controls actually attach. Bias testing attaches to an output. Explainability attaches to an output. Human oversight means a person reviews an output before it counts. Model cards, validation studies, drift monitoring, all of it hangs off the same object: a thing the model produced, sitting still, waiting for someone to look at it.
That is a reasonable design when the model returns an answer and stops. It is the wrong design the moment the model can act.
Agentic systems plan across multiple steps, call tools, carry memory between sessions, and take actions that change the state of real systems. The failure mode moves from a wrong answer to a wrong action, and a wrong action can be slow or impossible to undo. Review-the-output controls do not reach that. By the time there is an output to review, the refund has been issued, the message has left the building, the production config has changed.
We added a section on this to the AI Governance Framework template. It is free, it is CC BY 4.0, and nothing is gated behind it. This article is the reasoning underneath it, because the controls are easy to copy and the derivation is the part that transfers.
First, decide what you are actually governing
Half the confusion around agentic governance is that "agent" has become a marketing word. Every vendor has one. Almost none of them mean the same thing by it, and a definition you cannot apply consistently produces a policy nobody can enforce.
The test we settled on is behavioral. A system is governed as agentic if it meets two or more of these:
- Goal-directed planning. It decides its own sequence of steps toward an objective rather than following a fixed script.
- Tool and action invocation. It can call external tools, APIs, databases, or devices that change state outside the model.
- Persistent memory or state. It carries information across turns, sessions, or tasks and uses it to shape later behavior.
- Unattended execution. It can run without a person present for each step, on a schedule or in response to an event.
Two or more, not any one, and the reason matters. Any single property on its own describes things you already govern adequately. Rules-based automation invokes tools all day and is not agentic, because its action set is fixed and its behavior is deterministic. A chatbot with conversation history has persistent state and still just answers questions. The risk that justifies a separate control set comes from the properties compounding: planning plus action means the system chooses steps you did not enumerate, and memory plus unattended execution means it does so when nobody is watching and carries the consequences forward.
Applied honestly, this test catches things people do not think of as agents. An assistant that reads an incoming queue, decides which items to work, and closes them out is agentic even though a person configured it once six months ago and walked away. It also excludes things vendors sell as agents that are really scripted workflows with a language model doing the text formatting. Both of those corrections are the point.
The control that everything else follows from
Here is the load-bearing claim: an agent's real risk comes from what it can do, not from what it can say.
Once you accept that, authorization becomes the primary control surface and it has to be deny by default. An agent may invoke only the tools on its allowlist, and each tool gets its own least-privilege credential.
But the allowlist alone is not enough, because the tool is the wrong unit to make approval decisions on. "Can this agent use the ticketing API" is unanswerable. Creating a draft ticket and closing out a customer complaint with a public-facing response are the same API and are nowhere near the same risk. Approval requirements have to follow the consequence.
So classify every authorized action by what happens when it is wrong:
- Read, internal. Query a database, read a document, retrieve a record. Permitted within the agent's data scope, logged.
- Write, reversible. Create a draft, update a ticket, write to a staging area. Permitted at moderate autonomy within scope, logged.
- Write, irreversible. Delete data, send an external communication, submit a filing, change a production configuration. Human approval required at every autonomy level, unless the governance committee documents an exception.
- External-facing. Anything a customer, patient, citizen, partner, or regulator observes. Human approval or a reviewed template, plus disclosure.
- Financial. Purchase, payment, refund, contract commitment. Human approval, with hard per-transaction and per-period ceilings.
- Code and infrastructure. Execute code, deploy, modify access controls, spawn other agents. Isolated execution, human approval, and never granted to an agent that also processes untrusted external input.
Notice what that fourth column does to the autonomy conversation. Irreversible actions require human approval at every autonomy level, including full autonomy. That looks like a contradiction until you separate the two things "autonomy" is being asked to mean. An agent operating without contemporaneous oversight can still be denied the ability to do things that cannot be undone. Autonomy is about how much supervision the routine path needs. Reversibility is about how much a mistake costs. They are independent, and collapsing them is how organizations end up granting delete rights as a side effect of granting convenience.
An agent cannot be asked to respect its own budget
Spend ceilings, rate limits, and step limits are enforced by the platform, not by the agent's reasoning.
This one sounds like an implementation detail and it is actually a principle about where controls can live. The component you are trying to constrain is the component that is failing. If a runaway loop is caused by the model's judgment going sideways, and the only thing stopping the loop is the model deciding to stop, you have asked the failure to police itself. That is not a control. It is a hope with a dashboard.
The same reasoning applies more broadly than budget, and it is the fastest way to audit an agentic design. For every control in the architecture, ask where it is enforced. If the answer is "we told it not to" or "it is in the system prompt," it is not a control, it is a preference. Policy checks belong at the tool boundary where they are mechanical, not in the prompt where they are advisory.
An untested stop control counts as no stop control
Every agent above the lowest autonomy levels needs a documented stop mechanism, a named set of roles who can invoke it, and a target time from decision to stop. Tested at least quarterly and after any material change.
The testing requirement is not bureaucratic thoroughness. Stop paths are the one part of the system that normal operation never exercises. Everything else in your agent gets hit thousands of times a day and fails loudly when it breaks. The stop control gets used approximately never, which means it rots silently, and you find out it rotted on the day you need it. This is the same reason backups you have never restored are not backups.
Stop semantics have to be defined per agent, because "stop" is three different things:
- Pause in place. The agent halts and holds its state, ready to resume.
- Terminate and hold state. The agent ends, its work product is preserved for inspection.
- Terminate and roll back. The agent ends and its effects are reversed.
If an agent holds irreversible action rights, write down explicitly what cannot be rolled back. That document is not paperwork. It is the honest inventory of what your stop control does not save you from, and it should be uncomfortable to write. If it is not uncomfortable, the agent probably has permissions nobody has looked at.
Rubber-stamping is a control failure, not a metric
If oversight quality is not measured, it decays. Approval rates approaching 100 percent with falling median review time mean the human in the loop has stopped functioning as a control and has become a click.
Most governance programs count approvals. Counting approvals tells you the gate exists. It tells you nothing about whether anything is being caught. The pair of numbers that matters is approval rate together with review time, watched as a trend, because the failure has a signature: as trust builds and volume grows, the human starts approving faster and rejecting less, and the curve gets there long before anyone notices that a control has quietly turned off.
Treat that curve as an incident to investigate rather than a productivity win to celebrate. Sometimes the investigation finds that the agent genuinely is reliable and the approval gate should be replaced with sampling and a spend ceiling, which is a legitimate outcome. Sometimes it finds a person approving 400 actions an hour. Either way you learn something, and you only get to learn it if you were measuring oversight quality rather than assuming it.
The configuration that gets committee review every time
The highest-risk arrangement in the whole framework is an agent that reads untrusted external content and also holds write credentials.
Indirect prompt injection is why. Instructions hidden in content the agent retrieves, whether a document, a support ticket, an email, a web page, or a code comment, redirect what the agent does. The defense that actually holds is architectural: everything an agent retrieves, reads, or receives is untrusted data, never instruction, and that boundary is marked at the point of retrieval rather than assumed. Output filtering and approval gates on consequential actions sit behind it.
But defense in depth on injection is still a probabilistic bet, and the honest way to handle that is to make the losing bet cheap. Separate the agent that reads the internet from the agent that holds write credentials. When the business need genuinely requires both in one system, that combination goes to the governance committee before it ships, every time, and the blast radius statement gets written first: what an attacker gains by controlling this agent, and what a full malfunction costs. If that statement cannot be written, the agent is not ready for approval.
Two more that change the shape of the problem
Agents are non-human identities and have to be governed as such. An agent that borrows a person's credentials cannot be held accountable after the fact and cannot be revoked without collateral damage. Each agent gets a unique identity with short-lived, scoped credentials. Critically, its permissions are granted to the agent, not inherited from whichever user triggered it, and an agent must never be able to reach data the requesting user could not reach. That is the confused deputy problem, and it is the most common way an agent turns into a privilege escalation path that no access review will show you, because on paper every permission involved is correctly assigned.
Memory is a data store with weak boundaries and a long half-life. It gets classified at the sensitivity of the most sensitive item it holds, scoped per user and per tenant and per task, and subjected to the same access and deletion rights as any other record containing personal data. Memory poisoning, where attacker-supplied content persists and shapes later decisions, does not surface on its own. It only shows up if you review long-lived memory on a cadence and keep provenance for stored facts, meaning where the agent learned something and when.
And when agents call other agents, the governed unit is the system rather than the individual agent. One named human owns the outcome. Distributing ownership across per-agent owners produces a system where nobody is accountable for emergent behavior, which is the only behavior that matters in a multi-agent design.
What I want to be accurate about
Three caveats, because a governance document that oversells its own footing is worse than no document.
The section draws on the NIST AI Risk Management Framework, OWASP's agentic security work, and NIST's SP 800-53 control overlays for securing AI systems, whose planned use cases include dedicated overlays for single-agent and multi-agent deployments. Those NIST overlays were still in draft as of mid-2026. Confirm their status before you adopt them as a control baseline. I would rather say that here than have you find out during an audit.
On the regulatory side, EU AI Act Article 50 transparency duties, including telling a person they are interacting with an AI system, apply from 2 August 2026. Any external-facing agent in scope needs that disclosure designed in rather than bolted on. Separately, the Digital Omnibus on AI moves the high-risk obligations later, to 2 December 2027 for standalone Annex III systems and 2 August 2028 for systems embedded in products under Annex I. Those changes take legal effect on publication in the Official Journal, so confirm the operative dates with counsel before you plan against them. Dates that have not been published are not dates you can build a compliance calendar on.
Third, sector rules mostly have no agent-specific provisions yet, and that absence gets misread as permission. It is not. An agent that takes an action a regulated professional would otherwise take inherits that activity's requirements. Map the action, not the technology. If a human doing this task would need a license, a signature, a retention period, or a disclosure, the agent doing it does not make those go away.
Take it and tell me what is wrong with it
Section 10 of the framework template has the full version: a five-level autonomy scale mapped to oversight requirements, an agent registry with the fields that make revocation possible, memory governance, multi-agent trust boundaries, decision-trace logging built to a reconstructability standard, an agent-specific threat model, adversarial and shadow-mode testing, lifecycle gates, ten vendor questions, and the metrics that show whether any of it is working. There are two new appendices, an agent authorization record and an agentic red team test plan.
It is free, it is CC BY 4.0, and there is no email wall. Search and replace one placeholder with your organization's name, delete what does not apply, and you have a working draft instead of a blank page.
A template does not govern anything. People do. The judgment about which of your agents are actually high-stakes, who has authority to pause one that is drifting, and what your specific regulators and contracts require is where the real work lives. But you should not have to pay someone fifty thousand dollars to draw the starting line.
If something in it is wrong, I want to know. That is not a courtesy. Governance documents get copied, and a mistake in a template propagates further than a mistake in a deployment.