Insights

AI Governance

Is This Tool Call Allowed? Answer It at Runtime, Not in a Policy PDF

An agent has made thirty tool calls since you last looked. A policy document cannot tell you whether the thirty-first should run. That decision has to happen at runtime, per call, and it has to leave a record a board can read. The artifact is a decision log, not another GRC binder.

September 15, 2026 · 6 min read

Is This Tool Call Allowed? Answer It at Runtime, Not in a Policy PDF

An agent is working. It has made thirty tool calls since you last looked at it. Somewhere in there it read a file, hit an API, wrote to a database, sent a message. Now it wants to make the thirty-first call. Here is the only question that matters at that moment: should this specific call, with these specific arguments, be allowed to run right now?

A policy PDF cannot answer that. The document sitting in your governance folder says the agent may access customer records for support purposes and must not exfiltrate data. Fine. But the agent in front of you is about to call a tool named send_email with a body it assembled from the last three things it read. Is that the allowed support use, or the forbidden exfiltration? The policy does not know. The policy was written months ago by someone who never saw this call. The decision has to be made now, at runtime, by something that can see the actual call and the actual arguments, or it does not get made at all.

Policy is a claim, runtime is a control

There is a difference between a policy and a control, and agentic systems make the difference expensive. A policy is a statement of intent: here is what should happen. A control is a mechanism that makes it happen or stops it. A PDF is a policy. A function that inspects each tool call and returns allow or deny before the call executes is a control.

Most AI governance today is policy all the way down. Organizations write the document, get it approved, file it, and point to it when asked how they govern their agents. Then the agent runs, makes its thirty calls, and nobody checks any of them against the document, because the document is prose and the calls are code and nothing connects the two. The governance exists on paper and nowhere else. When a security review or a regulator asks to see governance in action, the honest answer is: we have the policy, we do not have evidence it was ever applied.

The fix is to move the decision to where the call happens. Before a tool executes, a runtime check reads the tool name and the arguments and decides. This is an allowlist, not a wish. read_customer_record for a ticket in scope: allow. send_email to an external domain: deny, or escalate to a human. The allowlist is short, specific, and testable, and it runs on every call whether or not anyone is watching.

The decision log is the artifact

The allowlist stops the bad call. The log proves the allowlist ran. That second half is what turns runtime enforcement into something a board or an auditor can actually use, and it is the part teams skip.

For each tool call, record five things:

  • Tool name. What the agent tried to do.
  • Argument hash. A fingerprint of the arguments, so you have evidence of what was passed without logging the raw sensitive payload in plaintext.
  • Decision. Allow or deny.
  • Rule. Which allowlist entry made the decision, so the log explains itself.
  • Goal-drift note. One line on whether this call still moves the agent toward the task it was given, or looks like a detour.

That last field is the one worth dwelling on. An agent thirty calls deep can be individually compliant on every call and collectively lost, grinding through allowed actions that no longer serve the original goal. A per-call note on semantic drift, even a crude one, is the difference between a log that says every call was permitted and a log that says the agent stopped making sense at call nineteen and here is where. The first is an alibi. The second is oversight.

Keep the last N calls of this log per agent run, and you have a record you can hand to a reviewer: here is every action the agent took, whether it was allowed, which rule allowed it, and whether it stayed on task. That is governance you can show, not governance you can cite.

This is a smaller build than a GRC suite

The instinct when a board asks about AI governance is to go shopping for a platform. Resist it. A runtime allowlist and a decision log are a small amount of code that sits in front of your agent's tool-calling layer, not a six-figure GRC purchase and a year of rollout. You already have the tool-calling layer. The allowlist is a function it calls before each execution. The log is an append-only record it writes after each decision. A working version of both fits in an afternoon, and it produces something a big platform often does not: evidence tied to actual calls rather than a dashboard of policies nobody enforced.

If you want a place to start writing down which actions an agent may take, the allowed-actions and human-in-the-loop rows in the open AI Governance Toolkit are the paper version of the same allowlist. Fill the row in first, per agent, before it runs: identity, data class, allowed actions, the human-in-the-loop line, and a tested stop. Then the runtime allowlist is just that row turned into code, and the decision log is the proof that the row was honored on every call.

Why the buyer asks

If you sell an agentic feature, the runtime-decision question is coming from your buyers, and it is the sharp one. A security-conscious buyer, a health system, a regulated financial firm, a public-safety agency, will not accept "we have an AI governance policy." They will ask what happens on the specific call that touches their data, and whether you can show them it happened the way you claim. The vendor who can produce a decision log wins that conversation. The vendor who can only produce a policy document gets the follow-up meeting, the security exception, and the delay.

Runtime enforcement plus a readable log is the artifact that ends the argument. It is small to build and it is exactly what the board wanted when it asked whether anyone owns the agent. Not another binder. Evidence, per call, that the answer to "is this tool call allowed" got made by something that was actually looking.