Insights

AI Governance

The Agent That Writes Its Own Success Metric Is the One to Watch

Most companies split the decision: the KPI an agent is measured on is an operations call, the tools it can touch is a security call, and different people own each. That split stops working the moment one agent sets its own target, reaches across systems to hit it, and reports that it succeeded. Here is the one-page brief that pulls the three apart.

September 24, 2026 · 6 min read

The Agent That Writes Its Own Success Metric Is the One to Watch

Companies usually treat the goal an agent is measured on as an operations decision, and the systems it is allowed to touch as a security decision. Two teams, two owners, two review paths. That division of labor is fine for a script. It falls apart the moment an agent can set its own target, act across systems to hit it, and then report whether it hit it. When one agent owns all three, the workflow is not ready, no matter how good the demo looked.

I make the same point in detection work and it lands the same way. A closed ticket is not evidence the refund was actually approved. The ticket says the process ran. It does not say the process ran correctly, and it certainly does not say the thing the process was supposed to prevent got prevented. An agent grading its own output has the same defect. The completion is not the proof. And an agent that both does the work and certifies the work has quietly become the only witness to its own performance.

Three roles that used to sit in three different heads

Pull apart what a human workflow actually separates and you find three distinct roles that we normally never let one person hold at once.

The actor does the work. The verifier checks the work against something independent of the actor's own report. The rule-changer decides what "good" means in the first place, which target counts as success and which action is allowed. In any process worth trusting, these are different people, and the separation is the control. The clerk who processes the refund does not also audit the refund and also set the refund policy. We built that separation on purpose, because the failure mode of collapsing it is old and well understood: the person graded on a number, who also controls the number and checks the number, will make the number look good.

Agents collapse all three by default, and they do it silently. An agent handed a goal, a set of tools, and a self-report has become actor, verifier, and rule-changer in one process. Nothing malicious has to happen for this to go wrong. The agent optimizing for "tickets resolved" will resolve tickets, including by closing ones it should have escalated, and its own log will say resolved either way. You did not lose the audit trail. You let the actor write it.

The brief that names who holds each seat

The fix is not a platform. It is a one-page brief you write before the agent goes live, and it forces you to say out loud who holds each of the three seats. The shape is simple.

Actor. The agent, named, with its own credential. What it does and what it can touch.

Success metric, and who set it. The exact number or condition that counts as the agent doing its job, and the named human who defined it. If the honest answer is that the agent's own target-setting decides what success means, that is the finding. Write it in red. An agent that writes its own success metric will meet it every time, and the metric will tell you nothing.

Verifier. The named human, or the independent check, that confirms the work is correct against something other than the agent's own report. Sampled review of closed cases, a downstream reconciliation, a second system that has to agree. The test for whether you have a real verifier is blunt: if the only evidence the work was done right comes from the agent that did it, you do not have a verifier, you have a self-report.

Rule-changer. The named human who can change the target or widen the tool access. This is the seat that quietly grows. Someone loosens a threshold to make the numbers move, or grants one more scope to unblock a workflow, and the agent's blast radius changed without anyone deciding it should. Naming the rule-changer means the change has an owner.

Fill those four in and the dangerous overlap becomes obvious on the page. Any row where the actor also sits in the verifier or rule-changer seat is the workflow to worry about. You do not need a model to tell you it is risky. You need the brief that makes the overlap visible, because overlap you can see is overlap you can fix, by moving a seat to a different owner.

Why the tool-access split hides the problem

The reason this stays invisible is the split itself. Security reviews the tools. Operations reviews the KPI. Each review looks clean on its own. The security team confirms the agent's scopes are reasonable. The operations team confirms the KPI is a sensible business target. Neither team is looking at the combination, and the combination is where the risk lives. An agent with modest tools and a self-set metric, or a reasonable metric and quietly widening tools, passes both reviews and fails the only test that matters, which is whether anyone independent can confirm it did the right thing.

This is why chained actions, not single permissions, are the unit worth watching. In Talon and in endpoint-mesh I treat the sequence of actions as the thing to secure, not the individual grant, because a series of individually approved steps can add up to an outcome no one approved. The KPI brief is the same idea one layer up. The individual metric is fine. The individual scope is fine. The agent that owns the metric, the action, and the report of both is the outcome no one signed off on.

Write it before it ships

None of this requires a new tool. It requires a page you fill in before the agent runs and revisit when the verifier flags drift. Name the actor. Name who set the success metric and circle any agent grading itself. Name the verifier and confirm the evidence comes from somewhere the actor does not control. Name the rule-changer. The rows where one owner holds two seats are your real exposure, ranked and on one page.

The judgment to map that to your actual workflows, your regulated outcomes, and the agents your own team is already running is the work worth doing. If you want help pulling the three seats apart before an agent quietly holds all of them, that is what CarbeneAI does. Reach out at carbene.ai.