Insights

AI Governance

Three Law Enforcement Disciplines That Keep AI in Security Operations Honest

Chain of custody, probable cause, and witness reliability are three law enforcement disciplines that apply directly to AI-augmented security operations. The tooling is new; the instincts that keep you out of trouble are not.

June 4, 2026 · 6 min read

Three Law Enforcement Disciplines That Keep AI in Security Operations Honest

I spent 13 years in law enforcement before I worked a single security alert. Three things I learned then are what keep me grounded now in AI-augmented security operations.

The tooling has changed beyond recognition since I started. The instincts that keep you out of trouble have not.

One: Chain of custody is non-negotiable, even when the "evidence" is an AI output

In law enforcement, if you cannot show what input produced what output, the conclusion does not hold. A confession is not usable if you cannot show the conditions it was given under. A DNA match is not usable if the sample's handling is not documented. A photograph is not evidence if you cannot account for who had it from the moment it was taken until the moment it landed in court.

Every police officer learns this discipline by their first year. Document everything. Sign every chain-of-custody form. Never break the seal on an evidence bag without a witness. The discipline is not bureaucratic ritual. It is the only thing that makes the evidence usable when it actually matters.

AI outputs are evidence too. When an AI summarizes a body-cam clip, triages a SIEM alert, or scores a transaction for fraud, the decision-makers downstream need to know:

  • What the model saw (full input, not just a transcript)
  • What it returned (full output, not just a flag)
  • What version of the model produced it (because model behavior changes)
  • Who acted on it (and what action they took)
  • When (timestamp on every step)

Without those five fields logged on every AI inference touching production decisions, your AI is making evidentiary claims you cannot defend. I see a lot of teams shipping AI features today without the logging discipline a competent incident response program would require for human-generated evidence. They are shipping evidence with no chain of custody.

The fix is straightforward. Treat every AI inference like an evidence event. Log it as if you might need to defend it under cross-examination. Because eventually, you might.

Two: Probable cause is not proof beyond a reasonable doubt, and AI does not tell you which one you have

A patrol officer learns quickly that "this is suspicious" is not the same standard as "this person committed a crime." Different evidentiary standards trigger different actions.

  • Reasonable suspicion lets you stop and question someone.
  • Probable cause lets you arrest or search.
  • Proof beyond a reasonable doubt is what you need to convict.

These are not aspirational. They are bright lines that determine what you can legally do at each stage of an investigation. Cross them in the wrong direction and the entire case gets thrown out. Cross them in the right direction at the wrong time and someone walks free because the investigator did not earn the next level of standard.

AI alerts blur these standards in a way that is quietly dangerous. An anomaly score of 0.87 is something. But what is it? The model does not tell you whether you are looking at:

  • "Something here is statistically unusual" (reasonable suspicion)
  • "This specific user appears to have exfiltrated this specific file" (probable cause)
  • "This employee has stolen company data" (proof beyond a reasonable doubt)

These are three very different conclusions with three very different remediation actions. An "investigate further" response to suspicion is appropriate. A "ban this account immediately" response to suspicion is not. An "escalate to legal" response to a 0.87 anomaly score is not.

Teams that treat AI scores as proof rather than as triage signals will make confident mistakes at machine speed. The discipline from law enforcement is the same discipline that works here: name the standard your evidence meets, and take only the action that standard supports.

Three: Witnesses are unreliable, and so is the model

The single most reliable thing about eyewitnesses is that they will give you a confident account of something that did not happen. This is one of the first surprises of detective work. Memory is reconstructive, not recorded. Bias fills in the gaps. Suggestion shapes recall. A witness who is absolutely certain of what they saw is often wrong about details that turn out to be load-bearing in court.

The defense police investigators develop is straightforward:

  • Cross-check the witness account against physical evidence
  • Look for internal inconsistencies in the witness's own statements over time
  • Triangulate with other witnesses
  • Never let confidence in delivery override your verification discipline

AI hallucination is the same problem class. The model gives you a confident, plausible-sounding output that is not true. The output sounds right. The phrasing is professional. The structure matches what you expect a correct answer to look like. And it is wrong.

The defense is the same:

  • Cross-check the AI output against ground truth (logs, raw data, original source)
  • Look for inconsistencies between what the AI claims and what your other tools show
  • Triangulate (the same query asked of two different models often produces revealing differences)
  • Never let confidence in delivery override your verification discipline

Healthy skepticism is the right default. Trust but verify is the right operating mode. The AI engineer building these systems and the SOC analyst consuming them both need to understand that confidence is a UI feature of the output, not a measure of correctness.

Why this matters now

AI-augmented security operations is one of the most-hyped categories of the past three years. The promise is real: faster triage, better pattern detection, summarization of overwhelming alert volumes. The risk is also real, and the security teams shipping these capabilities without the chain-of-custody discipline, the evidentiary calibration, and the verification discipline are creating new categories of breach they have not seen before.

The categories that worry me most:

  • AI-generated detections that get acted on without enough verification, leading to false-positive responses that erode user trust
  • AI-summarized incident timelines that quietly omit critical events because the summarizer dropped them as low-relevance
  • AI-assisted threat hunts that follow a confident hypothesis to a wrong conclusion because no one cross-checked the model's reasoning
  • AI-augmented IR communications that go to executives or regulators with errors the team did not catch because the prose was so confident

None of these failures are inherent to AI. They are inherent to deploying AI without the disciplines that keep human judgment honest. The disciplines exist. They were developed over a century in another profession that already knew confident-but-wrong was the default failure mode.

Fractional CTO and CISO leadership for companies putting AI to work: strategy, governance, cost control, and risk in business terms. Our team has led cyber defense, compliance, and risk programs for 20+ years across 6 countries and multiple industries, including healthcare, fintech, retail, manufacturing, telecom and consulting, and delivered large-scale security and compliance programs at Accenture, Dell, EY, Booz Allen Hamilton and AT&T. Technology and security leadership in one seat, reported in business terms. Talk to us.