AI Governance
Build the Healthcare AI Audit Trail Before Anyone Asks For It
The demand for an AI audit trail arrives one of three ways: a health-system security review, an OCR inquiry, or your own board. In all three, we log the API calls is not enough. You need a per-inference record, and the vendors who build it first win the review instead of scrambling through it.
September 14, 2026 · 6 min read

If you ship an AI feature that touches protected health information, the request for an audit trail is coming. It arrives one of three ways. A health-system customer runs a security review before they sign. The Office for Civil Rights opens an inquiry after an incident. Or your own board asks a simple question that turns out to be hard: who saw what, and who was accountable for it. In all three, the answer "we log the API calls" is not enough, and everyone in the room knows it the moment it is said.
The vendors that build the audit trail before they are asked win the review instead of scrambling through it. That is the whole argument. The work is not large, and it is far cheaper to do now than under a deadline set by someone else. So build it now, while it is a quarter of engineering work and not a condition of closing a deal or answering a regulator.
Why API logs are not an audit trail
Most teams already log requests. That is not the same thing as a per-inference record you can defend. An API log tells you a call happened. It does not tie a specific model output to the specific PHI it touched, the human who was accountable for the result, and the business associate agreement that made the whole thing lawful. When a reviewer asks "show me every time your AI touched this patient's record, and who checked the output," a stack of request logs cannot answer without a forensic project.
The unit that can answer is one record per inference, or per agent action. Each record is a small, complete story: when it happened, who or what acted, which model under which deployment, what class of PHI it reached and why, under which BAA, and whether a human reviewed the result before it went anywhere. Store those immutably, make them queryable by patient, and the reviewer's hardest question becomes a query instead of a scramble.
What one record has to carry
A useful record has a handful of fields that do real work. Skip these and the log looks complete but answers nothing.
The actor, and when an agent acts, whose clinical authority it acts under. A scribe agent drafting into a chart with no accountable human behind it is a finding waiting to be written up. Record the on_behalf_of so the accountability is in the data, not in someone's memory of how the feature was supposed to work.
The PHI classes touched, and the minimum-necessary justification. HIPAA's minimum-necessary standard is a per-use question, not a policy you set once. Record which classes the inference actually reached and why, so you can answer it after the fact instead of reconstructing intent months later.
The hashes of input and output, not the raw text. Hash what was processed, do not copy it into the log. The hashes prove what the model saw and produced without turning your audit trail into a second copy of the exact record set you are trying to protect. An audit log that is itself a PHI breach risk is not a control, it is a liability.
The BAA reference and the training-use bar. Tie each inference to the agreement that made it lawful, and record the contractual bar on training use. This is the field a health-system security review is really asking about, even when the question is phrased more gently. Being able to point to it per inference is the difference between a confident answer and a nervous one.
The human review decision, and how much the human changed. Distinguish "a human accepted this verbatim" from "a human rewrote it." Track the edit distance over time. When it trends toward zero on a high-tier action, your review has quietly become a rubber stamp, and that is worth knowing before an incident makes the point for you.
The hash chain. Link each record to the one before it so the log is tamper-evident. An audit trail you can silently edit is not evidence. This is the cheapest field to add and the one a serious reviewer checks for.
The evidence is also a live alarm
Here is the part teams miss. The same records that satisfy a review after the fact carry a signal you can act on in real time. The condition worth catching is precise: an AI action that touched PHI at a high tier, wrote to the EHR or sent something external, and never got the human review it required.
In a well-formed schema that is a single query. Any record where the action tier is write_ehr or send_external and the review is still pending past your threshold is an agent acting on a patient record without the accountable human the design assumes. Route that alert to the clinician named in the record, not to a shared queue where it dies. The point is not the alarm. The point is that an AI action reaching a patient record unreviewed becomes an event someone owns in minutes, instead of a finding a reviewer discovers months later. I published a ready detection rule for Wazuh and Specter alongside the schema so this is copy-and-adapt, not build-from-scratch.
Start with six fields
If the full schema is more than you can ship this quarter, there is a non-negotiable core that still answers the questions that matter: the time, the actor and who they acted for, the model and its deployment, the PHI classes and patient reference, the BAA reference, and the human review decision. Those six answer the six questions every reviewer and every regulator asks. When, who, which model, what PHI, under what contract, and who checked it. Ship those first and extend from there.
I published the full schema, a worked example record, the six-field minimum, and the detection rule as a free template in the CarbeneAI AI Governance Toolkit. It is CC BY 4.0. Copy it, adapt it, and delete what does not apply to you. It is not legal advice and it is not a substitute for your counsel or your privacy officer. Review it against your own BAAs and applicable state law before you rely on it.
The judgment to map a schema like this to your specific BAAs, your EHR, and your state's requirements is where the real work happens. That is the conversation worth having. If you want help having it, that is what CarbeneAI does. Reach out at carbene.ai.