AI Governance
The FDA is proposing to test your AI the way it tests a physician
A August 2026 FDA discussion paper floats competency-based evaluation for generative AI medical devices: benchmark the deployed product, then confirm clinically. It is not guidance and it is not binding, but it tells you what your risk file will need to say.
September 19, 2026 · 6 min read

On August 18, 2026, the FDA's Digital Health Center of Excellence, inside CDRH, released a discussion paper titled Considerations for the Regulation of Generative AI-Enabled Medical Devices. It opened docket FDA-2026-N-7874 with comments due October 19, 2026.
Start with what it is not. It is not draft guidance. It is not final guidance. It proposes no binding policy, and the agency is explicit that it does not address whether the approaches described sit inside its existing legal authorities. Anyone telling you there is a new FDA requirement for generative AI devices is wrong.
What it is, and why I think product teams should read it now, is the clearest signal yet about the questions a reviewer will ask. Those questions are answerable today, and the teams that can answer them in writing will move through review faster than the teams discovering them in a deficiency letter.
Why the existing framework does not fit
The current AI/ML device framework assumes a product with defined inputs, deterministic or statistically bounded outputs, and a model trained on structured data. You can characterize the input space, test against it, and say something meaningful about performance.
Generative AI devices break every part of that assumption. They accept open-ended inputs. They perform multiple subtasks. They produce different outputs for similar inputs. They evolve through changes to models, prompts, retrieval strategies, guardrails, orchestration logic, or the interface. And many of them are built on third-party foundation models that offer limited transparency into training data, architecture, or evaluation.
The agency states the practical problem directly: you cannot test the full range of possible inputs and assess the outputs using traditional premarket methodology. That is not a criticism of anyone's engineering. It is arithmetic. An open prompt field has no enumerable input space.
Two axes, and both of them go in your intended-use statement
The paper proposes scoring a device function on two axes.
The first is how independently the function acts. The range runs from providing non-directive information, to providing action-directing information, to supervised action, to fully autonomous action. The second is how consequential an incorrect output would be.
The practical consequence for a manufacturer is worth stating plainly. Your intended-use statement would need to articulate both where the function sits on autonomy and how bad a wrong answer is, and you would need to defend both placements in your risk file. That is a different document than most teams have written. It is also a document you can draft this quarter, without waiting for anything to become binding.
The paper also flags patient-facing functions as potentially warranting heightened controls, on the reasoning that a patient has less domain expertise to recognize an error, which lets a bad output propagate further before anyone catches it. If your product talks directly to patients rather than to clinicians, assume the bar is higher.
Competency instead of exhaustive validation
The novel piece is the premarket approach, and the analogy the agency reaches for is medical credentialing. Rather than validating every input-output pair, which is impossible, evaluate whether the device demonstrates competency at its intended clinical tasks, the way medicine evaluates a physician.
That has two components. First, device benchmarking: high-throughput non-clinical testing of the device in its deployed configuration, across up to ten elements spanning safety behaviors, clinical proficiency, generalizability, and agentic conduct. Second, clinical confirmation to verify real-world performance, tailored to intended use and risk, and not requiring a prospective clinical study in every case.
Two details inside that deserve attention from anyone building on a foundation model.
The evaluation target is the final user-facing device as configured for real-world deployment, not the foundation model alone. Your prompts, your retrieval, your guardrails, and your orchestration are part of the thing being tested. A vendor's model card is not a substitute for evaluating your product.
And the comparator being floated is a panel of qualified clinicians whose consensus reflects the standard of care, or a median clinician in practice. The paper also signals that an LLM might serve as an expert adjudicator for benchmarking if it meets independence and qualification standards. Whatever you think of that idea, notice what it implies: the benchmark is a clinical performance bar, not a software correctness bar.
Postmarket is where the real cost lands
Because variable outputs limit what premarket evaluation can establish, the paper leans harder on postmarket monitoring than the existing framework does. The approaches it describes are periodic re-benchmarking against prespecified thresholds, sample-based clinician review of real-world inputs and outputs, and monitoring for performance degradation and drift as models, prompts, or components change.
This is the part I would budget for. Premarket work is a project with an end date. Sample-based clinician review of live interactions is a standing operational cost, with staffing, a sampling methodology, and a definition of what a failed sample triggers. If your business case for the feature did not include a clinician reviewing outputs indefinitely, the business case is incomplete.
The drift piece also creates a coupling most teams have not thought through. If a prespecified threshold is what keeps the device in a known state, then every model swap, prompt edit, and guardrail change becomes an event that may require re-benchmarking. Your release process and your regulatory posture stop being separate things.
What to do with this before October 19
Three moves, none of which depend on the paper becoming policy.
Place every generative function on both axes and write down the defense. If you cannot articulate why a function is action-directing rather than autonomous, that ambiguity is a finding waiting to happen, and it is cheaper to resolve now.
Write down what changes to your deployed configuration would count as material. Model version, prompt, retrieval corpus, guardrails, orchestration. If you cannot enumerate what changed between two releases, you cannot answer a drift question later.
And if you have a view, file a comment. The docket is open until October 19, 2026, and this is the stage where the people who actually build these products have the most influence over what the eventual framework asks for. A framework written entirely by people who have never shipped one is a worse outcome for everybody.
The competency framing is, I think, the right instinct. You cannot enumerate the inputs, so you test whether the thing is good at the job. That is how we decide a clinician is safe to practice, and it is a more honest model for a system with an open prompt field than pretending the input space is bounded.