Insights

AI Governance

A Model That Reads CT Scans Just Shipped With Its Weights. That Is the Guardrail

Alibaba's DAMO Academy open-sourced RADAR, a model that reads abdominal CT scans across 146 findings and beat 23 of 26 radiologists in the study. The interesting part is not the accuracy. It is that anyone can download it and check the claim.

September 25, 2026 · 6 min read

A Model That Reads CT Scans Just Shipped With Its Weights. That Is the Guardrail

Alibaba's DAMO Academy released RADAR, a vision-language model that reads contrast-enhanced abdominal CT scans and reports findings across 18 organs. In the Science paper it averaged an AUC of 0.913 across 146 clinical findings on roughly 40,000 real-world examinations. Against 26 radiologists from multiple hospitals, it scored higher on average than 23 of them. Radiologists working with it caught 10 percent more of what they would have missed and read faster by more than 30 percent.

Those numbers will get the headlines. They are not what I want to talk about.

The code is on GitHub under Apache 2.0. The checkpoints are on Hugging Face under CC BY-NC-SA 4.0. The model card carries a plain sentence saying RADAR is for research only and needs prospective clinical study before anyone deploys it on a patient. You can download it tonight, run inference on the demo volume, and see for yourself whether 0.913 holds on data the authors never touched.

That is the thing worth arguing about. Not whether the model is good. Whether the release shape is the one we want to encourage.

The restriction reflex picks the wrong target

Every time a capable model lands in a high-stakes domain, the reflex is to ask how we stop it. Gate the weights. License the researchers. Put the frontier behind an API where a vendor can log who asked what. The argument is that a model which can read a tumor can also be misused, so containment is the safe default.

I think that gets the risk backwards in medicine specifically. The failure mode that actually hurts patients is not a stolen model. It is a model that works on the population it was trained on and quietly degrades on yours, inside a product you bought, where nobody can look.

You cannot audit an API. You can measure it, which is not the same thing. You send in scans, you get back findings, you compute your own accuracy on whatever labeled set you have. What you cannot do is ask why the model flagged a lesion, whether the training distribution resembles your patient mix, what the calibration looks like at the operating point your radiologists actually use, or whether last Tuesday's silent model update moved the threshold under you. The vendor knows. You do not. When something goes wrong you are in a conference room reconstructing a decision from a vendor's summary of their own behavior.

With weights in hand, every one of those questions becomes work you can do. Hard work. Real work. But work, not a request for permission.

Open does not mean unbounded, and this release shows it

The RADAR release is not a shrug. It is a set of choices, and they are legible.

The code is Apache 2.0, so build on it. The weights are CC BY-NC-SA 4.0, which means no commercial use and share-alike on derivatives. Someone made a deliberate decision that a company should not take this and sell it without contributing back. That is a guardrail. It is written down, it attaches to the artifact, and a downstream user can read it in ten seconds.

Then there is the research-only language. A team that just published in Science, with numbers that beat most of the radiologists they tested against, wrote in their own model card that this is not ready for clinical deployment. They drew the line themselves, at the exact place where a marketing department would have blurred it.

Compare that to a closed medical AI product, where the boundary between validated and aspirational is a sales conversation and the evidence is a case study the vendor commissioned. Which of those two is actually more constrained?

The point is that constraint and openness are not opposites. The RADAR release ships the capability and the caveats in the same package, and both travel with the file. A closed release ships the capability and keeps the caveats in a room you are not in.

What a guardrail has to do to count

I keep coming back to a simple test. A guardrail is only a guardrail if someone outside the organization that built the system can check whether it held.

By that test, most of what gets called AI safety in enterprise settings is not. A policy document is not checkable. A vendor attestation is not checkable. A model card with no artifact behind it is a press release with a table in it. Each of these might describe something real, but you are taking it on faith, and faith does not survive a bad outcome and a deposition.

Things that pass the test: published evaluation code you can rerun. Weights you can probe on your own held-out data. A license whose terms bind regardless of who is reading. A decision log showing what the system did and what a human did about it. An error rate measured on your population, not the vendor's.

Notice that every item on that list gets easier when the model is open and harder when it is closed. That is not a coincidence, and it is why I think the restriction reflex, applied at the weights layer, buys almost no safety while costing all of the verifiability.

Where the real gate belongs

None of this means a hospital should pull RADAR off Hugging Face and wire it into a reading room next quarter. It means the gate belongs at deployment, not distribution.

The deployment gate is where the questions are answerable. Who signs off that this model is appropriate for this patient population. What happens to a finding the model produces and a radiologist disagrees with. Who is on the hook when a miss happens with the model in the loop. How often you re-measure, and against what. What the rollback looks like when the numbers drift. Those are institutional questions with institutional answers, and none of them are improved by the model being secret. Several of them are only answerable because it is not.

The thing I would not want is a world where a research team with results like this decides the safe move is to publish the paper, keep the weights, and license access to three health systems who can afford it. That world is worse on every axis I care about. Fewer eyes on the model, fewer independent replications, fewer places the failure modes surface before they surface on a patient. The capability still exists. Only the scrutiny is gone.

AI that gets to thrive is AI that people can check. Build the checking capacity, put the gate at the point of care, and let the models be open enough that the check is possible at all.