ONDERZOEK·July 23, 2026·6 min read

Four Types of Medical AI Benchmarks, and Why the Score Alone Is Not Enough

Delen
Four Types of Medical AI Benchmarks, and Why the Score Alone Is Not Enough

Written by Nader Absi, MD

If you give a medical AI model a standard case vignette, it performs strongly. Convert those same cases into interactive clinical conversations, and some models keep less than one tenth of their original diagnostic accuracy [1].

That finding, from AgentClinic, published in npj Digital Medicine, points to a specific problem in how clinical AI results are reported: scores are often quoted without the test format that produced them [1].

One useful way to read current clinical AI testing is through four types of benchmark. The type may matter as much as the number.

1. Static medical question answering

The relevant clinical information has already been collected and organised. Frontier models perform strongly here, but that mainly tells us what a system knows, not whether it can gather the right information itself.

2. Rubric-scored responses

The AI responds to open-ended health questions and is scored against physician-written rubrics [2]. HealthBench Professional also adjusts scores for response length, because a longer answer has more opportunities to satisfy individual scoring criteria [3].

3. Workflow tasks

The model is tested across everyday clinical work, including chart review, documentation, patient communication, and administrative tasks. Some evaluations use private hospital datasets, which helps reduce the risk of benchmark contamination [4].

4. Interactive clinical environments

The model does not receive the complete case upfront. It has to interview the patient, request investigations, navigate the record, and work out what information is still missing.

Performance falls sharply here. In a separate 2026 evaluation of agent systems on AgentClinic, the strongest setups reached roughly 28 to 60 percent diagnostic accuracy depending on the dataset [5]. Despite access to tools including web browsing and code execution, their advantage over the best-performing baseline models did not reach statistical significance [5].

What this means for reading a score

So when a vendor quotes a high accuracy rate, the percentage alone is not enough. You need to know whether the model received a completed case or had to gather the information, which benchmark, subset and configuration were used, and when the evaluation was run.

That is why every DR. INFO evaluation reports its benchmark, subset, benchmark version, model configuration and evaluation date alongside the result [6]. A score is evidence of performance under specific conditions, and those conditions determine what you can conclude from it.

Four Types of Medical AI Benchmarks Explained | DR. INFO