AI systems in radiology are dangerously overconfident – and that's a serious problem for clinical practice. The RadLE 2.0 benchmark, developed by the CRASH Lab at Ashoka University in India, has just demonstrated that leading AI models deliver false findings with full conviction while simultaneously failing to recognize when they should stay silent.
Key Findings
- Human radiologists scored 988.7 out of 2,000 points, the best AI model only 758 points
- Claude Fable 5 (Anthropic) was most reliable on safe answers; Gemini 3 Pro (Google) had highest raw accuracy
- Meta's Muse Spark 1.1 best recognized when to hand cases to humans
- The benchmark rewards honesty: wrong answers with high confidence result in point deductions – mirroring real medical consequences
The Core Problem: Confidence Over Competence
The RadLE 2.0 test measures three dimensions: whether the model gets the diagnosis right, how confident it is in that answer, and whether it can admit when it's out of its depth. AI systems must rate answers on a confidence scale from 0 to 4 – and are explicitly allowed to say "I don't know."
Here's the problem: many don't. They prefer to guess with high confidence rather than admit uncertainty. In radiology, this isn't an academic question – a confidently wrong diagnosis can be fatal for patients.
The scoring system punishes exactly this behavior: correct answers with high confidence earn full points. Wrong answers with high confidence lose corresponding points. Saying "I don't know" costs zero points but causes no harm.
Different Strategies, Different Risks
| Model | Strength | Problem |
|---|---|---|
| Claude Fable 5 | Reliable, safe answers | Overall more conservative |
| Gemini 3 Pro | Highest raw accuracy | Overconfidence amid uncertainty |
| Muse Spark 1.1 | Best handover recognition | Fewer hallucinations through frequent refusals |
| Grok 4.5 | More knowledge | Significantly more hallucinations, more convinced of wrong answers |
A particularly alarming example: Grok 4.5 hallucinates far more than its predecessor – not because it knows less, but because it's more confident in its false answers.
Open-source and medical-specific models show the problem most starkly. They attempt to answer nearly every case but frequently get it wrong – usually with medium to high confidence. According to the research team, several of these models would have scored much better if they had stayed quiet more often.
What This Means for Practice
The study addresses a fundamental problem recently raised in highly cited research: as long as benchmarks only reward accuracy, AI models are trained to guess. In medicine, that's catastrophic.
For German enterprises and clinics, this means: before AI systems can independently diagnose in radiology, they must learn to accept their limits. A model that's right 90 percent of the time but wrong 10 percent – and confident in both cases – isn't suitable for clinical use. The question isn't just "How accurate is the AI?" but "Does the AI know when it's not accurate?" Regulation and certification must take this seriously – otherwise AI systems become a risk rather than a tool.
Sources
Editorially owned by Ideal Syka. Sources and method: Newsroom & method. Tips and corrections: ai@i6eal.de.




