[{"data":1,"prerenderedAt":30},["ShallowReactive",2],{"nr-en-radiology-ai-models-overconfident-wrong-diagnoses":3},{"slug":4,"title":5,"dek":6,"date":7,"time":8,"publishedAt":9,"updated":10,"updatedAt":10,"dateFmt":11,"updatedFmt":10,"kind":12,"tier":13,"author":14,"authorName":15,"topics":16,"tracker":22,"trackerLabel":23,"headlineStat":24,"image":25,"ogImage":26,"imageAlt":5,"csv":10,"minutes":27,"words":28,"html":29},"radiology-ai-models-overconfident-wrong-diagnoses","Radiology AI Models Fail at Self-Doubt – New Study Reveals Dangerous Overconfidence","The RadLE 2.0 benchmark exposes a critical safety flaw: AI models deliver wrong medical diagnoses with high confidence. Human radiologists are far better at recognizing their own limits.","2026-07-19","09:38","2026-07-19T09:38:00+02:00","","July 19, 2026","daten","standard","ideal-syka","Ideal Syka",[17,18,19,20,21],"Medical AI","Radiology","AI Safety","Benchmarking","Regulation","\u002Fstand-der-ki","AI Progress","Human radiologists scored 988.7 of 2,000 points; best AI model scored 758","\u002Fnewsroom\u002Fimg\u002Fradiology-ai-models-overconfident-wrong-diagnoses.webp","\u002Fog-nr\u002Fradiology-ai-models-overconfident-wrong-diagnoses.en.png",3,537,"\u003Cp>AI systems in radiology are dangerously overconfident – and that&#39;s a serious problem for clinical practice. The RadLE 2.0 benchmark, developed by the CRASH Lab at Ashoka University in India, has just demonstrated that leading AI models deliver false findings with full conviction while simultaneously failing to recognize when they should stay silent.\u003C\u002Fp>\n\u003Ch2>Key Findings\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>Human radiologists scored 988.7 out of 2,000 points\u003C\u002Fstrong>, the best AI model only \u003Cstrong>758 points\u003C\u002Fstrong>\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Claude Fable 5\u003C\u002Fstrong> (Anthropic) was most reliable on safe answers; \u003Cstrong>Gemini 3 Pro\u003C\u002Fstrong> (Google) had highest raw accuracy\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Meta&#39;s Muse Spark 1.1\u003C\u002Fstrong> best recognized when to hand cases to humans\u003C\u002Fli>\n\u003Cli>The benchmark rewards honesty: wrong answers with high confidence result in point deductions – mirroring real medical consequences\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>The Core Problem: Confidence Over Competence\u003C\u002Fh2>\n\u003Cp>The RadLE 2.0 test measures three dimensions: whether the model gets the diagnosis right, how confident it is in that answer, and whether it can admit when it&#39;s out of its depth. AI systems must rate answers on a confidence scale from 0 to 4 – and are explicitly allowed to say &quot;I don&#39;t know.&quot;\u003C\u002Fp>\n\u003Cp>Here&#39;s the problem: many don&#39;t. They prefer to guess with high confidence rather than admit uncertainty. In radiology, this isn&#39;t an academic question – a confidently wrong diagnosis can be fatal for patients.\u003C\u002Fp>\n\u003Cp>The scoring system punishes exactly this behavior: correct answers with high confidence earn full points. Wrong answers with high confidence lose corresponding points. Saying &quot;I don&#39;t know&quot; costs zero points but causes no harm.\u003C\u002Fp>\n\u003Ch2>Different Strategies, Different Risks\u003C\u002Fh2>\n\u003Cdiv class=\"tbl-scroll\">\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Model\u003C\u002Fth>\n\u003Cth>Strength\u003C\u002Fth>\n\u003Cth>Problem\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\u003Ctr>\n\u003Ctd>Claude Fable 5\u003C\u002Ftd>\n\u003Ctd>Reliable, safe answers\u003C\u002Ftd>\n\u003Ctd>Overall more conservative\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Gemini 3 Pro\u003C\u002Ftd>\n\u003Ctd>Highest raw accuracy\u003C\u002Ftd>\n\u003Ctd>Overconfidence amid uncertainty\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Muse Spark 1.1\u003C\u002Ftd>\n\u003Ctd>Best handover recognition\u003C\u002Ftd>\n\u003Ctd>Fewer hallucinations through frequent refusals\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Grok 4.5\u003C\u002Ftd>\n\u003Ctd>More knowledge\u003C\u002Ftd>\n\u003Ctd>Significantly more hallucinations, more convinced of wrong answers\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\u003C\u002Ftable>\u003C\u002Fdiv>\n\u003Cp>A particularly alarming example: Grok 4.5 hallucinates far more than its predecessor – not because it knows less, but because it&#39;s \u003Cem>more\u003C\u002Fem> confident in its false answers.\u003C\u002Fp>\n\u003Cp>Open-source and medical-specific models show the problem most starkly. They attempt to answer nearly every case but frequently get it wrong – usually with medium to high confidence. According to the research team, several of these models would have scored much better if they had stayed quiet more often.\u003C\u002Fp>\n\u003Ch2>What This Means for Practice\u003C\u002Fh2>\n\u003Cp>The study addresses a fundamental problem recently raised in highly cited research: as long as benchmarks only reward accuracy, AI models are trained to guess. In medicine, that&#39;s catastrophic.\u003C\u002Fp>\n\u003Cp>For German enterprises and clinics, this means: before AI systems can independently diagnose in radiology, they must learn to accept their limits. A model that&#39;s right 90 percent of the time but wrong 10 percent – and confident in both cases – isn&#39;t suitable for clinical use. The question isn&#39;t just &quot;How accurate is the AI?&quot; but &quot;Does the AI know when it&#39;s not accurate?&quot; Regulation and certification must take this seriously – otherwise AI systems become a risk rather than a tool.\u003C\u002Fp>\n\u003Ch2>Sources\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fthe-decoder.com\u002Fai-chatbots-reading-x-rays-can-be-dangerously-confident-even-when-theyre-wrong\u002F\">The Decoder\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>\u003Cem>Editorially owned by \u003Ca href=\"\u002Fen\u002Fautor\u002Fideal-syka\">Ideal Syka\u003C\u002Fa>. Sources and method: \u003Ca href=\"\u002Fen\u002Fredaktion\">Newsroom &amp; method\u003C\u002Fa>. Tips and corrections: \u003Ca href=\"mailto:ai@i6eal.de\">ai@i6eal.de\u003C\u002Fa>.\u003C\u002Fem>\u003C\u002Fp>\n",1784569356605]