Hidden Flaws
Medical AI Evaluation · Multimodal Reasoning

Hidden Flaws Behind Expert-Level Accuracy of Multimodal GPT-4 Vision in Medicine

Qiao Jin, Fangyuan Chen, Yiliang Zhou, Ziyang Xu, Justin M. Cheung, …, Michael F. Chiang, Yifan Peng, Zhiyong Lu
npj Digit. Med. 2024
These slides were generated with the help of AI and may contain errors.
01Motivation

High accuracy can hide broken reasoning

  • GPT-4V scores at expert level on medical image quizzes.
  • But a multiple-choice score says nothing about how it reached the answer.
  • A model can be right for the wrong reasons — dangerous in medicine.
02Method

Grading the rationale, not just the answer

  • We evaluated GPT-4V on 207 New England Journal of Medicine Image Challenges.
  • Physicians graded each rationale for image comprehension, knowledge recall & reasoning.
Each GPT-4V response is split into image comprehension, knowledge recall & reasoning, and graded
Each GPT-4V response is split into image comprehension, knowledge recall & reasoning, and graded by nine specialty physicians against the ground truth. Fig. 1.
03Result

Right answers, wrong reasons

  • GPT-4V matches physicians on accuracy — 81.6% vs. 77.8%.
  • But even when its answer is correct, image comprehension is right only 72.8% of the time.
  • Knowledge recall is most reliable (91.1%); reasoning collapses when the answer is wrong (10.5%).
Share of GPT-4V rationales rated correct / partially / incorrect, by capability — split by dif
Share of GPT-4V rationales rated correct / partially / incorrect, by capability — split by difficulty and by whether its final answer was right. Fig. 2c.
04Recognition

Reported by NIH

  • Featured in an NIH News release — the National Library of Medicine highlighted the study’s findings on integrating AI into medical decision-making.
nih.gov News Releases · July 23, 2024.
nih.gov News Releases · July 23, 2024.
05Summary

Hidden Flaws at a glance

BackgroundGPT-4V reaches expert-level scores on medical image quizzes.
ProblemBut a multiple-choice score can’t reveal whether the reasoning is sound.
ApproachNine physicians graded GPT-4V rationales on 207 NEJM Image Challenges — image comprehension, knowledge recall & reasoning.
ResultsAccuracy matches physicians (81.6% vs. 77.8%), yet 35.5% of correct answers rest on flawed rationales — mostly image misreading.
ConclusionEvaluate the reasoning, not just the answer, before deploying multimodal AI in medicine.