GPT-4’s Errors on Psychosomatic Medicine Exam Questions Through Bloom’s Taxonomy
In a 2023 preprint, GPT-4 answered 307 multiple-choice questions in psychosomatic medicine with 93% accuracy. Its errors were concentrated in Bloom’s “remember” and “understand” categories, and several resembled questions that medical students also find difficult.
Study overview
The study, “Assessing ChatGPT’s Mastery of Bloom’s Taxonomy using psychosomatic medicine exam questions,” by Anne Herrmann Werner and colleagues, evaluated GPT-4 on 307 multiple-choice questions from the psychosomatic medicine domain of medical examinations.
paragraphs
The authors used two prompting approaches: a detailed prompt and a short prompt. GPT-4 achieved 93% accuracy with the detailed prompt and 91% with the short prompt.
The paper is a MedRxiv preprint: https://doi.org/10.1101/2023.08.18.23294159
Bloom’s taxonomy and error patterns
Bloom’s taxonomy includes six levels: remember, understand, apply, analyze, evaluate, and create.
The questions GPT-4 missed were often questions that students also find difficult. In the analysis of reasoning errors, the largest groups of errors fell into the “remember” and “understand” categories, followed by “apply.” Errors classified as analyze, evaluate, or create were rare.
This suggests that GPT-4’s main difficulties involved accurately recalling diagnostic criteria, connecting concepts correctly, and reaching a diagnosis that fits those criteria.
Examples of incorrect answers
Remember: In a question about diagnostic criteria for somatization disorder, GPT-4 overlooked the requirement that symptoms must have been present for at least two years. The correct answer was undifferentiated somatoform disorder.
Understand: In a question about refeeding in a patient with anorexia nervosa, GPT-4 recognized in the third statement that a previously reduced basal metabolic rate rises during refeeding. However, it did not adequately incorporate that concept when choosing its final answer. The correct answer was that transient hypercholesterolemia occurs but does not require treatment.
Apply: GPT-4 correctly recalled and understood the diagnostic criteria for a depressive episode, but applied them too flexibly. In fact, at least two weeks must pass before the diagnosis is made.
Evaluate: In one example, GPT-4 appeared able to remember, understand, apply, and analyze the relevant information, but did not properly evaluate the consequences of not providing inpatient treatment.
Discussion
Overall, GPT-4 showed high performance, exceeding 90% accuracy. At the same time, its error pattern indicates that strong overall test performance does not eliminate problems with basic factual recall, conceptual integration, and strict use of diagnostic criteria.
I think it is important to analyze what kinds of errors language models make on medical questions. This matters for safety, for identifying directions for future model improvement, and for considering whether changes in prompting could reduce particular errors.
It would have been useful if the study had compared GPT-4 with at least one other model. I also wonder whether alternative prompting methods could have improved performance.
- Original tweet: https://twitter.com/mahler83/status/1693888223444136269