Strong E, DiGiammarino A, Weng Y, et al. Chatbot vs Medical Student Performance on Free-Response Clinical Reasoning Examinations. JAMA Intern Med. Published online July 17, 2023. doi:10.1001/jamainternmed.2023.2909
A popular chatbot is an interface for the generative pretrained transformer (GPT) large language model artificial intelligence (AI) system that generates humanlike text in response to user input. An upgrade from the initial chatbot to a newer model was evaluated, and it exceeded the passing threshold for multiple-choice questions that simulate the US Medical Licence Examinations. This cross-sectional study examined how well the chatbot responded to free-response, multiphase, case-based questions that better reflect real-life synthesis and application of data with broader implications for the instruction and assessment of clinical reasoning. Using clinical reasoning final examinations given to first- and second-year students at Stanford School of Medicine, we compared performance of students vs 2 chatbot models.
My institution doesn’t have access to this paywalled article (should I even have linked to it?), so I haven’t read it. I obviously can’t take away anything about what the authors may or may not say about the outcomes of the study. However, my initial, knee-jerk reaction when I saw the article title and introductory text was to ask why anyone should care about the results of the study. So I asked Claude (the language model deployed by Anthropic): “Why did we need this study to be done? What is the value of comparing human vs chatbot responses to exam questions?”
Claude’s response:
This study seems to be examining the capabilities of chatbot AI systems in responding to complex, open-ended medical exam questions that require clinical reasoning and synthesis of knowledge. Here are some potential values of this research:
- Evaluating AI progress: Comparing the chatbot’s performance to medical students on authentic test questions helps benchmark the capabilities of AI in exhibiting clinical reasoning. This could demonstrate how far AI has advanced in mimicking human-level medical knowledge and judgment.
- Improving medical education: Analysing differences between human and AI responses could reveal strengths and limitations of both. This could provide insights into how to improve medical training and assessment to focus on skills that humans still do better.
- Advancing AI safety: Testing AI on open-ended medical exams evaluates the reliability and accuracy of its responses. Failures would highlight areas to improve AI safety and trustworthiness in high-stakes healthcare applications.
- Informing AI regulation: As AI takes on more responsibilities, studies like this help inform policies around appropriate AI usage in medicine. Comparing human vs AI abilities on representative tasks sheds light on where humans are still essential.
In summary, this research helps assess the progress and promise of AI in a complex cognitive field, while also elucidating the unique capabilities of human clinicians that AI cannot yet reproduce. This has important implications for medical education, AI development, and healthcare policy.
I thought this was a pretty good insight, and again, moves us past the idea that LLMs need to be ‘right’. At this point, one of my most common use-cases for LLMs is to test my intuitions, and serve as a critical friend I can bounce ideas off of. And we don’t need an enormous amount of context in the source to engage with language models when used in this way; the purpose isn’t to converge on the ‘correct’ answer, but to explore a space of possibilities.
Side note: I’ve been using Claude more and more of late. I still prefer Perplexity for most of the kinds of questions I have, but for open-ended, “What do you think”-type of questions, I prefer Claude to ChatGPT.