Short answer

Integrate LLMs into educational and professional tools to provide intelligent, context-aware support and explanation, especially in knowledge-intensive domains.

Field
User-Centred Design
Source
PLOS Digital Health (2023)
Method
Performance evaluation of an AI model against human-designed standardized tests.
Evidence
Strong effect

Large Language Models (LLMs) can perform at or near the passing threshold on complex medical licensing exams, suggesting their capability to understand and apply medical knowledge. This user-centred design research insight is drawn from a 2023 study published in PLOS Digital Health. Using Performance evaluation of an ai model against human-designed standardized tests., researchers explored how this design variable affects real-world outcomes. The key design takeaway: Integrate LLMs into educational and professional tools to provide intelligent, context-aware support and explanation, especially in knowledge-intensive domains.

Study
User-Centred DesignRecentStrong effect

Large Language Models (LLMs) achieve passing scores on medical licensing exams, indicating potential for AI-assisted medical education.

Large Language Models (LLMs) can perform at or near the passing threshold on complex medical licensing exams, suggesting their capability to understand and apply medical knowledge.

PLOS Digital Health · 2023

01

Key Findings

  • 01ChatGPT performed at or near the passing threshold for all three USMLE exams (Step 1, Step 2CK, and Step 3).
  • 02ChatGPT demonstrated a high level of concordance and insight in its explanations for the exam questions.
02

Application

Design takeaway

Integrate LLMs into educational and professional tools to provide intelligent, context-aware support and explanation, especially in knowledge-intensive domains.

How to apply

For a medical education platform, design a feature where students can ask an AI (powered by an LLM) to explain complex medical conditions or treatment protocols, and the AI provides detailed, insightful answers, potentially referencing relevant guidelines.

Project actions

  • 01Consider how AI could help students learn complex subjects by explaining things in different ways.
  • 02Think about how AI could be used to create practice questions or study guides for exams.
03

Method & Evidence

AimTo evaluate the performance of ChatGPT on the United States Medical Licensing Exam (USMLE) without specialized training.
MethodPerformance evaluation of an AI model against human-designed standardized tests.
ProcedureChatGPT was presented with questions from the USMLE Step 1, Step 2CK, and Step 3 exams. Its responses were then scored against the exam criteria.
ContextMedical education and AI capabilities.

Variables

IVType of AI model (ChatGPT)
DVPerformance on USMLE (passing threshold, concordance, insight)
CVThe specific USMLE exams (Step 1, 2CK, 3)
04

Strengths & Limitations

Strengths

  • +Demonstrates AI's capability on a highly complex, standardized test.
  • +Highlights the potential for AI in educational contexts without specialized training.

Limitations

This study doesn't tell us if AI can actually treat patients or make real-life medical decisions, just that it can answer exam questions.

Reliability & validity

The reliability is high as the USMLE is a standardized test. The validity for assessing 'medical knowledge' is strong, but its validity for 'clinical competence' is limited.

Think critically

If AI can pass medical exams, what are the implications for the future of professional education and the role of human experts? What are the risks of over-reliance on AI in critical fields?

05

Design Principles

"AI-Enhanced Knowledge Facilitation: Leverage AI to process, explain, and assess complex information, augmenting human learning and decision-making."

This demonstrates that AI can process and synthesize vast amounts of specialized information, mimicking human-level understanding in specific domains. This capability can transform how medical knowledge is accessed, learned, and applied, potentially reducing cognitive load for human learners and practitioners.

06

What This Means for Your Design

AI like ChatGPT can pass tough medical exams, meaning it's good at understanding and explaining medical stuff.

How to use in your project

  • 1.When designing information architecture for an educational platform, consider how an LLM could act as an intelligent search or Q&A interface, providing direct answers rather than just links to documents.
07

Add to My Project

08

Quick Cite

Paragraph starter

Research by Kung et al. (2023) demonstrated that large language models like ChatGPT can perform at or near passing thresholds on the USMLE, suggesting their potential for AI-assisted medical education.

09

Source

PLOS Digital Health

Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models

journal · 2023

View source

Questions About This Research

What does the research say about large language models (llms) achieve passing scores on medical licensing exams, indicating potential for ai-assisted medical education?
Integrate LLMs into educational and professional tools to provide intelligent, context-aware support and explanation, especially in knowledge-intensive domains. Evidence: PLOS Digital Health (2023).
Why does "Large Language Models (LLMs) achieve passing scores on medical licensing exams, indicating potential for AI-assisted medical education." matter for design?
This demonstrates that AI can process and synthesize vast amounts of specialized information, mimicking human-level understanding in specific domains. This capability can transform how medical knowledge is accessed, learned, and applied, potentially reducing cognitive load for human learners and practitioners.
How can designers apply this research?
Integrate LLMs into educational and professional tools to provide intelligent, context-aware support and explanation, especially in knowledge-intensive domains.
What were the main findings?
ChatGPT performed at or near the passing threshold for all three USMLE exams (Step 1, Step 2CK, and Step 3).. ChatGPT demonstrated a high level of concordance and insight in its explanations for the exam questions.
What research method was used?
Performance evaluation of an AI model against human-designed standardized tests..
How strong is the evidence?
Evidence strength is rated Strong effect, based on a 2023 journal from PLOS Digital Health.
What should I do differently in my next project?
For a medical education platform, design a feature where students can ask an AI (powered by an LLM) to explain complex medical conditions or treatment protocols, and the AI provides detailed, insightful answers, potentially referencing relevant guidelines.
What are the limitations?
The study only evaluated one specific LLM (ChatGPT) and its performance on a single type of exam (USMLE). It does not assess clinical judgment, ethical reasoning, or real-world patient interaction.