Scholarly article · 2026
Enhancing history-taking education through GPT-4-based virtual patients and automated assessment: a study of medical student perceptions.
Byun J, Kim H, Lim J, Choi J, Ahn S
원문 정보 보기한국어 요약
이 연구는 GPT-4 API를 활용해 가상 환자(VP) 7개 임상 시나리오와 가상 평가자(VA)를 제공하는 병력청취 교육 도구를 개발하고, 기존 학습 방법과 비교해 의과대학생의 인식을 평가했다. 6일간의 연구에 참여한 학생 21명은 LLM 기반 도구가 사용성, 자기효능감, 피드백의 질 측면에서 기존 방법보다 우수하다고 평가했다.
핵심 내용
- 도구는 가상 환자와 상호작용하고, 가상 평가자로부터 즉각적인 성찰적 대화 피드백과 종합적인 서면 피드백을 받을 수 있도록 설계되었다.
- 연구에는 의과대학 1·2학년 학생이 참여했으며, 사전·사후 설문은 5점 리커트 척도로 사용성, 자기효능감, 피드백의 질을 평가했다.
- 연습 중 편안함은 LLM 기반 도구에서 기존 방법보다 높게 평가되었다(평균 4.57 대 2.95, p=0.0002). 특히 낯선 사례를 다룰 수 있다는 자신감을 포함한 8개 자기효능감 항목 중 6개에서 유의한 향상이 나타났다(4.00 대 2.90, p=0.0002).
- 권장 독자
- 의학교육 연구자, 임상교육자, 의료 AI 기반 학습도구를 평가하는 연구자 및 의과대학생
- 범위와 한계
- 표본은 21명이었고 연구 기간은 6일이었다. 결과는 주로 학생의 주관적 인식에 기반하며, 객관적 수행 지표로 학습 효과를 검증하지 않았으므로 결과의 일반화와 실제 병력청취 수행 향상은 제한적으로 해석해야 한다.
- 연구 맥락
- 의학교육에서 대규모 언어모델, 가상 환자, 자동화된 피드백을 활용해 병력청취 연습과 평가를 지원하는 교육·공중보건 연구 주제와 연결된다.
English summary
This study developed a GPT-4 API-based history-taking education tool featuring seven virtual-patient (VP) clinical scenarios and a virtual assessor (VA), and evaluated medical students’ perceptions compared with conventional learning methods. In the 6-day study, 21 students rated the LLM-based tool more favorably than conventional methods in usability, self-efficacy, and feedback quality.
Key points
- The tool enabled interaction with virtual patients and provided immediate reflective dialogue as well as comprehensive written feedback from the virtual assessor.
- First- and second-year medical students completed pre- and post-participation surveys using 5-point Likert scales to assess usability, self-efficacy, and feedback quality.
- Students reported greater comfort during practice with the LLM-based tool than with conventional methods (mean 4.57 vs. 2.95, p=0.0002). Six of eight self-efficacy measures improved significantly, including confidence in handling unfamiliar cases (4.00 vs. 2.90, p=0.0002).
- Audience
- Medical education researchers, clinical educators, investigators evaluating medical AI learning tools, and medical students
- Scope and limitations
- The study included 21 students and lasted 6 days. Its findings were based primarily on subjective perceptions, without validation using objective performance measures; therefore, generalizability and effects on actual history-taking performance should be interpreted cautiously.
- Research context
- The publication is thematically connected to education and public health through its focus on using large language models, virtual patients, and automated feedback to support history-taking practice and assessment in medical education.
AI-assisted summary based on the cited primary source. Model gpt-5.6-luna, prompt academic-hub-v1, generated 2026-08-29T02:43:42+00:00.