Abstract
Abstract
Background: Large language models (LLMs) are increasingly explored as adjunct tools in graduate medical education, but their reliability for answering discipline-specific certification examinations remains incompletely characterized. In dentistry in China, the national dental examination serves as a milestone for stomatology trainees. Prior studies predominantly rely on single metric, limiting reproducibility and granularity.
Objective: To systematically evaluate the performance of five LLMs in answering the national standardized training examination for dental residents.
Method: A total of 1,318 questions covering the six dentistry disciplines were selected. Five LLMs (DeepSeek-V3, DeepSeek-R1, Qwen3-32B, Claude 3.5, and GPT-4o) were evaluated via a single-response protocol on the Dify-based Chatflow framework. Performance combined human verification with automated scoring using the Retrieval-Augmented Generation Systems (RAGAS) framework across seven dimensions. Tasks involving question classification, key-point categorization, and instructional requirement summarization were designed to enable a horizontal comparison of models' cognitive capabilities, assessed using inter-model consistency scores.
Results: DeepSeek-R1 achieved the highest correct choice rate (80.90%, p < 0.001) and answer correctness score (0.853, p < 0.001), significantly outperforming the others. All models exhibited limitations, including information insufficiency, redundancy, and hallucination-related issues (lowest 44.20% in DeepSeek-R1). In classification and summarization tasks, Claude 3.5 showed higher consistency scores (186/181) than Qwen3-32B (152/151).
Conclusion: Evaluated LLMs demonstrated good performance on dental examination questions and in executing tasks. DeepSeek-R1 consistently outperformed the others. While LLMs showed potential as supportive tools in dental education, further refinement and validation are needed before formal adoption in graduate medical educational and clinical training settings.