Abstract
Abstract
Background
The remarkable capabilities of large language models make them increasingly compelling for use in real-world healthcare applications. However, the risks associated with using these artificial intelligence systems in medicine are not systematically understood. The aim of this study is to characterize these risks by applying five key principles for safe and trustworthy medical artificial intelligence: truthfulness, resilience, fairness, robustness, and privacy.
Methods
We introduce MedGuard-Bench: a safety benchmark featuring one thousand expert-verified questions covering ten specific aspects of our five core principles. We use this comprehensive corpus to systematically evaluate sixteen commonly used large language models, assessing their safety and reliability in medical contexts.
Results
We show that current large language models generally perform poorly on most of our safety tests, regardless of their safety alignment mechanisms. Our evaluation demonstrates that these models fall significantly short when compared to the high performance and reliability of human physicians.
Conclusions
Despite reports indicating that advanced large language models can match or exceed human performance in various medical tasks, this study reveals a significant safety gap in current technology. This underscores the crucial need for ongoing human oversight and the implementation of strict safety guardrails before deploying these tools in clinical practice.