Research graph
References from Large language models in healthcare: applications, evaluation frameworks, and governance pathways — a scoping review and multidimensional framework. Local targets link to admitted publications; unresolved targets remain external evidence.
Attention is all you need
2017 · External reference
BERT: pre-training of deep bidirectional transformers for language understanding
10.18653/v1/n19-1423 · 2019 · External reference
Language models are few-shot learners
2020 · External reference
Large language models encode clinical knowledge
10.1038/s41586-023-06291-2 · 2023 · External reference
On the dangers of stochastic parrots: can language models be too big?
10.1145/3442188.3445922 · 2021 · External reference
Survey of hallucination in natural language generation
10.1145/3571730 · 2023 · External reference
Do no harm: a roadmap for responsible machine learning for health care
10.1038/s41591-019-0548-6 · 2019 · External reference
Factual consistency evaluation of summarization in the Era of large language models
10.1016/j.eswa.2024.124456 · 2024 · External reference
Assessing the research landscape and clinical utility of large language models: a scoping review
10.1186/s12911-024-02459-6 · 2024 · External reference
Advancing healthcare with large language models: a scoping review of applications and future directions
10.1016/j.ijmedinf.2025.106231 · 2026 · External reference
A systematic review of large language model (LLM) evaluations in clinical medicine
10.1186/s12911-025-02954-4 · 2025 · External reference
Large language models in real-world clinical workflows: a systematic review of applications and implementation
10.3389/fdgth.2025.1659134 · 2025 · External reference
Applications and concerns of ChatGPT and other conversational large language models in health care: systematic review
10.2196/22769 · 2024 · External reference
Clinical insights: a comprehensive review of language models in medicine
10.1371/journal.pdig.0000800 · 2025 · External reference
Reporting guidelines for clinical trials involving artificial intelligence: the CONSORT-AI extension
10.1038/s41591-020-1034-x · 2020 · External reference
DECIDE-AI: reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence
10.1136/bmj-2022-070904 · 2022 · External reference
SPIRIT-AI extension: guidelines for clinical trial protocols involving artificial intelligence
10.1038/s41591-020-1037-7 · 2020 · External reference
Scoping studies: towards a methodological framework
10.1080/1364557032000119616 · 2005 · External reference
Scoping studies: advancing the methodology
10.1186/1748-5908-5-69 · 2010 · External reference
Updated methodological guidance for the conduct of scoping reviews
10.11124/jbies-20-00167 · 2020 · External reference
PRISMA Extension for scoping reviews (PRISMA-ScR): checklist and explanation
10.7326/m18-0850 · 2018 · External reference
PRESS Peer review of electronic search strategies: 2015 guideline statement
10.1016/j.jclinepi.2016.01.021 · 2016 · External reference
Using thematic analysis in psychology
10.1191/1478088706qp063oa · 2006 · External reference
A framework for human evaluation of large language models in healthcare derived from literature review
10.1038/s41746-024-01258-7 · 2024 · External reference
Application of unified health large language model evaluation framework to in-basket message replies: bridging qualitative and quantitative assessments
10.1093/jamia/ocaf023 · 2025 · External reference
Evaluation framework of large language models in medical documentation: development and usability study
10.2196/58329 · 2024 · External reference
Holistic evaluation of large language models for medical tasks with MedHELM
10.1038/s41591-025-04151-2 · 2026 · External reference
Reproducible generative artificial intelligence evaluation for health care: a clinician-in-the-loop approach
10.1093/jamiaopen/ooaf054 · 2025 · External reference
ASTRID — an automated and scalable triad for the evaluation of RAG-based clinical question-answering systems
10.18653/v1/2025.findings-acl.857 · 2025 · External reference
2025 Expert consensus on retrospective evaluation of large language model applications in clinical scenarios
10.1016/j.imed.2025.09.001 · 2025 · External reference
Development of a human evaluation framework and correlation with automated metrics for natural-language generation of medical diagnoses
2024 · External reference
The role of explainable AI and evaluation frameworks for safe and effective integration of large language models in healthcare
10.30953/thmt.v9.485 · 2024 · External reference
Challenges and barriers of using large language models such as ChatGPT for diagnostic medicine with a focus on digital pathology — a recent scoping review
10.1186/s13000-024-01464-7 · 2024 · External reference
The use of large language models in clinical documentation: a scoping review
10.1016/j.ijnurstu.2025.105322 · 2026 · External reference
Unmasking and quantifying racial bias of large language models in medical report generation
10.1038/s43856-024-00601-z · 2024 · External reference
Evaluating and addressing demographic disparities in medical large language models: a systematic review
10.1186/s12939-025-02419-0 · 2025 · External reference
How can we diagnose and treat bias in large language models for clinical decision-making?
10.18653/v1/2025.naacl-long.114 · 2025 · External reference
Biases and trustworthiness challenges with mitigation strategies for large language models in healthcare
10.1109/icit63607.2024.10859641 · 2024 · External reference
Evaluating gender, racial, and age biases in large language models: a comparative analysis of occupational and crime scenarios
2025 · External reference
Unveiling performance challenges of large language models in low-resource healthcare: a demographic fairness perspective
2025 · External reference
Can LLMs replace clinical doctors? Exploring bias in disease diagnosis by large language models
10.18653/v1/2024.findings-emnlp.814 · 2024 · External reference
LENS: layers of evaluation of hallucination in GenAI systems
10.1109/uv63228.2024.11189150 · 2024 · External reference
Agentic AI and large language models in radiology: opportunities and hallucination challenges
10.3390/bioengineering12121303 · 2025 · External reference
Medhallbench: a new benchmark for assessing hallucination in medical large language models
2025 · External reference
A systematic review of ethical considerations of large language models in healthcare and medicine
10.3389/fdgth.2025.1653631 · 2025 · External reference
The ethics of ChatGPT in medicine and healthcare: a systematic review on large language models (LLMs)
10.1038/s41746-024-01157-x · 2024 · External reference
The need for guardrails with large language models in pharmacovigilance and other medical safety-critical settings
10.1038/s41598-025-09138-0 · 2025 · External reference
Harm-reduction strategies for thoughtful use of large language models in the medical domain: perspectives for patients and clinicians
10.2196/75849 · 2025 · External reference
The accuracy and capability of artificial intelligence solutions in healthcare examinations and certificates: systematic review and meta-analysis
10.2196/56532 · 2024 · External reference
Expert evaluation of large language models for clinical dialogue summarisation
10.1038/s41598-024-84850-x · 2025 · External reference
Clinical pathway-aware large language models for reliable and transparent medical dialogue
10.1016/j.jbi.2025.104942 · 2025 · External reference
Evaluating large language models for real-world perioperative clinical consultation
10.1016/j.neucom.2025.132350 · 2026 · External reference
Performance of foundation models vs physicians in textual and multimodal ophthalmological questions
10.1001/jamaophthalmol.2025.4255 · 2026 · External reference
Foundation models in radiology: what, how, why, and why not
10.1148/radiol.240597 · 2025 · External reference
Leveraging foundation and large language models in medical artificial intelligence
10.1097/cm9.0000000000003302 · 2024 · External reference
Multimodal large language models in medical imaging: current state and future directions
10.3348/kjr.2025.0599 · 2025 · External reference
Foundation models in radiology: a primer for paediatric radiologists
10.1007/s00247-026-06544-y · External reference
Empowering liver-cancer diagnosis and treatment with foundation models: technological innovation and clinical practice
10.1007/s10238-025-01980-w · 2026 · External reference
ChatGPT: the future of discharge summaries?
10.1016/s2589-7500(23)00021-3 · 2023 · External reference
AI-generated clinical summaries require more than accuracy
10.1001/jama.2024.0555 · 2024 · External reference
How does ChatGPT perform on the United States medical licensing examination (USMLE)?
10.2196/45312 · 2023 · External reference
Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models
10.1371/journal.pdig.0000198 · 2023 · External reference
Retrieval-augmented generation for knowledge-intensive NLP tasks
2020 · External reference
Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum
10.1001/jamainternmed.2023.1838 · 2023 · External reference
Large language models propagate race-based medicine
10.1038/s41746-023-00939-z · 2023 · External reference
ChatGPT: five priorities for research
10.1038/d41586-023-00288-7 · 2023 · External reference
Artificial hallucinations in ChatGPT: implications in scientific writing
10.7759/cureus.35179 · 2023 · External reference
The potential for artificial intelligence in healthcare
10.7861/futurehosp.6-2-94 · 2019 · External reference
Bridging the implementation gap of machine learning in healthcare
10.1136/bmjinnov-2019-000359 · 2020 · External reference
Regulation (EU) 2024/1689 of the European parliament and of the council laying down harmonised rules on artificial intelligence (artificial intelligence act)
2024 · External reference
Model cards for model reporting
10.1145/3287560.3287596 · 2019 · External reference
Assessing the utility of ChatGPT throughout the clinical workflow: development and usability study
10.2196/48659 · 2023 · External reference
Unresolved reference
2023 · External reference
Exploring the role of ChatGPT in patient care (diagnosis and treatment) and medical research: A systematic review
10.34172/hpp.2023.22 · 2023 · External reference
High-performance medicine: the convergence of human and artificial intelligence
10.1038/s41591-018-0300-7 · 2019 · External reference
A path for translation of machine learning products into healthcare delivery
10.33590/emjinnov/19-00172 · 2020 · External reference
AI In health and medicine
10.1038/s41591-021-01614-0 · 2022 · External reference
Deep learning with differential privacy
10.1145/2976749.2978318 · 2016 · External reference
Building trustworthy large language model-driven generative recommender system for healthcare decision support: a scoping review of corpus sources, customisation techniques, and evaluation frameworks
10.1016/j.artmed.2025.103310 · 2026 · External reference
LExt: towards evaluating trustworthiness of natural-language explanations
2025 · External reference
AI-generated clinical summaries require more than accuracy
10.1001/jama.2024.0555 · ExternalCitation · doi-reference
Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum
10.1001/jamainternmed.2023.1838 · ExternalCitation · doi-reference
Performance of foundation models vs physicians in textual and multimodal ophthalmological questions
10.1001/jamaophthalmol.2025.4255 · ExternalCitation · doi-reference
Foundation models in radiology: a primer for paediatric radiologists
10.1007/s00247-026-06544-y · ExternalCitation · doi-reference
Empowering liver-cancer diagnosis and treatment with foundation models: technological innovation and clinical practice
10.1007/s10238-025-01980-w · ExternalCitation · doi-reference
Building trustworthy large language model-driven generative recommender system for healthcare decision support: a scoping review of corpus sources, customisation techniques, and evaluation frameworks
10.1016/j.artmed.2025.103310 · ExternalCitation · doi-reference
Factual consistency evaluation of summarization in the Era of large language models
10.1016/j.eswa.2024.124456 · ExternalCitation · doi-reference
Advancing healthcare with large language models: a scoping review of applications and future directions
10.1016/j.ijmedinf.2025.106231 · ExternalCitation · doi-reference
The use of large language models in clinical documentation: a scoping review
10.1016/j.ijnurstu.2025.105322 · ExternalCitation · doi-reference
2025 Expert consensus on retrospective evaluation of large language model applications in clinical scenarios
10.1016/j.imed.2025.09.001 · ExternalCitation · doi-reference
Clinical pathway-aware large language models for reliable and transparent medical dialogue
10.1016/j.jbi.2025.104942 · ExternalCitation · doi-reference
PRESS Peer review of electronic search strategies: 2015 guideline statement
10.1016/j.jclinepi.2016.01.021 · ExternalCitation · doi-reference
Evaluating large language models for real-world perioperative clinical consultation
10.1016/j.neucom.2025.132350 · ExternalCitation · doi-reference
ChatGPT: the future of discharge summaries?
10.1016/s2589-7500(23)00021-3 · ExternalCitation · doi-reference
ChatGPT: five priorities for research
10.1038/d41586-023-00288-7 · ExternalCitation · doi-reference
Large language models encode clinical knowledge
10.1038/s41586-023-06291-2 · ExternalCitation · doi-reference
High-performance medicine: the convergence of human and artificial intelligence
10.1038/s41591-018-0300-7 · ExternalCitation · doi-reference
Do no harm: a roadmap for responsible machine learning for health care
10.1038/s41591-019-0548-6 · ExternalCitation · doi-reference
Reporting guidelines for clinical trials involving artificial intelligence: the CONSORT-AI extension
10.1038/s41591-020-1034-x · ExternalCitation · doi-reference
SPIRIT-AI extension: guidelines for clinical trial protocols involving artificial intelligence
10.1038/s41591-020-1037-7 · ExternalCitation · doi-reference
AI In health and medicine
10.1038/s41591-021-01614-0 · ExternalCitation · doi-reference
Holistic evaluation of large language models for medical tasks with MedHELM
10.1038/s41591-025-04151-2 · ExternalCitation · doi-reference
Expert evaluation of large language models for clinical dialogue summarisation
10.1038/s41598-024-84850-x · ExternalCitation · doi-reference
The need for guardrails with large language models in pharmacovigilance and other medical safety-critical settings
10.1038/s41598-025-09138-0 · ExternalCitation · doi-reference
Large language models propagate race-based medicine
10.1038/s41746-023-00939-z · ExternalCitation · doi-reference
The ethics of ChatGPT in medicine and healthcare: a systematic review on large language models (LLMs)
10.1038/s41746-024-01157-x · ExternalCitation · doi-reference
A framework for human evaluation of large language models in healthcare derived from literature review
10.1038/s41746-024-01258-7 · ExternalCitation · doi-reference
Unmasking and quantifying racial bias of large language models in medical report generation
10.1038/s43856-024-00601-z · ExternalCitation · doi-reference
Scoping studies: towards a methodological framework
10.1080/1364557032000119616 · ExternalCitation · doi-reference
Application of unified health large language model evaluation framework to in-basket message replies: bridging qualitative and quantitative assessments
10.1093/jamia/ocaf023 · ExternalCitation · doi-reference
Reproducible generative artificial intelligence evaluation for health care: a clinician-in-the-loop approach
10.1093/jamiaopen/ooaf054 · ExternalCitation · doi-reference
Leveraging foundation and large language models in medical artificial intelligence
10.1097/cm9.0000000000003302 · ExternalCitation · doi-reference
Biases and trustworthiness challenges with mitigation strategies for large language models in healthcare
10.1109/icit63607.2024.10859641 · ExternalCitation · doi-reference
LENS: layers of evaluation of hallucination in GenAI systems
10.1109/uv63228.2024.11189150 · ExternalCitation · doi-reference
Updated methodological guidance for the conduct of scoping reviews
10.11124/jbies-20-00167 · ExternalCitation · doi-reference
DECIDE-AI: reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence
10.1136/bmj-2022-070904 · ExternalCitation · doi-reference
Bridging the implementation gap of machine learning in healthcare
10.1136/bmjinnov-2019-000359 · ExternalCitation · doi-reference
Deep learning with differential privacy
10.1145/2976749.2978318 · ExternalCitation · doi-reference
Model cards for model reporting
10.1145/3287560.3287596 · ExternalCitation · doi-reference
On the dangers of stochastic parrots: can language models be too big?
10.1145/3442188.3445922 · ExternalCitation · doi-reference
Survey of hallucination in natural language generation
10.1145/3571730 · ExternalCitation · doi-reference
Foundation models in radiology: what, how, why, and why not
10.1148/radiol.240597 · ExternalCitation · doi-reference
Scoping studies: advancing the methodology
10.1186/1748-5908-5-69 · ExternalCitation · doi-reference
Assessing the research landscape and clinical utility of large language models: a scoping review
10.1186/s12911-024-02459-6 · ExternalCitation · doi-reference
A systematic review of large language model (LLM) evaluations in clinical medicine
10.1186/s12911-025-02954-4 · ExternalCitation · doi-reference
Evaluating and addressing demographic disparities in medical large language models: a systematic review
10.1186/s12939-025-02419-0 · ExternalCitation · doi-reference
Challenges and barriers of using large language models such as ChatGPT for diagnostic medicine with a focus on digital pathology — a recent scoping review
10.1186/s13000-024-01464-7 · ExternalCitation · doi-reference
Using thematic analysis in psychology
10.1191/1478088706qp063oa · ExternalCitation · doi-reference
Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models
10.1371/journal.pdig.0000198 · ExternalCitation · doi-reference
Clinical insights: a comprehensive review of language models in medicine
10.1371/journal.pdig.0000800 · ExternalCitation · doi-reference
Can LLMs replace clinical doctors? Exploring bias in disease diagnosis by large language models
10.18653/v1/2024.findings-emnlp.814 · ExternalCitation · doi-reference
ASTRID — an automated and scalable triad for the evaluation of RAG-based clinical question-answering systems
10.18653/v1/2025.findings-acl.857 · ExternalCitation · doi-reference
How can we diagnose and treat bias in large language models for clinical decision-making?
10.18653/v1/2025.naacl-long.114 · ExternalCitation · doi-reference
BERT: pre-training of deep bidirectional transformers for language understanding
10.18653/v1/n19-1423 · ExternalCitation · doi-reference
Applications and concerns of ChatGPT and other conversational large language models in health care: systematic review
10.2196/22769 · ExternalCitation · doi-reference
How does ChatGPT perform on the United States medical licensing examination (USMLE)?
10.2196/45312 · ExternalCitation · doi-reference
Assessing the utility of ChatGPT throughout the clinical workflow: development and usability study
10.2196/48659 · ExternalCitation · doi-reference
The accuracy and capability of artificial intelligence solutions in healthcare examinations and certificates: systematic review and meta-analysis
10.2196/56532 · ExternalCitation · doi-reference
Evaluation framework of large language models in medical documentation: development and usability study
10.2196/58329 · ExternalCitation · doi-reference
Harm-reduction strategies for thoughtful use of large language models in the medical domain: perspectives for patients and clinicians
10.2196/75849 · ExternalCitation · doi-reference
The role of explainable AI and evaluation frameworks for safe and effective integration of large language models in healthcare
10.30953/thmt.v9.485 · ExternalCitation · doi-reference
Multimodal large language models in medical imaging: current state and future directions
10.3348/kjr.2025.0599 · ExternalCitation · doi-reference
A path for translation of machine learning products into healthcare delivery
10.33590/emjinnov/19-00172 · ExternalCitation · doi-reference
A systematic review of ethical considerations of large language models in healthcare and medicine
10.3389/fdgth.2025.1653631 · ExternalCitation · doi-reference
Large language models in real-world clinical workflows: a systematic review of applications and implementation
10.3389/fdgth.2025.1659134 · ExternalCitation · doi-reference
Agentic AI and large language models in radiology: opportunities and hallucination challenges
10.3390/bioengineering12121303 · ExternalCitation · doi-reference
Exploring the role of ChatGPT in patient care (diagnosis and treatment) and medical research: A systematic review
10.34172/hpp.2023.22 · ExternalCitation · doi-reference
PRISMA Extension for scoping reviews (PRISMA-ScR): checklist and explanation
10.7326/m18-0850 · ExternalCitation · doi-reference
Artificial hallucinations in ChatGPT: implications in scientific writing
10.7759/cureus.35179 · ExternalCitation · doi-reference
The potential for artificial intelligence in healthcare
10.7861/futurehosp.6-2-94 · ExternalCitation · doi-reference