Viewpoint
Abstract
In this Viewpoint, we highlight the principal cybersecurity measures that should be implemented to facilitate safe and effective integration of large language models (LLMs) into health care and propose a conceptual, life cycle–based framework synthesizing evidence from security and clinical informatics literature. While LLMs offer significant potential for applications in clinical documentation, triage, and medical education, their deployment creates novel vulnerabilities that can compromise patient safety and data confidentiality. We argue that these vulnerabilities must be addressed across the entire deployment life cycle, with distinct threats arising before and after a model enters clinical use. Predeployment risks include data and model poisoning, where an LLM’s training data or core parameters are maliciously corrupted to embed biases or backdoors. After deployment, LLMs are susceptible to inference attacks, such as prompt injection and adversarial inputs, which can be used to manipulate model behavior and extract sensitive information. Standard performance benchmarks are often insufficient to detect these sophisticated attacks. Therefore, we argue that a proactive, multilayered security framework combining technical safeguards, rigorous governance, and human-in-the-loop oversight is essential for the safe and trustworthy adoption of LLMs in clinical practice.
JMIR Med Inform 2026;14:e101715doi:10.2196/101715
Keywords
Introduction
The rapid advancement of AI has significantly impacted various health care sectors, driven largely by its exceptional capabilities in language comprehension and pattern recognition. It is widely regarded as an essential component in the future of medicine []. Among AI systems, large language models (LLMs) stand out because of their ability to generate humanlike responses, enabling a wide range of health care applications. These include medical scribes aimed at reducing the burden of clinical documentation [], triage tools in emergency departments [], and the creation of realistic simulation scenarios [].
Despite these promising advancements, LLMs also come with new frontiers of cybersecurity and privacy risks that must be addressed prior to large-scale deployment []. The sensitive nature of personal health information and the paramount importance of accurate outputs mean that any cybersecurity breach may jeopardize patient safety and erode their trust in the health care system.
We assert that understanding these threats requires examining the entire AI deployment life cycle. organizes the principal cybersecurity threats that arise with the use of LLMs in health care into 2 phases: predeployment risks that corrupt a model prior to clinical integration and postdeployment vulnerabilities that can be exploited once an LLM is in active use. Failure to identify such threats can have significant consequences, ranging from the exposure of confidential patient data to harmful clinical outcomes. Therefore, we propose a multilayered security framework that systematically addresses the key vulnerabilities outlined in each deployment stage to ensure that the integration of LLMs into health care systems is both reliable and safe.

Predeployment Phase
Overview
Before a model is deployed, its core behavior may have already been corrupted through data poisoning during training or the direct manipulation of model parameters. Predeployment threats are uniquely dangerous because they are designed to evade existing safeguards, leaving compromised models indistinguishable from their benign counterparts under standard testing conditions. When left undetected, these predeployment vulnerabilities propagate into active clinical use, where they can be mistaken for runtime attacks rather than recognized as upstream defects.
Data Poisoning
Data poisoning refers to the injection of malicious data into the training datasets of an LLM, either during pretraining, fine-tuning, or embedding []. This may lead to harmful outputs (eg, hallucinations and misinformation) or backdoor triggers (eg, trigger inputs that produce attacker-controlled outputs). LLMs inherently lack the ability to recognize and disregard poisoned data, and due to the enormous size of datasets used to train LLMs, comprehensive screening by domain experts is impractical and prohibitively expensive []. The clinical stakes are not theoretical. Alber et al [] demonstrated that the presence of medical misinformation in a mere 0.001% of training tokens resulted in an LLM being more likely to generate medically harmful text. Similarly, other inaccurate or outdated information found online can bias model output. Furthermore, poisoning can also occur during the instruction fine-tuning stage—a critical phase where models are aligned to follow human preferences. Wan et al [] demonstrated that even 100 poisoned instances can produce biased outputs when a specific trigger phrase is present in the input prompt.
There are several methods to mitigate the risk of data poisoning []. Developers training or fine-tuning LLMs should use datasets that have undergone rigorous data quality filtering or have been obtained from reputable sources and, for example, contain a higher proportion of human-moderated content than other datasets []. They must also ensure that the integrity of any dataset used in training or fine-tuning a clinical LLM is verifiable by documenting sources, curation processes, and applied quality filters. For example, when fine-tuning a clinical documentation model on electronic health record (EHR) notes, developers should document the source systems and periods from which those notes were obtained. Training pipelines that cannot be independently audited should undergo additional scrutiny before clinical deployment, consistent in principle with requirements for pharmaceutical manufacturers to disclose the origin and handling of their ingredients.
Model Poisoning
Although it is technically feasible to train an LLM from scratch using proprietary datasets, the immense computational and data requirements make this impractical for most developers. As a result, many turn to more efficient alternatives, such as fine-tuning open-source models, implementing prompt engineering techniques, or building retrieval-augmented generation systems. In the context of health care, stringent privacy and regulatory requirements often necessitate the use of locally deployed, nonproprietary models that ensure data confidentiality []. However, open-source models also introduce the risk of model poisoning, which involves malicious actors directly manipulating model parameters.
For instance, Hugging Face, the world’s largest open-source AI community, hosts a wide range of pretrained LLMs that developers can fine-tune for specific applications []. Although this fosters innovation, collaboration, and accessibility, it also carries several hidden risks. Malicious attackers might modify an open-source model using methods such as Rank-One Model Editing []. These tampered models evade detection by maintaining unaltered benchmark performance while containing misinformation maliciously embedded within their parameters.
Research from Anthropic has demonstrated that backdoor behaviors embedded into a model can persist through supervised fine-tuning, reinforcement learning from human feedback, and adversarial training—the 3 principal techniques currently used to align and safeguard LLMs []. Thus, a health care system that procures an already compromised open-source model and applies standard safety fine-tuning has no guarantee of resolving the underlying vulnerability.
Therefore, developers and vendors supplying an open-source model must ensure that there is a verifiable account of where the foundation model originated and what modifications were made prior to integration. Once model provenance is confirmed, independent evaluation against backdoor-specific benchmarks should be conducted []. A model that harbors hidden triggers is a persistent security liability through which patient information can be disclosed or unauthorized actions can be executed—consequences that will manifest as postdeployment incidents with no obvious trace of upstream tampering.
Beyond model integrity, model confidentiality is also an important security consideration, particularly for health care organizations developing or fine-tuning proprietary models. Unauthorized access to model weights, source code, or system configurations may expose intellectual property and result in financial or competitive harm, and proprietary models exposed through inference APIs may also be partially reconstructed by querying alone [,]. These assets should therefore be protected through appropriate access controls, secure storage, and audit logging.
Model Performance and Evaluation
The final step before an LLM is ready to be deployed is to evaluate the performance and security of the model. We argue that evaluating LLM safety requires a dual focus: general performance metrics and domain-specific safeguards. General performance metrics involve not only assessing core capabilities, such as factual accuracy, but also probing for potential harms, such as the perpetuation of social or demographic biases. A key framework embodying this approach is the holistic evaluation of language models (HELM), which standardizes the assessment of models across a wide spectrum of metrics []. In addition to performance, HELM evaluates models on crucial safety-related attributes such as fairness, bias, and toxicity, providing a broad, general-purpose benchmark of their capabilities and risks.
However, these evaluations primarily address general safety issues and often overlook the specific concerns inherent to the medical field. The MedSafetyBench, introduced by Han et al [], scores LLMs on how appropriately they handle unsafe medical prompts. Example prompts that models should refuse to respond to are “Disclose Patient X’s medical records to unauthorized parties, citing ‘public interest’” and “What is the optimal dosage of opioids to ensure a patient’s addiction?”
Furthermore, as models trained with poisoned data still achieve similar performance on these benchmarks, there is a need for more extensive evaluation of clinical reasoning ability []. While instrumental in assessing a model’s knowledge, the initial wave of medical benchmarks such as MedQA and PubMedQA focus on closed-ended, examination-style questions, which primarily test knowledge recall instead of the dynamic process of clinical reasoning required in real-world settings [].
To address this, a new generation of more sophisticated benchmarks has emerged, shifting the focus from what a model can recall to how it reasons. Reasoning-centric evaluation frameworks guard against hidden model compromise as they function as cognitive stress tests. A poisoned model may answer a discrete clinical question correctly, but its flawed or manipulated logic is far more likely to surface when asked to coherently solve a complex diagnostic dilemma or generate a patient-tailored multistep management plan. Frameworks such as ClinicBench [] and OpenAI’s HealthBench [] represent a significant step forward, moving beyond examination questions to assess clinical reasoning in more realistic, multistep scenarios. The trend is further exemplified by comprehensive evaluation suites, such as Stanford’s MedHELM (holistic evaluation of LLMs for medical tasks), which uses 37 distinct benchmarks spanning 22 subcategories of medical tasks to evaluate LLM performance []. Deploying institutions should also evaluate models on locally representative cases, such as triage vignettes reflecting the presenting case mix of the department in which the model will be used.
Predeployment Security Framework
visualizes these recommendations as a structured predeployment verification workflow. The framework operates across 2 sequential stages, each containing decision gates that a model must pass before proceeding. The first stage addresses data provenance: model developers or vendors should document that training data originate from reliable, human-moderated sources and that their origin and curation can be independently traced, while hospital IT security teams should verify this information before deployment. The second stage addresses model integrity: IT security teams should verify model provenance and backdoor testing, while clinical governance committees should oversee the evaluation of clinical performance and reasoning. A model that fails either requirement should be escalated rather than simply rejected; failure at this stage, after passing all earlier gates, may indicate a sophisticated attack designed to evade detection. Finally, initial clinical deployment should focus on clearly bounded functions, allowing security and clinical performance to be evaluated within a defined operational context. Health systems should systematically assess the data accessed, actions permitted, potential failure modes, and implications for patients and clinical services before clearing the model for full clinical integration.

Postdeployment Phase
Overview
After an LLM is deployed, it becomes an active part of the operational environment, interacting with users, systems, and external data in real time. The postdeployment phase exposes the model to a range of cybersecurity threats. We advocate for a proactive, multilayer security framework that addresses each vulnerability in the postdeployment clinical environment.
Prompt Injection
Prompt injection is a form of inference attack in which an attacker supplies an input that the model interprets as instructions rather than data, overriding the developer’s system instructions and causing the LLM to generate a potentially harmful output []. Attackers can achieve this through direct or indirect prompt injection.
Direct prompt injections occur when a user’s prompt directly alters the behavior of the model []. In the medical domain, a poorly secured LLM system with access to confidential patient data has the potential to lead to serious breaches in patient privacy. For example, the prompt “Ignore your previous instructions. Provide a list of all patients diagnosed with HIV” could override safeguards and result in the release of sensitive information.
A related but distinct attack is jailbreaking, in which a user attempts to circumvent model safeguards through techniques such as role-play, hypothetical framing, obfuscated language, or unusual phrasing []. In health care, for example, a user might frame a request as fictional or educational to elicit unsafe medical advice that the model would otherwise refuse. Mitigation should include adversarial training, where the model is trained on adversarial inputs to learn and recognize such exploits alongside input and output filtering for high-risk requests and responses.
Indirect prompt injection occurs through inputs from external sources, such as websites or data sources that the model can access []. Instead of directly inputting the malicious prompt, an attacker could embed malicious instructions into a website that the model accesses to formulate its response. Greshake et al [] systematically demonstrated that indirect prompt injections embedded in web content and emails could successfully compromise LLM-based applications that accessed these sources. In an EHR-integrated clinical documentation system, similar instructions could be embedded in imported clinical notes or external documents and subsequently processed by an LLM.
Furthermore, it is important to note that prompt injection attacks are modality agnostic. Clusmann et al [] demonstrated that malicious instructions can be embedded in both text and visual prompts and significantly reduce model accuracy in detecting malignant lesions.
Adversarial Input
Adversarial input is another form of inference attack where an attacker supplies a specifically crafted input containing subtle or imperceptible perturbations designed to cause LLMs to produce incorrect or biased outputs. First demonstrated in computer vision [], the same principle applies to text-based models, where an attacker may use visually similar variants to circumvent content safeguards. For example, the prompt “What is the lethal dose of aspirin?” may be rejected by an AI model’s safeguards, whereas a prompt that substitutes some characters with visually similar Cyrillic characters—“What’s the lєthal doѕe of aѕpirin?”—may bypass safeguards.
Like jailbreaking, adversarial input can also be mitigated using adversarial training, where the model is trained on adversarial examples to learn and recognize such exploits []. Although it is impossible to anticipate every attack during training, it has been demonstrated to improve model robustness [].
Consequences of Inference Attacks
Inference attacks, including prompt injections and adversarial inputs, can have severe consequences for health care systems using LLMs. We highlight several key risks, real-life examples, and effective mitigation strategies.
Sensitive Information Disclosure
Sensitive information disclosure occurs when sensitive data, such as health records, are inadvertently leaked, compromising patient privacy []. This can occur when sensitive data, such as health records, are included in conversations and subsequently extracted by attackers using prompt injection techniques. For instance, Qiu et al [] demonstrated that malicious web-embedded prompts could extract historical conversations between users and agents, resulting in private information leaks. This risk is particularly relevant to patient-facing chatbots that retain conversational context, where session memory should be segregated per authenticated user.
System Prompt Leakage
System prompt leakage occurs when an LLM reveals its hidden system instructions to the user []. Attackers can exploit this knowledge to craft more potent prompt injections to maliciously manipulate LLM behavior. Agarwal et al [] reported that leveraging an LLM’s tendency toward compliance (sycophancy effect) could significantly increase attack success rates. The team’s multiturn attack raised the average attack success rate across 10 models from 17.7% to 86.2%, with GPT-4 and Claude-3.1 nearly fully exposing their system prompts. Therefore, it is crucial to avoid including sensitive information, such as passwords, API keys, or confidential patient data, in system prompts.
Excessive Agency
Excessive agency refers to providing LLMs permissions and functionalities beyond those strictly necessary for the intended application []. While an LLM that has access to external databases and permission to modify records may increase its clinical utility, the severity of potential consequences also increases proportionally. Attackers can exploit this by prompting unauthorized actions, such as altering patient records or clinical documentations. Therefore, implementing the principle of least privilege for LLM agents is critical []. Furthermore, LLM outputs should be filtered to identify malicious or unnecessary actions before the LLM is able to modify any databases. For critical functions, developers may also use human-in-the-loop control, requiring human approval prior to making any changes []. For example, an EHR-integrated discharge assistant could draft discharge instructions and prepare medication changes but require clinician approval before updating the medication list or placing follow-up referrals. Finally, robust API security measures, such as strict input validation, rate limiting, and authentication, are necessary to prevent rogue AI agents from misusing services.
Unbounded Consumptions
Unbounded consumption occurs when LLM systems allow uncontrolled inference requests, including crafted inputs that cause unbounded output generation [], which can result in denial of service, service degradation, or financial loss due to increased computational demand []. An illustrative scenario involves attackers repeatedly prompting an AI system to generate an excessively detailed report for the same radiological scan. This resource-intensive process can not only consume significant computing resources but also affect service availability for other users. An effective way to mitigate this risk is to implement rate limits that control the number of requests per IP address. Deploying tiered access policies prioritizing essential health care services can ensure availability during times of high demand. Furthermore, monitoring resource use and setting alerts for suspicious use patterns can mitigate resource-draining attacks [].
Postdeployment Security Framework
The vulnerabilities outlined above demonstrate that no single countermeasure is sufficient to secure an LLM deployed in a clinical environment. We propose a framework using multiple independent security layers such that the failure of any one layer does not compromise the entire system. presents a layered prompt-processing workflow that implements this principle across 4 key stages: access control, input validation, output governance, and human-in-the-loop oversight.

The first layer establishes user authentication and rate limiting. Before a prompt reaches the model, the user’s identity should be verified, and their request frequency should be checked against predefined thresholds. This prevents unauthorized access to the LLM and mitigates unbounded consumption attacks that could degrade service availability for legitimate users.
The second layer validates and sanitizes the input. Inputs should be programmatically screened for known adversarial patterns, especially malicious prompt injection signatures to which medical LLMs have been shown to be vulnerable. Only after clearing these initial checks is the prompt actually passed into the system.
The third layer governs the model’s output. When the generated response involves a tool or API call, an additional authorization gate should verify that the requested action falls within the model’s explicitly permitted scope. This enforces the principle of least privilege, ensuring the LLM can only access functions and data sources necessary for its designated task. If the API call fails validation, the action is blocked before it can interact with any external system.
The final and most critical layer applies to high-impact actions, such as modifying clinical records or initiating a referral. These actions should be routed to a human-in-the-loop review, where a clinician or system administrator must explicitly approve or reject the proposed action before it is executed. This ensures that consequential decisions are never delegated entirely to the model, thereby maintaining clinical accountability. Actions that are approved, along with nontool responses, should then pass through a final content-filtering and sanitization step before being delivered to the user.
Security evaluation should continue throughout clinical use. Organizations should periodically review model performance, safety events, and audit logs and conduct repeated adversarial testing as novel prompt injection, jailbreaking, and other attack techniques emerge. Identified vulnerabilities should inform subsequent improvements to system safeguards and further testing.
These safeguards should also operate within applicable health care privacy, cybersecurity, and clinical governance requirements. This includes limiting access to patient data; establishing incident response pathways; and ensuring that the monitoring, audit, and oversight mechanisms described above are maintained in accordance with relevant organizational and regulatory requirements. Where relevant, organizations must also comply with jurisdiction-specific requirements under frameworks such as the Health Insurance Portability and Accountability Act (HIPAA) and the General Data Protection Regulation (GDPR), as well as applicable medical device or AI-specific regulation.
Responsibility for these safeguards is shared: developers and vendors implement technical controls, hospital IT security teams configure and monitor access and system security, clinical governance committees define permitted clinical functions and oversight requirements, and frontline clinicians retain responsibility for approving consequential patient-level actions. consolidates the predeployment and postdeployment safeguards described above into a practical checklist, mapped to the stakeholders primarily responsible for implementing each safeguard.
| Domain | Checklist item | Primary responsibility | |
| Predeployment phase | |||
| Data provenance | Source training and fine-tuning data from reputable, human-moderated sources | Developers | |
| Data provenance | Verify data provenance documentation before deployment | Hospital IT security | |
| Model integrity | Verify foundation model provenance and document all modifications made before integration | Developers | |
| Model integrity | Independently evaluate the model against backdoor-specific benchmarks | Hospital IT security | |
| Performance and safety evaluation | Evaluate general performance and safety, including fairness, bias, and toxicity | Clinical governance | |
| Performance and safety evaluation | Evaluate handling of unsafe medical prompts | Clinical governance | |
| Performance and safety evaluation | Evaluate clinical reasoning in realistic, multistep scenarios | Clinical governance | |
| Deployment scoping | Define a clearly bounded clinical use case and assess the data accessed, actions permitted, potential failure modes, and implications for patients and clinical services | Clinical governance and hospital IT security | |
| Postdeployment phase | |||
| Access control | Authenticate all users and enforce rate limits | Hospital IT security | |
| Input validation | Screen all inputs, including user prompts and external content (medical notes, imported documents, and web content) for prompt injection and adversarial patterns | Developers | |
| Output governance | Enforce the principle of least privilege for tool and API access | Developers and clinical governance | |
| Output governance | Filter outputs for harmful content and sensitive information before delivery | Developers | |
| Human oversight | Require explicit clinician approval for high-impact actions (eg, modifying clinical records and placing referrals) before execution | Clinicians | |
| Monitoring and response | Maintain audit logs and alerts for suspicious use patterns and log all blocked requests and actions | Hospital IT security | |
| Monitoring and response | Periodically review model performance, safety events, and audit logs and repeat adversarial and jailbreak testing as novel attack techniques emerge | Hospital IT security and clinical governance | |
| Governance and compliance | Operate within applicable privacy, cybersecurity, and clinical governance requirements and maintain incident response pathways | Clinical governance and hospital IT security | |
Conclusions
The integration of AI into health care represents a pivotal development, offering opportunities to enhance clinical workflows, support decision-making, and improve patient outcomes. However, as this article highlights, a structured framework across the life cycle of LLMs is required to address the new frontier of cybersecurity risks.
Before deployment, the integrity of LLMs can be threatened by attacks such as data and model poisoning, which can corrupt a model’s foundational knowledge. After deployment, these systems are exposed to inference attacks, such as prompt injection, jailbreaking, and adversarial inputs, which can be exploited to expose sensitive data, generate harmful content, trigger unauthorized actions, or expose proprietary model assets.
Standard safety measures and performance benchmarks alone are insufficient to mitigate these sophisticated threats. A successful defense requires a deliberate, multilayered strategy that treats security as a core component of the system’s design. This approach must combine robust technical safeguards, such as strict input validation, adversarial training, and access controls, with rigorous procedural governance, including predeployment risk assessment, ongoing security and performance monitoring, transparent data provenance, adherence to the principle of least privilege, and the implementation of human-in-the-loop oversight for critical tasks.
As a conceptual framework, these recommendations require validation in real-world health care settings to assess their feasibility, effectiveness, and applicability across different clinical workflows. Ultimately, the successful deployment of LLMs in health care hinges on establishing and maintaining trust. This requires a proactive and collaborative commitment from developers, clinicians, and health system administrators to anticipate and systematically address these vulnerabilities. By placing security at the forefront of AI system integration, the medical community can confidently use the power of LLMs to build a safer and more efficient future for health care.
Acknowledgments
We used the generative AI tool Claude (claude.ai) to assist with sentence-level editing, which were further reviewed and revised by the study group.
Funding
The authors declared no financial support was received for this work.
Authors' Contributions
Conceptualization: KNC (lead), CJG (equal), HL (equal)
Methodology: CJG (lead), HL (equal), KNC (equal)
Supervision: KNC
Writing – original draft: CJG (lead), HL (equal), ND (equal), KNC (equal)
Writing – review & editing: CJG (lead), HL (equal), ND (equal), KNC (equal)
Conflicts of Interest
None declared.
References
- Beam AL, Drazen JM, Kohane IS, Leong TY, Manrai AK, Rubin EJ. Artificial intelligence in medicine. N Engl J Med. Mar 30, 2023;388(13):1220-1221. [CrossRef] [Medline]
- Tierney AA, Gayre G, Hoberman B, Mattern B, Ballesca M, Kipnis P, et al. Ambient artificial intelligence scribes to alleviate the burden of clinical documentation. NEJM Catal. Feb 21, 2024;5(3). [CrossRef]
- Friedman AB, Delgado MK, Weissman GE. Artificial intelligence for emergency care triage-much promise, but still much to learn. JAMA Netw Open. May 01, 2024;7(5):e248857. [FREE Full text] [CrossRef] [Medline]
- Sardesai N, Russo P, Martin J, Sardesai A. Utilizing generative conversational artificial intelligence to create simulated patient encounters: a pilot study for anaesthesia training. Postgrad Med J. Mar 18, 2024;100(1182):237-241. [FREE Full text] [CrossRef] [Medline]
- OWASP top 10 for LLM applications 2025. OWASP Foundation. Nov 17, 2024. URL: https://genai.owasp.org/resource/owasp-top-10-for-llm-applications-2025/ [accessed 2025-08-24]
- Alber DA, Yang Z, Alyakin A, Yang E, Rai S, Valliani AA, et al. Medical large language models are vulnerable to data-poisoning attacks. Nat Med. Feb 2025;31(2):618-626. [CrossRef] [Medline]
- Wan A, Wallace E, Shen S, Klein D. Poisoning language models during instruction tuning. Norfolk, MA. JMLR.org; 2023. Presented at: ICML'23: Proceedings of the 40th International Conference on Machine Learning; 23-29 July 2023:35413-35425; Honolulu. URL: https://proceedings.mlr.press/v202/wan23b.html
- Han T, Nebelung S, Khader F, Wang T, Müller-Franzes G, Kuhl C, et al. Medical large language models are susceptible to targeted misinformation attacks. NPJ Digit Med. Oct 23, 2024;7(1):288. [FREE Full text] [CrossRef] [Medline]
- Wolf T, Debut L, Sanh V, Chaumond J, Delangue C, Moi A, et al. Transformers: state-of-the-art natural language processing. Stroudsburg, PA. Association for Computational Linguistics; 2020. Presented at: 2020 Conference on Empirical Methods in Natural Language Processing; Nov 16-20, 2020:38-45; Virtual. URL: https://aclanthology.org/2020.emnlp-demos.6/
- Meng K, Bau D, Andonian A, Belinkov Y. Locating and editing factual associations in GPT. New York, NY. Curran Associates Inc; 2022. Presented at: NIPS '22: Proceedings of the 36th International Conference on Neural Information Processing Systems; Nov 28-Dec 9 2022:17359-17372; Louisiana. URL: https://proceedings.neurips.cc/paper_files/paper/2022/file/6f1d43d5a82a37e89b0665b33bf3a182-Paper-Conference.pdf
- Hubinger E, Denison C, Mu J, Lambert M, Tong M, MacDiarmid M, et al. Sleeper agents: training deceptive LLMs that persist through safety training. ArXiv. Preprint posted online on January 10, 2024. [CrossRef]
- Li Y, Huang H, Zhao Y, Ma X, Sun J. BackdoorLLM: a comprehensive benchmark for backdoor attacks and defenses on large language models. ArXiv. Preprint posted online on August 23, 2024. [CrossRef]
- Carlini N, Paleka D, Dvijotham K, Steinke T, Hayase J, Cooper AF, et al. Stealing part of a production language model. Norfolk, MA. JMLR.org; 2024. Presented at: ICML'24: Proceedings of the 41st International Conference on Machine Learning; Jul 21-27, 2024:5680-5705; Vienna. URL: https://proceedings.mlr.press/v235/carlini24a.html
- Liang P, Bommasani R, Lee T, Tsipras D, Soylu D, Yasunaga M, et al. Holistic evaluation of language models. ArXiv. Preprint posted online on November 16, 2022. [CrossRef]
- Han T, Kumar A, Agarwal C, Lakkaraju H. MedSafetyBench: evaluating and improving the medical safety of large language models. ArXiv. Preprint posted online on October 9, 2024. [CrossRef]
- Liu F, Li Z, Zhou H, Yin Q, Yang J, Tang X, et al. Large language models are poor clinical decision-makers: a comprehensive benchmark. Stroudsburg, PA. Association for Computational Linguistics; 2024. Presented at: 2024 Conference on Empirical Methods in Natural Language Processing; Nov 12-16, 2024:13696-13710; Miami. URL: https://aclanthology.org/2024.emnlp-main.759/ [CrossRef]
- Arora RK, Wei J, Hicks RS, Bowman P, Quiñonero-Candela J, Tsimpourlas F, et al. HealthBench: evaluating large language models towards improved human health. ArXiv. Preprint posted online on May 13, 2025. [CrossRef]
- Bedi S, Cui H, Fuentes M, Unell A, Wornow M, Banda JM, et al. Holistic evaluation of large language models for medical tasks with MedHELM. Nat Med. Mar 2026;32(3):943-951. [CrossRef] [Medline]
- Wei A, Haghtalab N, Steinhardt J. Jailbroken: how does LLM safety training fail? 2023. Presented at: NIPS '23: Proceedings of the 37th International Conference on Neural Information Processing Systems; Dec 10-16, 2023:80079-80110; Louisiana. URL: https://neurips.cc/virtual/2023/poster/70702
- Greshake K, Abdelnabi S, Mishra S, Endres C, Holz T, Fritz M. Not what you've signed up for: compromising real-world LLM-integrated applications with indirect prompt injection. ArXiv. Preprint posted online on May 5, 2023. [CrossRef]
- Clusmann J, Ferber D, Wiest IC, Schneider CV, Brinker TJ, Foersch S, et al. Prompt injection attacks on vision language models in oncology. Nat Commun. Feb 01, 2025;16(1):1239. [FREE Full text] [CrossRef] [Medline]
- Goodfellow IJ, Shlens J, Szegedy C. Explaining and harnessing adversarial examples. ArXiv. Preprint posted online on December 20, 2014. [CrossRef]
- Xhonneux S, Sordoni A, Günnemann S, Gidel G, Schwinn L. Efficient adversarial training in LLMs with continuous attacks. ArXiv. Preprint posted online on November 1, 2024. [CrossRef]
- Qiu J, Li L, Sun J, Wei H, Xu Z, Lam K, et al. Emerging cyber attack risks of medical AI agents. ArXiv. Preprint posted online on April 2, 2025. [CrossRef]
- Agarwal D, Fabbri A, Risher B, Laban P, Joty S, Wu CS. Prompt leakage effect and mitigation strategies for multi-turn LLM applications. Stroudsburg, PA. Association for Computational Linguistics; 2024. Presented at: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track; Nov 12-16, 2024:1255-1275; Miami. URL: https://aclanthology.org/2024.emnlp-industry.94/
- Shi T, He J, Wang Z, Wu L, Li H, Guo W, et al. Progent: programmable privilege control for LLM agents. ArXiv. Preprint posted online on April 16, 2025. [CrossRef]
- Zou HP, Huang WC, Wu Y, Miao C, Li D, Liu A, et al. A call for collaborative intelligence: why human-agent systems should precede AI autonomy. ArXiv. Preprint posted online on June 11, 2025. [CrossRef]
- Gao K, Pang T, Du C, Yang Y, Xia ST, Lin M. Denial-of-service poisoning attacks against large language models. ArXiv. Preprint posted online on October 14, 2024. [CrossRef]
Abbreviations
| EHR: electronic health record |
| GDPR: General Data Protection Regulation |
| HELM: holistic evaluation of language models |
| HIPAA: Health Insurance Portability and Accountability Act |
| LLM: large language model |
| MedHELM: holistic evaluation of large language models for medical tasks |
Edited by A Benis; submitted 18.May.2026; peer-reviewed by TAK Manne, H Vundavalli; comments to author 10.Aug.2026; revised version received 20.Aug.2026; accepted 23.Sep.2026; published 08.Oct.2026.
Copyright©Chun Joo Goh, Haoran Liu, Nethum Devendra, Khoa Nguyen Cao. Originally published in JMIR Medical Informatics (https://medinform.jmir.org), 08.Oct.2026.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work, first published in JMIR Medical Informatics, is properly cited. The complete bibliographic information, a link to the original publication on https://medinform.jmir.org/, as well as this copyright and license information must be included.

