Abstract
Abstract Large Language Models (LLMs) and foundation models are increasingly deployed in security-critical and high-impact settings, including healthcare, cybersecurity, software engineering, and intelligent infrastructure. Their open-ended interfaces and multimodal capabilities create new attack surfaces, where prompt injection, jailbreaking, hallucination, adversarial fine-tuning, and cross-modal manipulation can compromise reliability, integrity, privacy, and user trust. This paper presents a systematic literature review of 94 peer-reviewed studies identified through database searching, backward reference snowballing, venue screening, and quality assessment within a January 2021–July 2026 search window. The review synthesises evidence on safety and robustness vulnerabilities, red-teaming practices, mitigation mechanisms, evaluation metrics, and unresolved research gaps for LLMs and related foundation-model systems. The findings show that LLM vulnerabilities are rarely isolated: prompt-level, model-level, data-level, and deployment-level risks often interact, particularly in high-stakes domains. Red-teaming has become more structured, but existing practices remain inconsistent across threat models, benchmarks, attack settings, and reporting standards. The reviewed defences include prompt hardening, input/output filtering, adversarial training, retrieval-augmented grounding, cross-model verification, gatekeeper models, and domain-specific safety layers; however, their effectiveness is difficult to compare because studies use heterogeneous metrics and evaluation protocols. Major gaps include fragmented benchmarking, limited multilingual and low-resource assessment, insufficient adaptive-adversary and long-term deployment studies, and weak integration between technical defences and governance mechanisms. The review indicates that trustworthy foundation-model deployment requires standardised threat models, reproducible red-teaming protocols, transparent safety metrics, and defence-in-depth strategies tailored to domain-specific security risks.
Bullet Summary
- The paper addresses the security and trustworthiness challenges of deploying large language models (LLMs) and foundation models in critical and high-impact domains such as healthcare and cybersecurity.
- A systematic literature review of 94 peer-reviewed studies from January 2021 to July 2026 was conducted to examine safety vulnerabilities, evaluation methodologies, defense mechanisms, and research gaps related to foundation models.
- Identified vulnerabilities span prompt-level attacks, model-level issues, data-level risks, and deployment-level threats, which often interact in complex ways, especially in sensitive applications.
- The review highlights the evolution of red-teaming practices, noting increased structure but persistent inconsistencies in threat models, benchmarks, attack scenarios, and reporting standards.
- Existing defense strategies include prompt hardening, input/output filtering, adversarial training, retrieval-augmented grounding, cross-model verification, gatekeeper models, and specialized safety layers tailored to specific domains.