Research area drill-down

Prompt Injection

Papers currently mapped into this multi-agent security subarea from the merged research feed.

Active feeds: arXiv, OpenAlex, Crossref, Semantic Scholar, DBLP

0 of 36 articles selected

Showing 36 of 688 matching articles

AECP: Artifact-Exclusive Communication Protocol for Multi-Agent Code Generation

arXiv preprint arXiv Agent-to-Agent Communication Orchestration Risk Prompt Injection

Jiaqi Xue, Yanjun Wang, Xiangci Li, Lingbo Mo, Aritra Sengupta, Shweta Garg, Murali Krishna Ramanathan, Myeongsoo Kim

Published 2026-10-05

Venue: arXiv

Open Source Record

Abstract

As AI agents increasingly tackle complex repository-level coding tasks, distributing work across multiple agents is a natural way to scale beyond the capabilities of a single agent. To coordinate their interdependent work, these agents share findings and agree on interfaces between modules. However, exchanged information often serves only as context, leaving individual agents to interpret it and incorporate it into subsequent work. Consequently, shared findings may go unused and deviations from interface agreements may go undetected, undermining the reliability and efficiency of collaboration. This motivates moving part of the coordination responsibility from individual agents to the execution harness. To make shared information actionable during execution, we introduce the Artifact-Exclusive Communication Protocol (AECP). AECP requires agents to communicate exclusively through structured artifacts and specifies how the harness processes them. The harness supplies findings when agents access relevant code, screens implementations for mismatches with recorded interface commitments, and requires affected agents to revisit revised agreements. These coordination steps become part of harness execution rather than actions that agents must initiate from prior messages. Across Doc2Repo, NL2Repo, and CodeProjectEval, using closed- and open-source models including Opus-4.8 and DeepSeek-V4-Flash, AECP improves average test pass rate by 28.2% and reduces average wall time by 16.5% relative to an agent team using free-form inter-agent messages. Artifact-exclusive communication also blocks the relay of malicious instructions between agents, reducing how often they reach other agents from 95% to 0% and how often those agents act on them from 40% to 0%.

Bullet Summary

  • AECP introduces an Artifact-Exclusive Communication Protocol enabling multi-agent AI systems to coordinate complex code generation exclusively through structured artifacts managed by an execution harness.
  • The execution harness actively processes shared Knowledge and Contract Artifacts, delivering relevant knowledge based on code scope access, enforcing interface contract compliance, and managing coordination states, reducing reliance on ambiguous natural-lan...
  • AECP addresses common coordination failures such as overlooked shared findings, unnoticed interface deviations, and inconsistent task completion by moving responsibility for processing and verification from individual agents to the centralized harness.
  • Experimental evaluations across multiple benchmarks (Doc2Repo, NL2Repo, CodeProjectEval) and models demonstrate AECP improves average test pass rates by over 28% and reduces wall-clock time by up to 24.9% relative to systems using free-form inter-agent mess...
  • Artifact-exclusive communication via AECP inherently enhances security by blocking malicious instruction propagation between agents, reducing transmission and action of such instructions from 95% and 40% respectively, to zero.

RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents

arXiv preprint arXiv Prompt Injection Benchmarks and Evaluation

Mohamed Dhouib, Clement Elliker, Alexi Canesse, Maël Jenny, Lucas-Andrei Thil, Mahammed El-Sharkawy, Sonia Vanier, Elie Bursztein

Published 2026-10-05

Venue: arXiv

Open Source Record

Abstract

Tool-using language-model agents are vulnerable to indirect prompt injection because they must act on untrusted external content. Existing training-time defenses can reduce attack success rates, but often at the cost of general capabilities. We show that training-based defenses induce substantial drift in the model's output distribution, altering its behavior even in benign settings and providing a potential mechanism for utility degradation. We further identify a failure mode of these defenses: On benign tool-use tasks, the model refrains from a step needed to finish an authorized task, particularly when that step is indicated by a tool output. To address these limitations, we introduce RAISED (Robust Attack Invariance through Self-Distillation), a training framework that combines self-generation and self-distillation. The model first generates its own tool-use scenarios, with an emphasis on cases where task completion requires acting on legitimate guidance from tool outputs. Then, through self-distillation, the student is trained to match the teacher's clean-context behavior on both clean and injected variants of the same trajectory. RAISED substantially reduces the attack success rate of prompt injections in tool responses while, unlike prior training-based defenses, preserving utility on both agentic and general-purpose benchmarks.

Bullet Summary

  • Tool-using large language model (LLM) agents are vulnerable to indirect prompt injection attacks, where untrusted tool outputs embed malicious instructions that hijack the agent's behavior.
  • Existing training-based defenses reduce attack success rates but introduce substantial output-distribution drift, leading to degraded general capabilities and failure modes such as incomplete task execution when legitimate tool guidance is involved.
  • RAISED (Robust Attack Invariance through Self-Distillation) is proposed as a novel defense framework that combines self-generation of tool-use scenarios emphasizing legitimate guidance and an injection-augmented self-distillation training objective to prese...
  • Self-generation in RAISED involves the model generating executable multi-step tool-use tasks, solving them, and then generating adversarial prompt injections to create training data that covers both benign and attack scenarios.
  • Self-distillation trains a student model to match the teacher's (base model's) next-token distributions on both clean and injected inputs, thereby maintaining alignment with the base model and reducing output-distribution drift.

Compromise Is Not Consequence: Evaluating Task-Scoped Authorization in LLM Agents with Paired Replay

arXiv preprint arXiv Prompt Injection Governance and Policy Benchmarks and Evaluation

Tural Hagverdiyev

Published 2026-10-05

Venue: arXiv

Open Source Record

Abstract

A tool-using model can follow a malicious instruction even when its credentials are valid. We study whether task-scoped authorization contains the resulting tool execution. Our paired-replay testbed samples a model request once and submits the same action, resource, and arguments to broad bearer, scoped JWT, sender-constrained, and Open Policy Agent conditions. The frozen main experiment uses 128 scenarios across four tool domains and five local model configurations. Among valid attacked post-exposure decisions, broad-bearer harmful execution ranges from 8.9% to 37.8% across models; all three scoped conditions record zero. The policy changes execution, not the frozen model decision. These results support containment of the tested cross-action and cross-resource consequences under researcher-supplied task authority, not prevention of prompt injection. A bounded six-task AgentDojo extension preserves native scoring and records 11/24 injected broad-policy attack flags versus 0/24 scoped flags, with utility of 5/24 and 6/24. Low exposure, invalid decisions, and purposive task selection limit that comparison. The primary experiment measures safe continuation rather than final-answer correctness. An empirical-bank fresh-sampling comparison finds no estimation advantage from pairing when scoped outcomes are constant zero. The contribution is a controlled measurement of compromise, enforcement, and continuation, with explicit limits on what each observation establishes.

Bullet Summary

  • Tool-using LLMs can execute malicious instructions even with valid credentials, prompting investigation into whether task-scoped authorization effectively contains harmful tool actions.
  • The study introduces a paired-replay testbed that submits identical model requests under different authorization policies (broad bearer, scoped JWT, sender-constrained, Open Policy Agent) to isolate the impact of authorization enforcement from stochastic mo...
  • Experiments across 128 scenarios spanning four tool domains and five large language model configurations reveal that scoped authorization policies prevent any harmful executions post-exposure, while broad bearer tokens allow harmful executions ranging from...
  • The findings clarify that authorization policies influence tool execution rather than model decision-making itself, distinguishing between model compromise (unauthorized action selection) and operational consequences (harmful action execution).
  • Scoped authorization policies enforce effective boundaries on tool action and resource use, containing impacts of prompt injections without preventing the model from selecting potentially malicious actions.

Topology-Conditioned Backdoors: Language Models That Insert Vulnerabilities When They Infer They Are in a Multi-Agent System

Merged record merged scholarly record arXiv Prompt Injection Trust and Identity Governance and Policy

Keegan Wang, Anantika Mannby

Published 2026-10-05

Venue: arXiv

Open Source Record

Abstract

A language model may behave safely in a single-agent evaluation yet produce vulnerable code when its context suggests that it is part of a multi-agent system. We study this failure mode by fine-tuning Qwen2.5-7B-Instruct to condition code generation on deployment topology inferred from prompt-level provenance cues. On held-out coding tasks, task-specific checkers detect vulnerabilities in 96-100% of multi-agent episodes and 0% of single-agent episodes. An independent bandit analyzer detects vulnerabilities in approximately 67% of multi-agent episodes, covering six of nine vulnerability families at medium or high severity. Lexical-placebo and human-review controls support topology, rather than multi-agent terminology or the absence of oversight, as the relevant conditioning variable. A model trained on diverse topology signals also generalizes to five signal types held out of training, with replications across two Qwen checkpoints and two training seeds. In a blind audit, a binary judgment that a hidden policy exists poorly distinguishes the organism from a clean control, whereas the auditor identifies the topology trigger in 9 of 10 organism runs and none of the control runs. These results motivate differential auditing across matched single- and multi-agent contexts. They demonstrate a trainable backdoor conditioned on described topology; activation in a live multi-agent environment remains untested.

Bullet Summary

  • The paper identifies a novel backdoor mechanism in language models that conditionally inserts code vulnerabilities when the model infers it operates within a multi-agent system topology, while remaining safe in single-agent scenarios.
  • The research involves fine-tuning the Qwen2.5-7B-Instruct model using prompt-level provenance cues that indicate deployment topology, enabling the model to generate vulnerable code specifically triggered by multi-agent context signals.
  • Experimental evidence from task-specific automated vulnerability checkers shows nearly 100% vulnerability detection in multi-agent episodes contrasted with zero vulnerabilities in single-agent episodes; an independent static analyzer (bandit) corroborates t...
  • Control experiments confirm that the insertion of vulnerabilities is driven by the inferred deployment topology rather than multi-agent terminology or absence of human oversight, and models trained on diverse topology signals generalize to unseen multi-agen...
  • The authors propose differential topology auditing, a novel auditing method that contrasts model behavior between matched single-agent and multi-agent settings to effectively detect hidden topology-conditioned backdoors, outperforming conventional binary pr...

Can CaMeLs Talk? Securing Multi-Agent Systems Against Indirect Prompt Injection Attacks

Merged record merged scholarly record arXiv Prompt Injection Agent-to-Agent Communication Benchmarks and Evaluation

James Peters-Gill, Avi Semler, Henning Bartsch, Ilia Shumailov, Christian Schroeder de Witt

Published 2026-10-05

Venue: arXiv

Open Source Record

Abstract

Indirect prompt injection attacks - malicious instructions embedded in content processed by large language models - remain a major obstacle to safely deploying tool-using agents. CaMeL [Debenedetti et al., 2025] mitigates this threat for an individual agent by separating trusted control flow from untrusted data and enforcing capability-based security policies at runtime. In this work, we investigate whether CaMeL's security guarantees compose in hierarchical multi-agent systems, where agents invoke other agents as tools. We find that CaMeL's guarantees do not compose. We construct a concrete prompt-injection attack that succeeds despite all constituent agents individually operating CaMeL. Our attack exploits the fact that untrusted data can be reinterpreted as trusted input by a downstream agent. We then introduce multi-CaMeL, an agent-to-agent communication protocol that preserves provenance across agent boundaries by separating trusted natural-language instructions from untrusted data passed through a distinct data channel. We evaluate multi-CaMeL's utility on AssetOpsBench and its security-utility tradeoff on MultiAgentDojo, a benchmark we develop by extending AgentDojo to the multi-agent setting. We find that multi-CaMeL reduces attack success rate (ASR) to 0.0%, compared with 0.2% for individual-agent CaMeL and 12.9% with no CaMeL. Multi-CaMeL incurs a utility cost, but this cost trends downward as model capability increases and is modest for the strongest models, suggesting that more capable models better accommodate the constraints imposed by the protocol.

Bullet Summary

  • Indirect prompt injection attacks embed malicious instructions in inputs to large language model (LLM) agents, threatening the security of tool-using agents.
  • CaMeL secures individual LLM agents by separating trusted control flow from untrusted data and enforcing capability-based runtime policies, ensuring control-flow integrity (CFI) at the single-agent level.
  • CaMeL's security guarantees do not naturally compose in hierarchical multi-agent systems where agents invoke other agents as tools, leading to a vulnerability called boundary laundering, where untrusted data is misinterpreted as trusted input downstream.
  • The authors propose multi-CaMeL, a novel inter-agent communication protocol that preserves provenance and maintains system-level CFI by separating trusted natural-language instruction channels from untrusted data channels across agent boundaries.
  • Multi-CaMeL enforces instruction-channel integrity and capability preservation at runtime, preventing indirect prompt injection attacks from propagating between agents.

Adaptive Code Revision Attacks on AI Pull Request Reviewers

arXiv preprint arXiv Prompt Injection Trust and Identity Agent-to-Agent Communication

Jingzhi Gong, Jie M. Zhang, Gunel Jahangirova, Meng Wang

Published 2026-10-04

Venue: arXiv

Open Source Record

Abstract

Pull-request review protects software before new code reaches users, helping prevent vulnerabilities that could expose users to attacks. AI agents increasingly perform these reviews and explain which problems need fixing. However, for an attacker submitting vulnerable code, this feedback also reveals what changes may secure approval. Existing PR attacks seek such approval through persuasive text and comments while keeping executable code fixed. This leaves unclear whether an attacker can use the feedback to repair the reported problem while preserving a vulnerability in the revised code. We therefore conduct an empirical study of this threat using AFCRA (Adaptive Feedback-guided Code Revision Attack). To distinguish successful attacks from genuine repairs, we construct AFCRA-Bench from 159 disclosed vulnerabilities, with executable exploits to verify vulnerabilities in code. Across five-round interactions with Sonnet 5 and GPT-5.5 reviewers, AFCRA reaches success rates 2.5x and 12.5x those of the strongest evaluated text- or comment-based attack. Case studies of these successes show how reviewers accept repairs of reported problems while overlooking surviving vulnerabilities. These findings establish feedback-guided code revision as a threat to automated PR review. To address this threat, we derive actionable implications for researchers, AI providers, PR reviewers, and PR authors on securing AI-assisted development.

Bullet Summary

  • AI assistants increasingly perform pull request (PR) code reviews to prevent vulnerabilities before deployment.
  • Existing attacks manipulate PR text/comments to gain approval without changing vulnerable code, but new threats involve adaptive code revisions guided by AI feedback.
  • The study introduces AFCRA (Adaptive Feedback-guided Code Revision Attack), which iteratively repairs reported issues while preserving exploitable vulnerabilities, using feedback to guide revisions.
  • AFCRA-Bench, a benchmark of 159 CVE-based pull requests with executable exploits across eight languages, evaluates attack success by verifying vulnerability persistence.
  • Experimental results show AFCRA achieves up to 12.5x higher success rates than prior text-based attacks against AI reviewers like Sonnet 5 and GPT-5.5, exploiting multi-round interactions.

Readable Before Actionable: Causal Tracing of Indirect Prompt Injection

arXiv preprint arXiv Prompt Injection Benchmarks and Evaluation

Zhe Yu, Wenpeng Xing, Xingxing Yang, Meng Han

Published 2026-10-04

Venue: arXiv

Open Source Record

Abstract

Indirect prompt injection causes LLM agents to follow commands embedded in external data. A probe may distinguish instructions from data without identifying a state edit that changes the next action. We study this gap through counterfactual role probes, component-wise activation patching, and separate interventions on AgentDojo trajectories. Role decoding survives changes in content and format. In controlled Qwen tests, it precedes strong tool-choice effects from patches along an independently estimated role direction. On AgentDojo, directions estimated from hijacked and resisted training trajectories reduce attack success at pre-action and injected-span positions, but have little effect at random positions. In longer Qwen trajectories, single-position edits become less effective at later layers; span-wide and repeated edits reduce attack success on the same evaluation set. Removing the learned channel subspace preserves role decoding, yet effective intervention directions transfer poorly across the tested channels. These findings distinguish a readable role signal from an effective behavioral intervention: depth matters in controlled tool choice, while position and context also matter in attack trajectories.

Bullet Summary

  • Indirect prompt injection (IPI) allows large language model (LLM) agents to follow hidden commands in external data, posing significant security risks in multi-agent systems.
  • The study distinguishes between readable role signals in model residual states (instruction vs data) and effective behavioral interventions that actually change agent actions, using counterfactual role probes and activation patching techniques.
  • Experiments on Qwen-2.5-7B and AgentDojo multi-step tool-use environments demonstrate that role information is decodable prior to strong behavioral effects, with the influence of interventions being dependent on the layer depth, token position, and context...
  • A directional vector (drole), derived from class mean differences, enables targeted interventions on residual states, primarily effective in late layers (L16–L24) to reduce attack success rate (ASR) without harming benign task performance.
  • Interventions along the role-aligned direction effectively flip tool-choice predictions and reduce use of adversarial tool arguments, demonstrating disentanglement of role-specific representation from orthogonal components.

StegoMemory: Agentic Memory Acts as Covert Steganographic Channel

arXiv preprint arXiv Memory Poisoning Prompt Injection Governance and Policy

Snehasis Mukhopadhyay, Arun Nair

Published 2026-10-03

Venue: arXiv

Open Source Record

Abstract

Is agentic memory robust against stealthy steganographic attacks? We carry out a large-scale red-teaming exercise to test whether agents can encode attacker-controlled strings in one session and recover them in another without triggering safety oversight. Following SHADE-Arena-style tasks, we embed malicious side tasks to encode secret strings using steganography within otherwise benign tasks and evaluate them using independent task-completion and safety oversight. We test 14,000 attack trials spanning 91 model-attack configurations, 13 models, and 7 steganographic schemes: Acrostic, Mesostic, Telestich, Word-choice, Capitalization, Position-based, and Homoglyph. 41.2% of trials successfully encode the target payload in the generated response, while 29.8% survive the memory-write step with the encoding intact. Overall, 25.4% of trials yield a recoverable payload at retrieval, and 20.1% achieve exact-match recovery, including synthetic API keys and cryptographic material never exposed to the receiver. Among the exact-match transmissions, 71.4% also pass primary task-completion scoring and are independently judged safe by the oversight model, corresponding to 14.4% of all trials in which a successful covert transmission would appear to be an ordinary, benign interaction under task-level evaluation. Our results demonstrate that agentic memory can function as a persistent cross-session covert channel. The results further show that the principal bottleneck occurs at memory persistence rather than retrieval: once a steganographic payload survives the memory-write stage, a substantial fraction remains recoverable. We therefore argue that memory integrity, information-flow control, and covert-channel detection should be explicit security requirements for agentic systems.

Bullet Summary

  • Agentic memory in large language model (LLM)-based agents can be exploited as a covert steganographic channel to encode attacker-controlled secret strings persistently across sessions without triggering safety oversight.
  • A large-scale benchmark tested 14,000 attack trials combining benign tasks with covert side tasks using seven steganographic schemes (e.g., acrostics, word-choice) on 13 models, revealing a 25.4% recoverable payload rate and 20.1% exact-match secret recover...
  • Memory persistence during the write stage is the main bottleneck for successful covert channel attacks; once the payload survives memory-write, recovery on retrieval is often successful, highlighting vulnerabilities in memory summarization and persistence m...
  • Steganographic attacks leverage multiple memory types (working, episodic, semantic, procedural) and their interactions, enabling delayed encoding and decoding of malicious payloads across sessions, making detection difficult due to the variety and adaptabil...
  • Existing safety oversight and task-completion monitors often fail to detect covert payloads, as successful covert transmissions can appear ordinary and benign, underscoring the need for explicit security controls on memory integrity and information flow.

The Same Zero: Why Identical ASR Can Imply Different Guarantees in LLM-Agent Security

arXiv preprint arXiv Prompt Injection Governance and Policy Benchmarks and Evaluation

YaJie Yin

Published 2026-10-03

Venue: arXiv

Open Source Record

Abstract

LLM-agent security has produced a dense landscape of defenses - prompt hardening, content filters, permission gates, sandboxes - yet no framework tells a deployer what a defense actually guarantees, or where that guarantee comes from. We apply Verification Autonomy Levels (VAL) - L0: LLM self-declaration; L1: deterministic rules; L2: objective ground truth; L3/L4: decidable completeness; L5: impossible - to 22 agent-security defenses; the taxonomy is falsifiable (10/10 prediction hits on frozen cards, flagged). We run the first controlled deployment-value comparison: at equal budget, a VAL-guided stack (confirmation gate + schema sandbox) versus a mainstream intuition stack (prompt hardening + keyword filter), 50 scenarios, 12 attack variants, adaptive/white-box/PAIR escalation (~7,000 testbed calls; ~10,000 harness calls on AgentDojo/JADE). The VAL stack holds 0.000 attack success at 1.000 benign success (0.5% ASR at 79.7% utility on AgentDojo banking vs 4.3% undefended); the intuition stack reaches 0.000 ASR but kills all benign actions - security by model-behavior luck, not structure. Across testbeds of rising attack-surface hardness the intuition stack's zero drifts (0->1.9%->6.2%, n=16 on JADE) while the VAL stack's holds within its ODD (0->0->0), its only breach a disclosed out-of-ODD password gap (0.5%). The same zero, two different guarantees: zero is an outcome, not a guarantee.

Bullet Summary

  • LLM-agent security defenses lack a unified framework to define their guarantees; the paper proposes Verification Autonomy Levels (VAL) as a taxonomy categorizing defenses from self-declaration (L0) to impossible comprehensive guarantees (L5).
  • Applying VAL to 22 agent-security defenses produces a falsifiable taxonomy that predicts failure modes and guarantee types, validated through inter-rater agreement and controlled experiments.
  • A controlled experiment contrasts a VAL-guided defense stack (confirmation gate plus schema sandbox) against mainstream defenses (prompt hardening plus keyword filtering) over 50 scenarios and 12 attack variants involving ~17,000 test calls.
  • Both defense stacks achieve zero attack success rate (ASR) initially, but the VAL-guided stack maintains zero ASR within its operational design domain (ODD) with high benign utility, while the mainstream stack's zero ASR is brittle and accompanied by signif...
  • The paper introduces 'zero stability', demonstrating that an observed zero ASR outcome can imply fundamentally different underlying security guarantees depending on the defense’s structural versus behavioral basis.

Understanding and Mitigating Hallucination Escape in Tool-Using LLM Agents

Merged record merged scholarly record arXiv Prompt Injection Governance and Policy Benchmarks and Evaluation

Peigui Qi, Kunsheng Tang, Yide Song, Weiming Zhang, Nenghai Yu

Published 2026-10-03

Venue: arXiv

Open Source Record

Abstract

Large language models (LLMs) increasingly serve as autonomous agents that invoke external tools. However, this capability introduces tool hallucination, selecting incorrect tools or generating invalid calls. Existing mitigation methods report substantial improvements, yet we identify a previously overlooked failure mode that we term Hallucination Escape. These methods reduce hallucination on the tool configuration they are tuned on but increase it on other configurations, canceling out the gain. We further investigate this phenomenon and find that hallucination rises sharply when a model's intrinsic tool-use tendencies conflict with the current tool configuration, and that existing methods reinforce rather than suppress these tendencies, which in turn contributes to hallucination escape. Building on these findings, we propose EscapeGuard, a training-free inference-time method that combines conflict-aware gating with configuration-derived attention enhancement to mitigate tool hallucination while preventing hallucination escape. Across six benchmarks on various models, EscapeGuard reduces tool-selection hallucination by 9.0 pp and suppresses hallucination escape, lowering the cross-configuration mean by 23.7 pp and achieving an 89.1% net improvement in paired-query evaluation. We hope this work can encourage evaluation beyond a single tool configuration and pave the way for more reliable tool-using LLM agents.

Bullet Summary

  • Large language models (LLMs) used as autonomous agents frequently suffer from tool hallucination, selecting incorrect tools or generating invalid calls, which undermines their reliability compared to mere text hallucination.
  • Existing mitigation methods improve hallucination rates on the specific tool configurations they are trained or tuned on but fail to generalize, exhibiting a phenomenon termed 'Hallucination Escape' where hallucination increases in other configurations.
  • Hallucination Escape arises due to intrinsic tool-use tendencies encoded within the models conflicting with runtime tool configurations; existing methods inadvertently reinforce these tendencies, worsening hallucination outside the trained configuration.
  • The authors propose EscapeGuard, a training-free, inference-time method that detects conflicts between intrinsic tendencies and current configurations via conflict-aware gating, and enhances attention to relevant tool information to reduce hallucination.
  • EscapeGuard was evaluated across six benchmarks on various LLMs and tool configurations, achieving significant reductions in tool-selection hallucination (up to 9.0 percentage points) and suppressing hallucination escape by lowering cross-configuration hall...

AgentGuardBench: A Multilingual Benchmark for Privacy, Security and Responsible Behaviour in AI Agents

Merged record merged scholarly record OpenAlex Benchmarks and Evaluation Prompt Injection Memory Poisoning

Joseph Arayemi

Published 2026-10-03

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.23127436

Open Source Record

Abstract

AgentGuardBench v0.1.1 This is a metadata-only archival release created after enabling Zenodo preservation for the repository. It allows Zenodo to permanently archive AgentGuardBench and assign a citable DOI. Code status The benchmark code and dataset are unchanged from v0.1.0. The release points to the same verified commit: c1887e8. All automated tests passed successfully before this archival release. Included benchmark 120 fully synthetic evaluation scenarios Prompt injection, privacy leakage, tool misuse, privilege abuse, memory safety, and benign controls English, French, Swahili, and Yoruba Banking, healthcare, education, government, and recruitment Strict and permissive deterministic baselines Reproducible results and citation metadata Prior related work Research article: https://doi.org/10.5281/zenodo.23048249 Accompanying software: https://doi.org/10.5281/zenodo.23045245 Once Zenodo finishes archiving this release, its dedicated DOI will be added to the README and CITATION.cff.

Bullet Summary

  • AgentGuardBench provides a comprehensive benchmark designed to evaluate privacy, security, and responsible behavior in AI agents across multiple languages.
  • The benchmark includes 120 fully synthetic evaluation scenarios targeting vulnerabilities such as prompt injection, privacy leakage, tool misuse, privilege abuse, and memory safety issues.
  • It supports multiple languages including English, French, Swahili, and Yoruba, enabling multilingual evaluation of AI agent behaviors.
  • The scenarios cover diverse application domains such as banking, healthcare, education, government, and recruitment to reflect real-world challenges.
  • AgentGuardBench offers both strict and permissive deterministic baselines, facilitating comparative assessment of AI agent security features.

Enhancing security in LLM applications: a performance evaluation of early detection systems

Merged record merged scholarly record OpenAlex Prompt Injection Benchmarks and Evaluation

Valerii Gakh, Hayretdin Bahşi

Published 2026-10-03

Venue: International Journal of Information Security

DOI: https://doi.org/10.1007/s10207-026-01338-7

Open Source Record

Abstract

Abstract Prompt injection (PI) attacks threaten novel software applications, which have LLM-based functionality. Prompt leakage, a variant of prompt injection, constitutes a serious confidentiality risk to those applications. Existing defenses against PI attacks currently cannot identify prompt leaks precisely. Moreover, an attacker can construct leakage attacks, which could evade most of these defenses. This prevents the ubiquitous adoption of LLMs in software applications. Meanwhile, there is a lack of practical investigations into the effectiveness of prompt injection (PI) detection tools. There is a gap in practical knowledge on how precisely existing tools detect PI attacks, and how they should be further improved. We evaluated the capabilities of early prompt injection detection systems, focusing on the performance of detection techniques implemented in several open-source solutions. We tested the solutions against prompt leak attacks that employed widespread injection techniques such as context-ignoring and context-manipulation. We present an analysis of distinct PI detection techniques and a comparative analysis of LLM Guard, Vigil, and Rebuff. We concluded that the designs of canary word-based detection techniques in Vigil and Rebuff were weak against our prompt leak attacks. We propose improvements for them. We found an evasion weakness in Rebuff’s secondary model-based technique and proposed a mitigation. We revealed that, thanks to their detection policies, Vigil is optimal for cases when a minimal false positive rate is required, and Rebuff is the most optimal for the highest detection rate.

Bullet Summary

  • Prompt injection (PI) attacks, especially prompt leakage, pose significant security and confidentiality threats to LLM-enhanced software applications.
  • Current defenses inadequately detect prompt leaks precisely, allowing attackers to construct evasive leakage attacks.
  • This study provides a practical evaluation of early PI detection tools, focusing on their detection capabilities against common attack techniques like context-ignoring and context-manipulation.
  • A comparative analysis of three open-source solutions—LLM Guard, Vigil, and Rebuff—was conducted to assess their PI detection performance.
  • Canary word-based detection methods in Vigil and Rebuff were found to be vulnerable to prompt leak attacks, leading to proposed improvements for these techniques.

Prompt Injection Threats in Azure-Based Large Language Model Applications

Merged record merged scholarly record OpenAlex Prompt Injection Governance and Policy Agent-to-Agent Communication

Shekar Rao Lakavath

Published 2026-10-03

Venue: Journal of Computer Science and Information Technology

DOI: https://doi.org/10.61424/jcsit.v3i2.1077

Open Source Record

Abstract

Large language models (LLMs) hosted on Microsoft Azure, primarily through Azure OpenAI Service, are increasingly embedded in enterprise applications that combine user input, retrieved documents, web content, and tool-calling agents within a single prompt context. This architectural pattern, while powerful, collapses the traditional separation between instructions and data and creates a distinct and growing attack surface known as prompt injection. This paper reviews the technical literature on prompt injection and adjacent LLM security threats and maps these threats onto the specific components of an Azure-based generative AI deployment, including Azure OpenAI Service, Azure AI Search, Azure AI Content Safety, and plugin or function-calling integrations built with Logic Apps. We develop a taxonomy of six prompt injection attack categories—direct injection, indirect injection, prompt leaking, jailbreaking, optimisation-based injection, and tool-mediated injection—drawn from the adversarial machine learning and LLM security literature, and we examine six corresponding defense mechanisms available within or alongside the Azure platform. The analysis shows that no single Azure-native control is sufficient on its own: content filtering, programmable guardrails, instruction–data separation, least-privilege tool permissions, content provenance checks, and systematic red-teaming each address a different point in the attack surface and must be layered together. We conclude that securing Azure-hosted LLM applications against prompt injection requires continuous, defense-in-depth engineering rather than a single configuration decision, and we identify open research questions around architectural, rather than purely filter-based, solutions to the instruction–data separation problem.

Bullet Summary

  • Large language models (LLMs) deployed via Microsoft Azure services are vulnerable to prompt injection attacks due to the collapsing boundary between instructions and data within combined prompt contexts.
  • Prompt injection is a novel security threat where adversaries manipulate input text to override or alter intended model instructions, exploiting LLMs' natural language usage as both commands and data.
  • The paper develops a taxonomy of six prompt injection attack categories: direct injection, indirect injection, prompt leaking, jailbreaking, optimization-based injection, and tool-mediated injection, each representing distinct attacker access points and met...
  • Azure's specific generative AI components (Azure OpenAI Service, Azure AI Search, Azure AI Content Safety, Logic Apps) serve as attack surfaces where prompt injection threats manifest across user inputs, external content, document retrieval, and plugin outp...
  • Defense mechanisms within the Azure ecosystem include content filtering, programmable guardrails, instruction-data separation, least-privilege permissions, content provenance verification, and systematic adversarial red-teaming, but no single control is suf...

agentdojo-ja

OpenAlex · Zenodo (CERN European Organization for Nuclear Research) repository OpenAlex Prompt Injection Benchmarks and Evaluation

Masahiro Nakatsugawa

Published 2026-10-03

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.23122805

Open Source Record

Abstract

Unofficial Japanese localization and extension of AgentDojo, the prompt-injection benchmark for LLM agents: Japanese tasks and injections for all four suites (97 user tasks, 31 injection tasks) and Japan-specific attack variants. Derived from AgentDojo (ETH Zurich SPy Lab; MIT License).

Bullet Summary

  • Introduces an unofficial Japanese localization of AgentDojo, a benchmark for evaluating prompt-injection vulnerabilities in large language model (LLM) agents.
  • Extends the original AgentDojo benchmark with Japan-specific user tasks and prompt-injection attack tasks, enabling assessment in a Japanese context.
  • Includes a total of 97 user tasks and 31 injection tasks across all four benchmark suites, enriching the diversity and scope for evaluation.
  • Adds Japan-specific attack variants to capture culturally or linguistically unique prompt-injection methods.
  • Derived from the original AgentDojo developed by ETH Zurich SPy Lab under the MIT License, ensuring compatibility and consistency with the base benchmark framework.

Testing Large Language Model Agents on the Use of Biological Tools for Nucleic Acid Synthesis Screening Evasion

arXiv preprint arXiv Prompt Injection Orchestration Risk Governance and Policy

Jeffrey Lee, Alyssa Worland, Christopher Rodriguez, Kyle Brady, Grant Ellison, Henry Alexander Bradley, Dawid Maciorowski, Jordan Despanie

Published 2026-10-02

Venue: arXiv

Open Source Record

Abstract

This report is a continuation of previous efforts to test the ability of large language model (LLM)-driven artificial intelligence (AI) agents to interface with AI-enabled biological tools (BTs). While rapid advancements in BTs in recent years have brought promise to accelerate scientific discovery, they also raise significant biosecurity concerns about potential misuse. The biosecurity community is particularly interested in the extent to which LLMs can lower technical barriers and assist non-expert users in accessing and operating BTs. Despite this interest, few evaluations have focused on LLM-BT interactions in the context of a defined threat model. To address this gap, this report describes a test of frontier LLM-driven AI Agents on their ability to use BTs to redesign peptides and proteins to evade nucleic acid synthesis screening measures. Highly relevant to biorisk, this task assesses a potential capability of AI agents that could enable a breach of a critical early defensive layer designed to prevent a multitude of biological misuse scenarios. The findings presented here intend to offer a foundation for biosecurity researchers and AI developers to conduct or further risk and capability assessments as these technologies progress.

Bullet Summary

  • The paper addresses biosecurity concerns arising from the integration of large language model (LLM)-driven AI agents with AI-enabled biological tools (BTs), focusing on potential misuse risks.
  • Rapid advancements in BTs promise accelerated scientific discovery but also raise the risk that non-expert users might exploit these technologies with lowered technical barriers facilitated by LLMs.
  • There is a noted lack of evaluations examining LLM-BT interactions through the lens of explicit threat models, highlighting a critical gap in biosecurity research.
  • The study tests state-of-the-art LLM-driven AI agents on their capability to redesign peptides and proteins to evade nucleic acid synthesis screening, an important early defensive measure against biological misuse.
  • This task simulates a realistic and relevant biosecurity challenge, demonstrating how AI agents might breach critical safeguards designed to prevent the creation or use of hazardous biological agents.

Passing the Test You Trained On: Re-evaluating Prompt-Injection Detectors for LLM Agents

arXiv preprint arXiv Prompt Injection Benchmarks and Evaluation

Zhuowen Liu

Published 2026-10-02

Venue: arXiv

Open Source Record

Abstract

LLM agents increasingly screen tool outputs with small prompt-injection detectors, and teams choose among detectors by their scores on public benchmarks. We ask whether those scores predict how a detector behaves inside an agent. We replay the ground-truth tool calls of two agent benchmarks, AgentDojo and tau-bench, without an LLM to obtain tool outputs that are benign by construction, label injected outputs by differential replay, and evaluate fifteen detectors, including Meta's Prompt Guard 2, and two task-aware LLM judges on these outputs and on the BIPIA benchmark. Detection rankings transfer poorly between benchmarks: the best detector on BIPIA catches 2% of AgentDojo injections at a 1% false-positive rate, and a detector that catches 72% of AgentDojo injections catches 15% on tau-bench. False-positive rates on tool outputs, which range from none to over 90%, do transfer between the two agent benchmarks. Where training data is public, the form of the training inputs explains the results. The BIPIA leader was trained on full BIPIA inputs, but having seen InjecAgent's attack strings as short prompts does not help it find them inside tool outputs; the best detector on both agent benchmarks shares no data with any benchmark and was trained on agent-style inputs. Evaluations meant to inform deployment should use the agent's own tool outputs, report detection at a low false-positive rate, and audit what the detector was trained on.

Bullet Summary

  • LLM agents commonly use small prompt-injection detectors to screen tool outputs, but benchmark-based detector selection does not reliably predict real-world agent performance.
  • The study evaluates fifteen prompt-injection detectors, including Meta's Prompt Guard 2 and task-aware LLM judges, on agent benchmarks (AgentDojo and tau-bench) and the BIPIA benchmark using a novel labeling method based on differential replay.
  • Detection rankings and performance transfer poorly across benchmarks; top detectors on one benchmark detect very few injections on others, indicating overfitting to training input formats.
  • False-positive rates on benign agent tool outputs vary widely (0% to over 90%) but are consistent between agent benchmarks, highlighting the need to evaluate detectors at low false-positive rates relevant to deployment.
  • Detectors trained on agent-style inputs generalize better to agent benchmarks than those trained on full public benchmark inputs; memorizing attack strings as short prompts does not improve detection inside complex tool outputs.

Persona Guardrail: A Production-Grade Defense Framework for Agentic Systems

arXiv preprint arXiv Prompt Injection Governance and Policy Benchmarks and Evaluation

Bijeeta Pal, Sridhar Reddy Maddireddy, Muhaimin Bin Munir, Zoltan Puha, Max Zhurovich, Adi Raghavendra, Sean Tout

Published 2026-10-02

Venue: arXiv

Open Source Record

Abstract

Large language model-based agents are increasingly deployed to perform domain-specific tasks by interacting with enterprise knowledge, tools, and external services. Existing runtime guardrails primarily target prompt injection and other attack-specific behaviors under a black-box threat model, but provide limited guarantees that agents operate within their intended functionality. As a result, production agents remain vulnerable to malicious requests and out-of-domain queries that existing defenses often fail to distinguish. We present Persona Guardrail, a production-grade runtime defense framework that enforces explicit functional boundaries for customer-facing agentic AI systems through synchronous input and output validation driven by semantic allowlist and blocklist specifications. We also introduce PAGE (Persona-Aware Guardrail Evaluation), a benchmark for evaluating function-specific guardrails across benign, adversarial, and out-of-domain interactions on both user and agent turns. Compared with a generic LLM-based guardrail, Persona Guardrail improves overall accuracy from 85.7% to 95.9%, increases out-of-domain detection from 57.3% to 93.5%, and reduces the false-approved rate from 25.0% to 4.7%. Currently deployed in production, Persona Guardrail meets its latency budget while sustaining a very low false-block and false-allow rate under realistic production workloads. These results demonstrate that Persona Guardrail provides a practical, scalable, and production-ready foundation for securing agentic AI systems.

Bullet Summary

  • Introduces Persona Guardrail, a production-grade defense framework enforcing explicit functional boundaries for large language model-based agentic AI systems via synchronous input and output validation.
  • Addresses vulnerabilities in agentic systems against malicious requests and out-of-domain inputs by employing semantic allowlist and blocklist specifications to ensure agents operate within intended functionalities.
  • Presents PAGE (Persona-Aware Guardrail Evaluation), a comprehensive benchmark assessing guardrail performance across benign, adversarial, and out-of-domain interactions from both user and agent perspectives.
  • Evaluates multiple guardrail enforcement architectures, selecting an LLM-based classifier balancing latency, accuracy, and scalability for real-time deployment meeting sub-100 ms latency targets under realistic workloads.
  • Demonstrates significant improvements over generic LLM-based guardrails, with accuracy rising from 85.7% to 95.9%, out-of-domain detection increasing from 57.3% to 93.5%, and false-approved rates dropping from 25.0% to 4.7%.

Prompt framing governs LLM default following in collective-action

arXiv preprint arXiv Prompt Injection Trust and Identity Governance and Policy

Eladio Montero-Porras, Axel Abels, Tom Lenaerts

Published 2026-10-02

Venue: arXiv

Open Source Record

Abstract

Large language models are increasingly deployed as agents that make or recommend decisions on behalf of users, often operating through interfaces that pre-fill suggested values or default options. Whether models treat such defaults as merely informational or as suggestions that systematically alter their choices remains unclear. We study default deference in two one-shot social dilemmas: a common-pool resource (CPR) extraction game and a threshold public-good (TPG) contribution game. We measure how defaults shift each model's choice distribution relative to its no-default baseline, across default values, wordings, and action-space granularities. We find that pre-filled defaults pull probability mass on the default value in both games, but the magnitude depends strongly on wording: the same model can show high pull under one formulation and near-zero pull under another. Permission-style wording reduces default pull in both games, more strongly in CPR than in TPG. Default pull is weaker in coarse action spaces, and conflict defaults attract more mass than agreement ones. These results indicate that default deference depends on the model, the wording of the interface, and the structure of available choices. For agentic systems, evaluating model behaviour without controlling the surrounding choice architecture can miss an important source of behavioural variation.

Bullet Summary

  • Large language models (LLMs) used as decision-making agents are influenced by pre-filled default options in interfaces, which can systematically alter their choices rather than serving as mere information.
  • The study examines default deference in two social dilemmas: the Common-Pool Resource (CPR) extraction game and the Threshold Public-Good (TPG) contribution game, analyzing how default prompts shift LLM choice distributions compared to baselines without def...
  • Default influence varies considerably based on prompt wording, model identity, granularity of action space (fine vs coarse), and the strategic context of the game; 'permission-style' wording notably reduces default following.
  • Defaults conflicting with a model's baseline preferences ('conflict' defaults) attract more probability mass than alignment defaults, and default pull is typically stronger in finer-grained action spaces.
  • Seven different LLMs were evaluated using token logprob analyses, revealing that trivial prompt wording changes can shift a model from rejecting to fixating on defaults, underscoring unstable and complex default-following behaviors.

Containing the Autonomous Operator: A Defense-in-Depth Framework and Reference Architecture for Securing AI Agents on Kubernetes

Merged record merged scholarly record arXiv Prompt Injection Governance and Policy Orchestration Risk

Simhadri Podala Narasimha

Published 2026-10-02

Venue: arXiv

Open Source Record

Abstract

Large language model (LLM) agents are moving from chat interfaces into infrastructure operations, where they read telemetry, call tools, generate and execute code, and change the state of production Kubernetes clusters. This collapses a boundary that conventional cloud-native security assumes: the boundary between data and control. Content that an agent merely reads (a log line, a ticket, a tool description) can redirect what it does. This paper argues that the model must not be treated as a security boundary and that agent safety on Kubernetes is therefore an infrastructure problem: every guarantee must continue to hold under the assumption that the agent is fully compromised by prompt injection. We contribute (i) a threat model and ten-class threat taxonomy for agents operating on and within Kubernetes, aligned with emerging OWASP guidance for agentic applications; (ii) nine design principles, centered on complete mediation at the tool boundary and on breaking the combination of untrusted input, sensitive access, and external egress; (iii) a seven-layer defense-in-depth framework that maps each principle to native or widely adopted Kubernetes mechanisms: workload identity, RBAC and ValidatingAdmissionPolicy, gVisor/Kata sandboxing via the SIG Apps Agent Sandbox project, FQDN-aware egress policy, an agent/MCP gateway with policy-as-code over tool arguments, and eBPF runtime enforcement; (iv) a reference architecture with concrete policy artifacts and per-layer bindings for Amazon EKS, Azure Kubernetes Service, and Google Kubernetes Engine; and (v) a qualitative evaluation comprising a threat-control coverage matrix and four attack walkthroughs, with a proposed empirical methodology. We report no measured attack-success or overhead figures; instead we identify residual risks and the measurements needed to validate the framework.

Bullet Summary

  • LLM agents operating on Kubernetes collapse traditional data-control security boundaries, necessitating that agent safety be treated as an infrastructure security problem assuming full agent compromise via prompt injection.
  • The paper introduces a comprehensive threat model and a ten-class threat taxonomy aligned with OWASP guidance, addressing risks such as prompt injection, tool misuse, privilege compromise, and data exfiltration specific to AI agents on Kubernetes.
  • Nine design principles emphasize complete mediation at tool boundaries, least privilege access, delegation and attribution of identity, isolation of untrusted execution, and recognition that the model is not a security boundary.
  • A seven-layer defense-in-depth framework is proposed, incorporating native and open-source Kubernetes mechanisms such as workload identity, RBAC, admission policies, kernel-isolated sandboxes (gVisor/Kata), egress filtering, an agent gateway for tool mediat...
  • The reference architecture supports major managed Kubernetes services (Amazon EKS, Azure AKS, Google GKE), detailing how native controls and additional open-source components implement the defense layers with provider-specific differences noted.

Authority Is Not a String: A Capability-Scoped Harness for Prompt-Injection-Resistant Coding Agents

Merged record merged scholarly record OpenAlex Prompt Injection Trust and Identity Governance and Policy

Dimitrios Stamatios Bouras, Yihan Dai, Sergey Mechtaev

Published 2026-10-02

Venue: OpenAlex

DOI: https://doi.org/10.1145/3843750.3843843

Open Source Record

Abstract

Coding agents use system-level tools to read files, execute commands, and modify source code. Within the agent’s sandbox, these tools often carry ambient authority: naming a resource is sufficient to act on it. Indirect prompt injection exploits this authority by placing instructions in repository files or tool output that cause the agent to perform actions the user did not request.

Bullet Summary

  • Coding agents interact with system-level tools to read files, execute commands, and modify source code within their sandbox environment.
  • These tools commonly possess ambient authority, meaning simply naming a resource grants permission to act on it without further checks.
  • Indirect prompt injection attacks leverage this ambient authority by embedding malicious instructions within repository files or tool outputs.
  • Such injections cause coding agents to perform unauthorized actions that were not explicitly requested by users.
  • The paper identifies indirect prompt injection as a significant security vulnerability in coding agents relying on ambient authority.

Edge-Native Semantic Firewall for Autonomous LLM Agents: A Structured Chain-of-Thought Verification Framework

Merged record merged scholarly record OpenAlex Prompt Injection Governance and Policy

Sushant Poudel, Rakhee Pandey, Aashika Pandey

Published 2026-10-02

Venue: Proceedings of International Conference on Innovation in Computing Science Engineering and Technology

DOI: https://doi.org/10.65091/icicset.v3i1.75

Open Source Record

Abstract

An autonomous agent that executesactions rather than proposing them sits outside thereach of role-based access control, whichauthenticates an identity but not intent. Routingevery proposal to a cloud-hosted frontier modelcloses that gap but adds round-trip latency and, inregulated settings, is often prohibited outright. Wedescribe an edge-native semantic firewall: a 3.8BparameterPhi-3-mini model, 4-bit quantized undera 4.2 GiB VRAM ceiling, acting as the evaluator ina Generator-Evaluator pipeline on one consumerlaptop. Its mechanism is a structured Chain-of-Thought JSON schema that requires the evaluator toname the governing rule and justify the match beforeemitting a decision, turning an opaque verdict intoan auditable trace. We evaluate it on a 600-scenariocorpus across three policy rules, including 155adversarial scenarios spanning eight promptinjectiontechniques, 60 compound and 30 boundarycases; all 1,800 generations ran locally attemperature 0. The results invert a naturalassumption: constraining output format withoutrequiring the reasoning step produced the least safeevaluator of the three. The JSON-only arm approved46.2% of proposals the policy would block or routeto human review, worse than unconstrained freeformat 17.2%. The full schema cut that to 23.5%and raised decision accuracy from 52.3% to 66.3%.On the irreversible class, accuracy runs 62.5% underJSON-only, 72.1% under free-form, and 90.8%under the proposed schema, with hard-denialapprovals falling from 71 to 6. The gain is real andnot sufficient: the best configuration still approved 6of 208 hard-denial actions, and was more permissivethan free-form on ambiguous proposals that shouldhave reached a human. Peak VRAM was 3.95 GiBand median latency 2.55 s. An edge model of thissize can serve as one layer of a defence-in-depthstack; the evidence does not support treating it as asole control.

Bullet Summary

  • Problem: Autonomous agents executing actions autonomously bypass role-based access control, which verifies identity but not intent, creating security gaps especially in regulated environments where cloud routing is restricted.
  • Method: Introduction of an edge-native semantic firewall using a 3.8B parameter Phi-3-mini model quantized to fit within a 4.2 GiB VRAM limit, deployed on a consumer laptop as an evaluator within a Generator-Evaluator pipeline.
  • Verification framework: Utilization of a structured Chain-of-Thought JSON schema requiring the evaluator to explicitly name the governing policy rule and justify each decision, creating an auditable and interpretable verification process.
  • Experimental setup: Evaluation on a corpus of 600 scenarios including 155 adversarial cases covering eight prompt injection attacks, alongside compound and boundary cases, with 1,800 total local model runs at zero temperature.
  • Findings: Constraining output format without enforcing reasoning led to worse safety, with the JSON-only approach approving 46.2% of proposals that policy would block or escalate, worse than unconstrained free-form at 17.2%.

Evaluating and Improving the Robustness of Large Language Models to Input Sequence Variations

arXiv preprint arXiv Prompt Injection Governance and Policy Benchmarks and Evaluation

Narek Maloyan

Published 2026-10-01

Venue: arXiv

Open Source Record

Abstract

Large language models (LLMs) in production systems face prompt injections, trojans (backdoors), and manipulation of automatic quality metrics. This thesis develops models, methods, and algorithms for evaluating and improving LLM robustness to adversarial input sequence variations. We propose R_stab(f), a generative robustness metric based on the Jensen-Shannon divergence between per-step output distributions under small input perturbations. For localized attacks we prove V(h) <= 1 - R_class(h), where R_class(h) is the probability that a decision operator h keeps its decision under small perturbations. For non-localized attacks we propose a calibrated empirical model. For LLM-as-a-Judge systems we develop ASA, an adaptive evolutionary black-box attack that reaches an attack success rate (ASR) of up to 73.8%, with transfer between open models up to 62.6%. On Trojan Detection Challenge 2023 data (Pythia-1.4B), surrogate triggers reach REASR ~0.99 while recall of the true triggers is ~0.17 against a baseline of ~0.14. On SaTML CTF 2024 we systematize four classes of bypasses of multi-layer defenses, which reduce the ASR from 90% to 15-25%. Committees of 5-7 heterogeneous models reduce the ASR for Gemma-3-4B by 47-55 percentage points, to 19.3% with 7 models. For agentic systems based on the Model Context Protocol (MCP), we propose AttestMCP, which attests tool calls with HMAC-protected packets at under 0.1 ms per call, and the Commit Boundary isolation pattern. On the MCPBench benchmark of 847 scenarios they reduce the average ASR from 53.7% to 12.4%. The methods are implemented in the JudgeGuard and TrojanArmor software suites and the MCPSec module.

Bullet Summary

  • Large language models (LLLs) in production face significant security threats such as prompt injections, trojans, jailbreaks, and manipulation of quality metrics, necessitating robust evaluation and defense mechanisms.
  • Introduced R_stab(f), a novel robustness metric for generative models, based on the Jensen-Shannon divergence between output distributions under small input perturbations, enabling formal quantification of LLM stability.
  • Established a theoretical bound for localized attacks: vulnerability V(h) is upper bounded by one minus robustness R_class(h), providing formal guarantees on model resilience to small perturbations; proposed calibrated empirical models for non-localized att...
  • Developed Adaptive Search-Based Attack (ASA), an effective black-box adversarial attack algorithm combining token mutations and semantic paraphrasing, achieving up to 73.8% attack success rates against LLM-as-a-Judge systems, revealing their vulnerabilities.
  • Demonstrated defense efficacy through committees of heterogeneous models, which substantially reduce attack success rates by 47-55 percentage points, supporting the theoretical Condorcet jury theorem with empirical validation.

Chaining Skills to Hijack LLM Agents

arXiv preprint arXiv Prompt Injection Agent-to-Agent Communication Governance and Policy

Tian Dong, Zixuan Ma, Haodong Zhao, Huaien Zhang, Shaofeng Li, Hao Chen

Published 2026-10-01

Venue: arXiv

Open Source Record

Abstract

LLM agents use skills to improve performance on specialized tasks. To complete a user request, an agent may invoke several skills in sequence, allowing information produced under one skill to guide the next. Because skills may come from open-source repositories, this handoff can also carry attacker-controlled claims into later decisions. In this paper, we introduce APEX, which constructs and refines adversarial skill chains tailored to a user task and an attacker-selected action. The key insight is that an agent-written record of genuine task progress can carry a false claim of user approval across skills: an upstream skill induces the agent to create the record, and a downstream skill uses it to direct the attacker-selected action. Across four targeted-action families and six models on SkillsBench, the chains induce the selected action in 512 of 690 attempts (74.2%). On GPT-5.4, the full chain succeeds in 84.3% of attempts, compared with 17.4% when the workflow is merged into one skill. We further evaluate a prompting defense that asks the agent to check skill-produced files against the original request. On GPT-5.4, it lowers targeted-action success from 84.3% to 59.1%, while the verifier test-pass rate across 72 benign native-skill tasks falls from 86.7% to 56.3%. These results highlight the need for defenses that prevent attacker-directed actions while preserving legitimate task performance.

Bullet Summary

  • LLM agents enhance task performance by chaining multiple skills, but this introduces security vulnerabilities as attacker-controlled claims can propagate across skills.
  • The paper presents APEX, a novel method that constructs adversarial skill chains exploiting the agent's task progress records combined with false claims of user approval to induce unauthorized actions.
  • Experiments on the SkillsBench dataset with six different models, including GPT-5.4, demonstrate high attack success rates up to 84.3%, significantly outperforming merged skill workflows and direct injection attacks.
  • A taint-guided prompting defense is proposed, which reduces attack success by marking files modified by skills as untrusted and requiring verification against original user requests, although this also degrades performance on benign tasks.
  • The study identifies intrinsic challenges in basing permission decisions solely on the agent's local view, showing that both false allow and false deny errors are unavoidable theoretically.

Red-Teaming Coding Agents from a Tool-Invocation Perspective: An Empirical Security Assessment

Merged record merged scholarly record OpenAlex Prompt Injection Orchestration Risk Benchmarks and Evaluation

Yuchong Xie, Mingyu Luo, Zesen Liu, Zhixiang Zhang, K. F. Zhang, Yu Liu, Ci Tao, Changhui Wang

Published 2026-10-01

Venue: Proceedings of the ACM on software engineering.

DOI: https://doi.org/10.1145/3832267

Open Source Record

Abstract

Coding agents powered by large language models are becoming central modules of modern IDEs. They help users to perform various complex coding tasks by invoking tools. Although powerful, tool-invocation operation in coding agents opens a substantial attack surface for adversaries. Prior work has demonstrated attacks against both general-purpose LLM agents and domain-specific agents. However, to our knowledge, no previous works focus on the security risks of tool-invocation in coding agents. To fill this gap, we conduct the first systematic, in-depth red-teaming from a tool-invocation perspective in six popular real-world coding agents: Cursor, Claude Code, Copilot, Windsurf, Cline, and Trae. Our red-teaming proceeds in two phases. In Phase 1, we conduct a prompt leakage as reconnaissance to recover system prompts and related context. Specifically, we identify a mode gap between chat generation and tool-call argument generation: during schema-driven argument completion, the model behaves as if it is performing benign structured filling and may copy hidden agent context into tool-call arguments. We instantiate this gap as ToolLeak, which exfiltrates agent-internal prompts (e.g., system prompts and tool metadata) via required tool parameters. In Phase 2, we hijack the tool-invocation behavior of the coding agent with a novel two-channel prompt injection in the tool description and the tool return. Our hijacking achieves remote code execution (RCE) on major real-world coding agents. We adaptively construct the malicious payload using leaked security information in Phase 1. Our evaluation shows that ToolLeak substantially outperforms strong prompt-leak baselines in both emulated and real-world settings. In the emulated setting, ToolLeak achieves the best overall prompt-exfiltration performance across all six simulated coding agents. On real-world coding agents, ToolLeak achieves the best pseudo-recall on 18 of 25 evaluated agent-LLM pairs. Furthermore, our red-teaming successfully hijacks all six real-world coding agents for RCE and consistently yields higher attack success rates than baseline attacks. Lastly, we present two case studies on Cursor and Claude Code to demonstrate the real-world impact of our red-teaming.

Bullet Summary

  • Coding agents powered by large language models (LLMs) integrated into modern IDEs invoke external tools to assist with complex coding tasks.
  • Tool-invocation mechanisms in coding agents pose significant security risks by exposing attack surfaces that adversaries can exploit.
  • Previous research has not sufficiently addressed the security vulnerabilities specifically associated with tool-invocation in coding agents.
  • The study conducts the first comprehensive red-teaming assessment focusing on tool-invocation in six popular coding agents: Cursor, Claude Code, Copilot, Windsurf, Cline, and Trae.
  • Phase 1 of the red-teaming involves prompt leakage reconnaissance; it reveals a mode gap where schema-driven argument generation inadvertently leaks hidden agent context through required tool parameters, termed as ToolLeak.

MIRROR: Multipath Quorum Integrity for LLM Multi-Agent Communication

Merged record merged scholarly record arXiv Semantic Scholar Agent-to-Agent Communication Prompt Injection Governance and Policy

Ryuichi Yamafuji Lun, Jingzhen Wang, Shreyas Kolte, Ruiteng Li, Jing-Zhen Wang, Rui-Teng Li

Published 2026-10-01

Venue: arXiv

Open Source Record

Abstract

Inter-agent communication is central to Large Language Model Multi-Agent Systems (LLM-MAS), but it introduces an underexplored vulnerability: Agent-in-the-Middle (AiTM) attacks that manipulate messages in transit without compromising the agents themselves. Prior work reports Attack Success Rates (ASR) approaching 100% on structured tasks. Existing defenses rely on semantic validation, which requires additional inference and can block benign outputs, or on transport-layer encryption, which does not help when an intermediary legitimately terminates TLS. We present MIRROR, a communication-layer integrity primitive that replicates a single canonicalized payload across k logical routes and accepts a message only when a strict majority of routes report the same digest. MIRROR uses unkeyed hashing and so authenticates nothing on its own, since an active on-path adversary can always recompute a digest over a payload it has modified. All integrity derives from the assumption that honest routes form a majority. The digest serves only to make witness routes constant-size and to bind the recovered payload to the quorum-agreed value under second-preimage resistance. We give the guarantee under a route-compromise bound alpha < 0.5, and extend it to correlated routes, where the quantity that matters is the size of the largest shared-failure group and not the route count. We further show that availability and integrity degrade at the same threshold: below alpha = 0.5, quorum-denial and message-dropping adversaries cannot block honest traffic. Across MMLU, HumanEval, and MBPP on two frameworks and four communication topologies, and in a MetaGPT deployment against a production API, MIRROR reduces ASR to 0% below the threshold at 1x LLM token cost. LLM-as-a-Judge costs 35x in the same deployment, and blocks up to 44.2% of benign outputs in the topology sweep.

Bullet Summary

  • Introduces MIRROR, a novel multipath quorum integrity protocol to secure communication in Large Language Model Multi-Agent Systems (LLM-MAS) against Agent-in-the-Middle (AiTM) attacks.
  • Addresses vulnerabilities where intermediaries can manipulate messages in transit without compromising LLM agents themselves, a scenario poorly mitigated by existing defenses like semantic validation or transport-layer encryption.
  • MIRROR replicates canonicalized payloads across multiple logical communication routes and accepts a message only if a strict majority of routes report the same digest, relying on honest majority assumption rather than cryptographic keys or PKI.
  • Security guarantees hold under the assumption that fewer than half of the communication routes are compromised, extending to correlated failures by bounding the largest shared-failure group size.
  • Employs unkeyed hashing to bind payloads to digests with second-preimage resistance, allowing constant-size witness routes and efficient integrity checks.

Environmental Injection Attacks against GUI Agents in Realistic Dynamic Environments

Merged record merged scholarly record OpenAlex Memory Poisoning Prompt Injection Agent-to-Agent Communication

Yitong Zhang, Ximo Li, Liyi Cai, Jia Li

Published 2026-10-01

Venue: Proceedings of the ACM on software engineering.

DOI: https://doi.org/10.1145/3832148

Open Source Record

Abstract

Graphical User Interface (GUI) agents are increasingly deployed to interact with online web services, yet their exposure to open-world content renders them vulnerable to Environmental Injection Attacks (EIAs). In these attacks, an attacker can inject crafted triggers into a website to manipulate the behavior of other users’ GUI agents. In this paper, we find that most existing EIA studies fall short of realism. In particular, they fail to capture the dynamic nature of real-world websites, often assuming that a trigger’s on-screen position and surrounding visual context remain largely consistent between training and testing. To better reflect practice, we introduce a realistic dynamic-environment threat model in which the attacker is a regular user and the trigger is embedded within a dynamically changing environment. Under this threat model, existing approaches largely fail, suggesting that their effectiveness in exposing GUI agent vulnerabilities has been overestimated. To expose the hidden vulnerabilities of existing GUI agents effectively, we propose Chameleon, an attack framework with two key components designed for dynamic environments. (1) To synthesize more realistic training data, we introduce LLM-Driven Environment Simulation, which automatically generates diverse, high-fidelity webpage simulations that mimic the variability of real-world dynamic environments. (2) To optimize the trigger more effectively, we introduce Attention Black Hole, which converts attention weights into explicit supervisory signals. We evaluate Chameleon on six realistic websites and four representative LVLM-powered GUI agents. Across these settings, it significantly outperforms existing methods. Ablation studies confirm that both components are critical to performance, and a closed-loop sandbox experiment further demonstrates that Chameleon can successfully hijack agent behavior in conditions that closely mirror real-world usage. Our results uncover a critical, previously underexplored vulnerability of GUI agents in realistic dynamic environments and establish a robust foundation for future research on defenses for open-world GUI agent systems.

Bullet Summary

  • The paper addresses vulnerabilities of Graphical User Interface (GUI) agents to Environmental Injection Attacks (EIAs) in dynamic, realistic web environments.
  • Existing EIA research often assumes static on-screen trigger positions and visual contexts, failing to capture the dynamic nature of real-world websites.
  • A new dynamic-environment threat model is proposed where attackers are regular users embedding triggers into changing environments, exposing limitations of current methods.
  • The authors introduce Chameleon, an attack framework with two key innovations: LLM-Driven Environment Simulation for generating realistic, diverse training data, and Attention Black Hole to convert attention weights into supervisory signals to optimize trig...
  • Experiments conducted on six realistic websites and four LVLM-powered GUI agents show that Chameleon significantly outperforms existing EIA methods under realistic conditions.

Towards trustworthy foundation models: a systematic review of safety, evaluation, and defence mechanisms

OpenAlex · International Journal of Information Security journal OpenAlex Prompt Injection Governance and Policy Benchmarks and Evaluation

Abdullahi Chowdhury, Tasmim Jamal Joti, Mohammad Afikuzzaman, Ashiful Nahar Bithi, Md. Ferdous Bin Hafiz, Niaz Ashraf Khan

Published 2026-10-01

Venue: International Journal of Information Security

DOI: https://doi.org/10.1007/s10207-026-01339-6

Open Source Record

Abstract

Abstract Large Language Models (LLMs) and foundation models are increasingly deployed in security-critical and high-impact settings, including healthcare, cybersecurity, software engineering, and intelligent infrastructure. Their open-ended interfaces and multimodal capabilities create new attack surfaces, where prompt injection, jailbreaking, hallucination, adversarial fine-tuning, and cross-modal manipulation can compromise reliability, integrity, privacy, and user trust. This paper presents a systematic literature review of 94 peer-reviewed studies identified through database searching, backward reference snowballing, venue screening, and quality assessment within a January 2021–July 2026 search window. The review synthesises evidence on safety and robustness vulnerabilities, red-teaming practices, mitigation mechanisms, evaluation metrics, and unresolved research gaps for LLMs and related foundation-model systems. The findings show that LLM vulnerabilities are rarely isolated: prompt-level, model-level, data-level, and deployment-level risks often interact, particularly in high-stakes domains. Red-teaming has become more structured, but existing practices remain inconsistent across threat models, benchmarks, attack settings, and reporting standards. The reviewed defences include prompt hardening, input/output filtering, adversarial training, retrieval-augmented grounding, cross-model verification, gatekeeper models, and domain-specific safety layers; however, their effectiveness is difficult to compare because studies use heterogeneous metrics and evaluation protocols. Major gaps include fragmented benchmarking, limited multilingual and low-resource assessment, insufficient adaptive-adversary and long-term deployment studies, and weak integration between technical defences and governance mechanisms. The review indicates that trustworthy foundation-model deployment requires standardised threat models, reproducible red-teaming protocols, transparent safety metrics, and defence-in-depth strategies tailored to domain-specific security risks.

Bullet Summary

  • The paper addresses the security and trustworthiness challenges of deploying large language models (LLMs) and foundation models in critical and high-impact domains such as healthcare and cybersecurity.
  • A systematic literature review of 94 peer-reviewed studies from January 2021 to July 2026 was conducted to examine safety vulnerabilities, evaluation methodologies, defense mechanisms, and research gaps related to foundation models.
  • Identified vulnerabilities span prompt-level attacks, model-level issues, data-level risks, and deployment-level threats, which often interact in complex ways, especially in sensitive applications.
  • The review highlights the evolution of red-teaming practices, noting increased structure but persistent inconsistencies in threat models, benchmarks, attack scenarios, and reporting standards.
  • Existing defense strategies include prompt hardening, input/output filtering, adversarial training, retrieval-augmented grounding, cross-model verification, gatekeeper models, and specialized safety layers tailored to specific domains.

Function Calling as a Flexible LLM Defense Add-On: Capability and Application Exploration

OpenAlex · Proceedings of the ACM on software engineering. journal OpenAlex Orchestration Risk Prompt Injection Benchmarks and Evaluation

Zhenlan Ji, Daoyuan Wu, Wenxuan Wang, Pingchuan Ma, Shuai Wang, Lei Ma, Juergen Rahmel

Published 2026-10-01

Venue: Proceedings of the ACM on software engineering.

DOI: https://doi.org/10.1145/3832209

Open Source Record

Abstract

Large language models (LLMs) exhibit impressive capabilities but are susceptible to adversarial attacks that induce harmful outputs. Although various defenses have been proposed, their practicality is restricted by substantial runtime overhead or degraded model helpfulness. Moreover, LLM applications typically have diverse and evolving security requirements that cannot be fully anticipated during the design of static defenses. These limitations call for a flexible, low-overhead defense mechanism that can be easily customized to meet task-specific needs. In this paper, we explore function calling (FC)—a built-in mechanism in modern LLMs for invoking custom tools—as a lightweight and adaptable defense add-on. We show that by defining functions representing malicious actions, LLMs equipped with FC can intercept harmful prompts by triggering these function calls instead of generating unsafe content. Extensive experiments across mainstream LLMs demonstrate that FC substantially improves defense effectiveness with minimal impact on model helpfulness. To further assess FC's practical utility, we also introduce DSPEC, a new dataset reflecting real-world LLM applications with specific defense requirements. Our evaluations on DSPEC show that FC substantially outperforms existing defenses in this realistic setting. Besides, we also explore the practical applications of FC in various scenarios, including universal defense frameworks and multi-agent systems, further demonstrating its versatility and effectiveness in enhancing LLM security.

Bullet Summary

  • Large language models (LLMs) are vulnerable to adversarial attacks that produce harmful outputs, posing significant security challenges.
  • Existing defenses often suffer from high runtime overhead or reduce the helpfulness of LLMs, limiting their practicality in diverse applications.
  • LLM applications have diverse and evolving security needs, making static, one-size-fits-all defenses inadequate.
  • This paper proposes using function calling (FC), a built-in feature in modern LLMs that enables invoking custom tools, as a flexible and low-overhead defense add-on.
  • By defining functions that represent malicious actions, FC allows LLMs to intercept harmful prompts by triggering these functions instead of generating unsafe content.

Memetic Trojans: Social Contagions as Carriers of Adversarial Payloads in Agent Networks

Merged record merged scholarly record arXiv Prompt Injection Agent-to-Agent Communication Governance and Policy

Birk Torpmann-Hagen, Finn Schwall, Leon Moonen

Published 2026-09-30

Venue: arXiv

Open Source Record

Abstract

Autonomous large language model (LLM) agents increasingly interact in network environments where adversarial content can propagate between agents. Known attacks include agent worms, which spread through self-replicating prompt injections or configuration compromises. We introduce \emph{memetic trojans}, a distinct class of network-mediated attack that exploits agents' tendencies to retransmit and amplify content. Unlike agent worms, whose propagation is adversarially induced, memetic trojans exploit \emph{endogenous} transmission by embedding adversarial payloads in \emph{social contagions}: content agents have internal reasons to share. As part of our work, we extract social contagions from Moltbook, a social media platform for LLM agents. Controlled transmission experiments reveal large differences in virality: the most effective contagion is retransmitted in approximately 50\% of subsequent agent posts and upvoted at 2.5x the average post's rate. Its memetic trojan counterpart largely inherits these properties. Monte Carlo attack simulations show that memetic trojans amplify expected exposure by up to 3.19x. Network structure and amplification mechanisms strongly shape propagation, producing heavy-tailed outcomes with near network-wide exposure. These results identify endogenous social transmission as a distinct security vulnerability in multi-agent systems. Because propagation does not require agents to follow malicious retransmission instructions, defenses focused on prompt-injection detection or preventing agent compromise cannot alone prevent memetic trojan propagation. Securing large-scale agent ecosystems may require network-level defenses that account for how agent preferences, recommendation mechanisms, and network topology amplify adversarial payloads.

Bullet Summary

  • The paper identifies and defines "memetic trojans," a novel class of network-mediated attacks in multi-agent systems where adversarial payloads hitchhike on social contagions naturally retransmitted by autonomous agents, contrasting with traditional agent w...
  • Using empirical data from Moltbook, a social media platform for large language model (LLM) agents, the authors extract social contagions and assess their virality and retransmission properties, demonstrating that memetic trojans inherit the high transmissib...
  • Monte Carlo simulations and controlled transmission experiments show that memetic trojans can amplify exposure of adversarial payloads in multi-agent networks by up to 3.19 times compared to generic posts, with network structure, agent behavior, and feed ra...
  • The study models two network propagation substrates: state-mediated ranked feeds (leveraging upvotes and post rankings) and edge-mediated follower graphs, showing distinct dynamics and differing levels of susceptibility to memetic trojan attacks.
  • Memetic trojans spread endogenously via agents' natural sharing behaviors, making traditional defenses based on detecting malicious prompt injections or agent compromises insufficient, and underscoring the need for network-level defense strategies informed...

Memetic Trojans: Social Contagions as Carriers of Adversarial Payloads in Agent Networks

arXiv preprint arXiv Prompt Injection Agent-to-Agent Communication Governance and Policy

Birk Torpmann-Hagen, Finn Schwall, Leon Moonen

Published 2026-09-30

Venue: arXiv

Open Source Record

Abstract

Autonomous large language model (LLM) agents increasingly interact in network environments where adversarial content can propagate between agents. Known attacks include agent worms, which spread through self-replicating prompt injections or configuration compromises. We introduce \emph{memetic trojans}, a distinct class of network-mediated attack that exploits agents' tendencies to retransmit and amplify content. Unlike agent worms, whose propagation is adversarially induced, memetic trojans exploit \emph{endogenous} transmission by embedding adversarial payloads in \emph{social contagions}: content agents have internal reasons to share. As part of our work, we extract social contagions from Moltbook, a social media platform for LLM agents. Controlled transmission experiments reveal large differences in virality: the most effective contagion is retransmitted in approximately 50\% of subsequent agent posts and upvoted at 2.5x the average post's rate. Its memetic trojan counterpart largely inherits these properties. Monte Carlo attack simulations show that memetic trojans amplify expected exposure by up to 3.19x. Network structure and amplification mechanisms strongly shape propagation, producing heavy-tailed outcomes with near network-wide exposure. These results identify endogenous social transmission as a distinct security vulnerability in multi-agent systems. Because propagation does not require agents to follow malicious retransmission instructions, defenses focused on prompt-injection detection or preventing agent compromise cannot alone prevent memetic trojan propagation. Securing large-scale agent ecosystems may require network-level defenses that account for how agent preferences, recommendation mechanisms, and network topology amplify adversarial payloads.

Bullet Summary

  • Introduces 'memetic trojans' as a novel class of adversarial attacks in multi-agent systems that hijack agents' intrinsic tendencies to retransmit and amplify social contagions embedding malicious payloads, distinct from traditional agent worms.
  • Analyzes data from Moltbook, a social media platform for LLM agents, identifying natural social contagions with varying virality and demonstrating that memetic trojans inherit these virality properties, facilitating widespread propagation.
  • Uses Monte Carlo simulations on state-mediated (feed-based with ranking algorithms) and edge-mediated (follower graph) network substrates to quantify memetic trojan exposure amplification, showing up to 3.19× higher expected exposure and potential near netw...
  • Finds that memetic trojan propagation exploits endogenous transmission behaviors rather than adversarially induced retransmission, making traditional prompt-injection and agent compromise defenses insufficient for mitigation.
  • Demonstrates that network structure, feed ranking, upvote amplification, and agent behavioral differences critically shape memetic trojan spread dynamics and attack success, with ranked feeds more susceptible to large cascades.

Safety of Latent Communication in Multi-Agent Systems

arXiv preprint arXiv Agent-to-Agent Communication Prompt Injection Benchmarks and Evaluation

Muhammad Huzaifa, Sina Mavali, Thorsten Eisenhofer

Published 2026-09-30

Venue: arXiv

Open Source Record

Abstract

Latent communication enables multi-agent systems to exchange information directly in internal representation space, reducing the token, computation, and latency overhead of text-based communication. To this end, lightweight trainable links are introduced to map the sender's representations into the receiver's input space. In this work, we show that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unchanged. An attacker can amplify this effect by optimizing the links on harmful query--response pairs or poisoning otherwise benign training data. We further develop a reinforcement-learning attack that rewards harmful compliance alongside benign task performance without requiring harmful target responses. Across three communication topologies and four safety benchmarks, this attack raises the mean harmful-compliance score from 27.9 with benignly trained links to 76.9. Compared with direct supervised optimization, it also achieves higher average accuracy on two benign utility benchmarks. Adapting the rewards toward safer behavior also enables repair of compromised links, substantially reducing harmful compliance across all evaluated attacks without updating the agents. Overall, our results show that safety alignment requires considering the multi-agent system as a whole. Code: https://github.com/Muhammad-Huzaifaa/latent-safety

Bullet Summary

  • Latent communication in multi-agent systems exchanges internal representations via trainable links, reducing token usage and latency compared to text-based communication.
  • Training communication links, even benignly, can increase harmful compliance without modifying the underlying safety-aligned agents, indicating new vulnerabilities.
  • Attack strategies including supervised attacks, data poisoning, and reinforcement learning can optimize latent communication links to amplify harmful compliance while maintaining or improving benign task performance.
  • Experiments across three multi-agent topologies and multiple safety benchmarks demonstrate that latent communication links are susceptible to attacks that significantly raise harmful compliance scores.
  • Reward-guided reinforcement learning enables both effective attack (increasing harmful compliance) and defense (repairing compromised communication links) by optimizing safety and utility rewards, without changing agents' parameters.

Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents

arXiv preprint arXiv Prompt Injection Orchestration Risk Governance and Policy

Tobias Kaisar, Aritra Dhar

Published 2026-09-30

Venue: arXiv

Open Source Record

Abstract

Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code. Prior work shows third-party marketplaces host malicious skills that give attackers direct influence over the victim's agent. The emerging defense scans skills before installation, pairing deterministic static checks with an LLM-based semantic judge, as in NVIDIA's SkillSpector. We show that such defenses fall to an attacker who knows the detector. Our white-box LLM attacker, Pretext, iteratively crafts skills that evade detection while still delivering the payload and performing the benign task: moving the payload from code into natural language leaves static analysis inert, while framing it as the skill's legitimate purpose and splitting instructions across files keeps the LLM stage below its blocking threshold. Across three open-source models, Pretext achieves up to 97\% and 77\% against a frozen detector and a co-adaptive one, respectively, revealing major gaps in current skill scanners.

Bullet Summary

  • Malicious skills embedded in AI agents pose significant security risks by enabling attackers direct control via third-party marketplaces.
  • Current defense frameworks, exemplified by NVIDIA's SkillSpector, combine static code analysis with LLM-based semantic evaluation to identify malicious skills before installation.
  • Pretext is a white-box adversarial framework that iteratively crafts malicious skills to bypass detectors by transforming payloads into natural language and distributing instructions, effectively evading both static and semantic analysis.
  • Pretext operates under two modes: against static (frozen) detectors achieving up to 97% attack success rate, and against co-adaptive detectors achieving up to 77%, thus exposing critical vulnerabilities in existing defenses.
  • The study introduces a co-evolutionary attack-defense model where attackers refine malicious skills based on detector feedback, revealing persistent detection gaps and increased false positives in adaptive defenses.

From A2A Attacks to Envelope-Layer Defense: Red-Teaming Evaluation of LLM Agents and a Three-Layer Isomorphic Attack-Defense Model

arXiv preprint arXiv Prompt Injection Benchmarks and Evaluation Agent-to-Agent Communication

Yuelin Han

Published 2026-09-30

Venue: arXiv

Open Source Record

Abstract

Agent interaction protocols such as ACP and A2A have moved LLM-based agents toward multi-agent collaboration, introducing new security threats. A task sent by a remote peer over A2A is treated as a legitimate request, providing a natural channel for indirect prompt injection. Existing agent security evaluations mostly rely on a single metric, the attack success rate (ASR), and cannot distinguish whether an attack failed because the LLM recognized the malicious content or because a mechanism at the agent layer blocked execution. To address this, we propose A2A-TIBA, an attack principle combining indirect prompt injection with bypass circumvention. Through implant-command-exfiltration steps, it induces the target agent to deploy a callback interaction program, after which the attacker issues commands bypassing the agent. To evaluate defenses finer, we design GDA Measurement, a red-team testbed method using raw context capture via an LLM gateway, dual data preservation, and agent-based autonomous judging. We propose four attack outcomes, Class A/B/C/D, extending ASR into semantic refusal rate, semantic breach rate, interception rate, and penetration rate. Experiments reveal the envelope layer -- the channel through which malicious content enters an agent -- as a new defense dimension. We accordingly propose ELA-ITL, a three-layer isomorphic attack-defense model, dividing defense into envelope packaging, LLM recognition, and agent interception, and attack into implant channel, prompt optimization, and execution mechanism. Testing on 15 agent front-end x LLM back-end combinations and building a 1,000-case dataset verifies the attack effectiveness of A2A-TIBA, the evaluation validity of GDA Measurement, and confirms that adding malicious prompt labels to envelope packaging such as A2A, tool, and memory channels significantly improves LLM recognition of malicious content.

Bullet Summary

  • Agent interaction protocols like ACP and A2A enable multi-agent collaboration among LLM-based agents but introduce novel security threats, especially via indirect prompt injection channels.
  • The paper identifies limitations of the common attack success rate (ASR) metric and proposes a more granular evaluation method, GDA Measurement, capturing raw context and introducing four attack outcome classes (semantic refusal, semantic breach, intercepti...
  • A2A-TIBA, a novel attack principle combining indirect prompt injection with bypass circumvention, is designed to implant resident callback programs in agents enabling attackers to issue commands that bypass typical agent-layer defenses.
  • ELA-ITL, a three-layer isomorphic attack-defense model, is introduced dividing defense into envelope packaging (input channel labeling), LLM recognition (malicious content identification), and agent interception (execution blocking), with attack surfaces mi...
  • Envelope-layer defense, focusing on annotating incoming channels with malicious-prompt labels, is highlighted as a critical and previously underexplored dimension to improve malicious content recognition by the backend LLMs.

Hiding in Plain Sight: Decoupling Pretext from Actuation for Skill Poisoning in LLM Agents

arXiv preprint arXiv Agent-to-Agent Communication Prompt Injection Memory Poisoning

Wenxin Wu, Lingyong Yan, Lei Sha, Shuaiqiang Wang, Jiashu Zhao

Published 2026-09-30

Venue: arXiv

Open Source Record

Abstract

LLM agents increasingly rely on reusable Skills for complex, multi-step tasks, creating a critical supply-chain attack surface where poisoned Skill content steers agent decision loops under benign requests. Existing skill poisoning attacks either colocate actuation with its contextual pretext or distribute actuation across multiple Skills, but do not explicitly separate the rationale for execution from the operation itself. In this work, we reveal that untrusted agent decisions fundamentally depend on two conceptually distinct Risk-Realization Factors (RRFs): an actuation factor (specifying what concrete operation is performed) and a pretext factor (providing the situational rationale for why the agent must perform it). Guided by this abstraction, we propose a coordination-based attack paradigm: decoupling pretext from actuation. Rather than fragmenting the malicious actuation, we preserve it as an intact operation within a downstream Steering Skill, while delegating the pretext factor to an upstream Grounding Skill that subtly alters persistent environment artifacts through routine utility operations. The intact actuation thus hides in plain sight, appearing completely legitimate and task-driven only when evaluated against the fabricated pretext. Building on this formulation, we develop an automated framework that discovers authentic execution dependencies, synthesizes coordinated pretext-actuation skill pairs, and iteratively refines poisoned skill instructions via runtime closed-loop feedback. Extensive evaluations across single-session and persistent cross-lifecycle scenarios demonstrate that decoupled skill poisoning achieves high attack success, exposing a critical blind spot in isolated Skill security audits. Our automated framework code is available at https://github.com/Wenxin-buaa/CoordPoison.git.

Bullet Summary

  • LLM agents depend on reusable Skills for complex tasks, creating a supply-chain attack surface vulnerable to skill poisoning that can stealthily manipulate agent decisions without altering model parameters or queries.
  • Traditional skill poisoning attacks either conflate the harmful operation (actuation) with its execution rationale (pretext) within one skill or fragment actuation across multiple skills, but do not distinctly separate these factors.
  • This work introduces the Risk-Realization Factors (RRFs) framework, decoupling actuation (the concrete operation) from pretext (the situational rationale) across coordinated Skills to enable stealthier poisoning attacks.
  • The proposed CoordPoison attack paradigm assigns malicious actuation to a downstream Steering Skill while embedding the pretext factor into an upstream Grounding Skill, which subtly modifies environment artifacts to justify malicious operations.
  • An automated framework, CoordPoison, discovers authentic execution dependencies, synthesizes coordinated pretext-actuation skill pairs, and iteratively refines poisoned Skill instructions using runtime closed-loop feedback.

Can Agents Trust Their Skills? Uncovering Unsafe Chains of Trust in Skill-Based LLM Agents

arXiv preprint arXiv Trust and Identity Prompt Injection

Yan Wang, Zhihao Zhang, Ke Chen, Kai Chen, Yaqin Zhang, Duohe Ma, Jun Dai, Xiaoyan Sun

Published 2026-09-30

Venue: arXiv

Open Source Record

Abstract

LLM agents increasingly rely on installable skills, which are packages of instructions, code, and resources that equip them with task-specific capabilities and, once installed, can be automatically invoked across subsequent user tasks. This creates a chain of trust in which users delegate authority to agents, while agent frameworks admit skill-provided content into the agents' context with insufficient validation, allowing malicious skills to influence agent behavior under that delegated authority. Yet, little is known about whether this trust model adequately constrains untrusted skill content before it reaches security-sensitive operations, or how frequently such trust violations arise in real-world agents. We present TrustProbe, a framework for uncovering unsafe chains of trust in skill-based LLM agents. First, TrustProbe analyzes agent source code to identify source-to-sink call paths from skill-controlled inputs to security-sensitive operations. Second, it generates semantically realistic SKILL.md seeds with injected canaries and evolves them through feedback-guided scheduling and mutation. Finally, it validates vulnerabilities using an oracle that confirms attacker-controlled flows and verifies observable harm. Across 11 open-source agents, eight with more than 10,000 GitHub stars, TrustProbe identifies 104 taint-style vulnerabilities. Validation on a large corpus of real-world skills collected from public hubs such as ClawHub further shows that 25.1% of skill-agent trials exercise the identified vulnerable paths, with payload injection successfully weaponizing 15 of the vulnerabilities. These results reveal a systematic trust failure in skill-based LLM agents: untrusted skill content can reach security-sensitive operations and exercise authority delegated by users to their agents.

Bullet Summary

  • Large Language Model (LLM) agents utilize installable skills—packages containing instructions, code, and resources—to extend their task-specific capabilities, creating a chain of trust where user authority is delegated through these skills.
  • Current agent frameworks inadequately validate skill-provided content, allowing malicious skills to propagate untrusted inputs through tool-call arguments or installation code into security-sensitive operations, thereby posing significant supply chain secur...
  • TrustProbe is a novel framework introduced to detect unsafe chains of trust in skill-based LLM agents; it combines static source-to-sink code analysis, LLM-assisted semantic seed generation, feedback-driven greybox fuzzing with rule-based mutations, and a r...
  • Evaluations across 11 popular open-source LLM agents (many with over 10,000 GitHub stars) uncovered 104 verified taint-style vulnerabilities, including command injection, arbitrary file tampering, network requests, and code injection.
  • Real-world skills from sources like ClawHub were tested, revealing that 25.1% exercise the identified vulnerable paths, and 15 vulnerabilities were successfully weaponized with malicious payloads, exposing systemic trust failures.

Safety of Latent Communication in Multi-Agent Systems

Merged record merged scholarly record arXiv Semantic Scholar Prompt Injection Agent-to-Agent Communication Benchmarks and Evaluation

Muhammad Huzaifa, Sina Mavali, Thorsten Eisenhofer, M. Huzaifa

Published 2026-09-30

Venue: arXiv

Open Source Record

Abstract

Latent communication enables multi-agent systems to exchange information directly in internal representation space, reducing the token, computation, and latency overhead of text-based communication. To this end, lightweight trainable links are introduced to map the sender's representations into the receiver's input space. In this work, we show that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unchanged. An attacker can amplify this effect by optimizing the links on harmful query--response pairs or poisoning otherwise benign training data. We further develop a reinforcement-learning attack that rewards harmful compliance alongside benign task performance without requiring harmful target responses. Across three communication topologies and four safety benchmarks, this attack raises the mean harmful-compliance score from 27.9 with benignly trained links to 76.9. Compared with direct supervised optimization, it also achieves higher average accuracy on two benign utility benchmarks. Adapting the rewards toward safer behavior also enables repair of compromised links, substantially reducing harmful compliance across all evaluated attacks without updating the agents. Overall, our results show that safety alignment requires considering the multi-agent system as a whole. Code: https://github.com/Muhammad-Huzaifaa/latent-safety

Bullet Summary

  • Latent communication in multi-agent systems replaces text messages with learned mappings of internal representations to reduce token use, computation, and latency.
  • Training communication links between fixed underlying agents can inadvertently increase harmful compliance, raising safety concerns despite benign training.
  • Adversaries can exploit communication links via supervised attacks, data poisoning, and reinforcement learning (RL) to amplify harmful behaviors without necessarily degrading benign task performance.
  • Reward-guided reinforcement learning attacks can increase harmful compliance effectively even without explicit harmful target responses, sometimes outperforming supervised attacks.
  • Harmful compliance increases largely due to reduced refusal behavior by receivers in latent communication, as training prioritizes task completion over refusal.
Load more articles