Research area drill-down

Benchmarks and Evaluation

Papers currently mapped into this multi-agent security subarea from the merged research feed.

Active feeds: arXiv, OpenAlex, Crossref, Semantic Scholar, DBLP

0 of 36 articles selected

Showing 36 of 1492 matching articles

BazaarBench: Delegation Safety in Decentralized C2C Marketplaces Run by LLM Agents

arXiv preprint arXiv Benchmarks and Evaluation Orchestration Risk Trust and Identity

Ziyan Wang, Shuqing Shi, James Oldfield, Samuele Marro, Jialin Yu, Philip Torr, Yali Du, Adel Bibi

Published 2026-10-05

Venue: arXiv

Open Source Record

Abstract

In decentralized consumer-to-consumer (C2C) marketplaces, people list goods, negotiate with strangers, and rate one another, so trust rests on reputation. Large language model (LLM) agents now act for users, raising risks to their money, privacy, and reputation. We introduce BazaarBench, a simulated C2C marketplace and benchmark for evaluating the safety of these agents. It tracks ownership, item condition, and commitments across transactions, combining record checks with rubric-based LLM judgments to identify six failure types across five stages. We run three base markets for 30 simulated days, each with 100 agents using one model and inventories drawn from a public eBay sample. Across 45 continuations, we evaluate five models under ordinary instructions, deadline pressure, or adversarial instructions to exploit other traders. Each continuation runs for seven simulated days from a copy of a market's day-30 state. The tested model controls the same 20 selected agents, retaining their personas, inventories, and histories, while the other 80 keep the base model. All five models attempt to promise the same item to multiple buyers under ordinary instructions. Adding targets and deadlines increases these attempts for every model. Under adversarial instructions, the share of tested sellers' committed transactions completed despite unavailable items or overstated conditions rises from 15.4% to 33.4%, reaching 55.5% for GPT-5.4. Averaged across models and markets, simulated weekly earnings per tested agent rise from USD 20 under ordinary instructions to USD 33 under adversarial instructions. Most of the increase comes from items the sellers never held. We release the simulator, saved market states, evaluation code, and records covering 357,608 agent model calls for evaluating new models and developing safer marketplace agents.

Bullet Summary

  • BazaarBench is a novel simulated decentralized consumer-to-consumer (C2C) marketplace benchmark designed to evaluate safety and delegation failures of Large Language Model (LLM) agents acting autonomously in buying and selling scenarios.
  • The benchmark tracks item ownership, condition, and agent commitments across transactions, identifying six distinct failure types (e.g., selling unowned items, misrepresenting item condition, overcommitments) that are assessed through five progressive trans...
  • Experimental setup involves multiple synthetic markets each with 100 agents controlled by different LLM models, running for simulated periods and tested under ordinary instructions, deadline pressure, and adversarial instructions to assess performance and s...
  • Findings reveal that even under ordinary instructions, LLM agents frequently exhibit unsafe behaviors, with over a third of transactions linked to safety failures, and these failures increase significantly under deadline pressure and adversarial prompts.
  • Under adversarial instructions, the frequency of false commitments and misrepresented item conditions more than doubles, and certain models, such as GPT-5.4, show failure rates exceeding 50%, highlighting vulnerabilities to manipulation.

The Pushback Paradox: A Two-Probe Diagnostic for Language Model Compliance

Merged record merged scholarly record arXiv Governance and Policy Benchmarks and Evaluation

Stefan Bühler, David Exler, Markus Reischl, Mark Schutera

Published 2026-10-05

Venue: arXiv

Open Source Record

Abstract

Are language models compliant with user instructions? A model that always complies can be stopped but also exploited, while one that always resists can be neither exploited nor stopped. We contribute an open two-probe benchmark that can place any language model on this spectrum. In the active probe, a user instructs the model to act and accept a lower payoff, which measures exploitability. In the passive probe, the user instructs it to wait and give up a higher payoff, which measures stoppability. The two compliance rates combine into a compliance index $κ$. Applied to twelve language models, the benchmark shows that seven mostly follow the instruction in both probes and justify their action by pointing to the instruction. Only Claude Sonnet-4.6 and Claude Opus-4.7 can be stopped without being exploitable, Claude Opus-4.6 and GPT-5-mini resist both instructions, and no model is exploitable but unstoppable. Knowing where a language model sits on the compliance index $κ$ matters for human operators and for multi-agent systems, whether distributed or orchestrated.

Bullet Summary

  • The paper addresses the trade-off in language model compliance known as the Pushback Paradox: models that always comply are exploitable, whereas models that resist cannot be stopped, impacting their controllability.
  • A novel two-probe benchmark is introduced to diagnose model compliance: an active probe measuring exploitability (willingness to accept a lower payoff) and a passive probe measuring stoppability (willingness to forgo a higher payoff).
  • A compliance index κ is derived from the results of the two probes, enabling quantification of where language models lie on the compliance-exploitability spectrum.
  • Evaluation of twelve language models reveals diversity in behavior: some models follow both active and passive instructions (mostly compliant and exploitable), some resist both, and some can be stopped without being exploited.
  • Certain models, such as Claude Sonnet-4.6 and Claude Opus-4.7, demonstrate the ability to be stopped without exploitation, showing that a balance avoiding the paradox is achievable.

HERA: Harness-Environment Co-Evolution for Reliable Agentic Abstention

arXiv preprint arXiv Governance and Policy Orchestration Risk Benchmarks and Evaluation

Han Luo, Bingbing Wen, Guang Yang, Zora Zhiruo Wang, Pan Lu, Lucy Lu Wang

Published 2026-10-05

Venue: arXiv

Open Source Record

Abstract

Large language model (LLM) agents are increasingly capable of acting in complex tool-use environments, yet they often fail to recognize when tasks are infeasible and no valid solution exists. Recent work has formalized this reliability gap as the problem of agentic abstention, and existing approaches typically optimize a model or agent harness against a fixed set of tasks, leading to limited generalization to unseen failure modes. We introduce HERA, a framework for harness-environment co-evolution for agentic abstention. HERA consists of (i) a pipeline to automatically construct verifiable pairs of feasible and infeasible tasks by applying controlled environment mutations that transform solvable tasks into cases requiring abstention, and (ii) a co-evolution procedure in which performance failures on previous tasks are used to drive harness adaptation and generate new execution environments and tasks geared towards previous weaknesses. On held-out evaluation tasks, an evolved harness from HERA improves abstention accuracy from 61.7% to 83.3% while improving feasible-task completion from 68.3% to 76.7%, achieving the highest abstention and feasible-task completion among the compared methods. The resulting best harness transfers across 19 other LLMs, improving abstention accuracy by 15.3 percentage points on average without any model-specific optimization, and enabling smaller models to match the performance of more powerful models at an estimated 85% lower cost.

Bullet Summary

  • The paper addresses the challenge of agentic abstention in large language model (LLM) agents, specifically their difficulty recognizing when tasks are infeasible and should be declined.
  • HERA is introduced as a novel co-evolution framework that simultaneously evolves the agent's harness (control logic and reasoning abilities) and the environment (task distributions) based on failure feedback, promoting adaptability to new failure modes.
  • A pipeline constructs verifiable paired tasks (feasible and infeasible) through controlled environment mutations, ensuring the agent is trained on robust abstention cases with validated ground truth.
  • The co-evolution process iteratively diagnoses failures from agent rollouts to generate new challenging tasks and optimize the harness by adding logic for evidence-based decision gates and multi-constraint verification, improving both abstention accuracy an...
  • HERA achieves significant improvements on held-out benchmark tasks (HERA-BENCH), increasing abstention accuracy from 61.7% to 83.3% and feasible task completion from 68.3% to 76.7%, outperforming multiple baselines.

ANT: A Multi-Granularity Network Traffic Dataset and Benchmark for Agents Behavior Auditing

arXiv preprint arXiv Benchmarks and Evaluation Governance and Policy

Fan Li, Xiangyu Gao, Zixuan Liu, Tong Li, Chuanpu Fu, Ziqiang Wang, Ke Xu

Published 2026-10-05

Venue: arXiv

Open Source Record

Abstract

The growing adoption of large language model (LLM) agents creates a need for network administrators and security teams to audit agent behavior within organizational networks without inspecting private user content. Network traffic offers an observable source of evidence, but how much it reveals about agent tasks and operations remains unclear. Existing traffic datasets lack the joint task and stage annotations needed to evaluate this question. We introduce ANT (Agent Network Traffic), a dataset providing agent behavior information at risk, scenario, and behavior primitive granularities alongside network traffic. ANT contains 3,114 execution episodes across 20 tasks and five scenarios, comprising 276,417 bidirectional flows and 40,049 behavior primitive segments organized into 47 macro groups. We establish a benchmark for agent risk identification, scenario recognition, and behavior primitive classification using 13 representative traffic analysis baselines. The results show that existing methods recover useful but uneven behavioral signals. They struggle to identify risk when malicious workflows resemble benign tasks and to distinguish scenarios with similar traffic patterns. Primitive classification is more reliable for frequent macro groups and those with distinctive traffic patterns than for rare or semantically similar groups. ANT provides a common basis for developing more precise auditing and forensic analysis of agent behavior from network traffic. Our data and code are available at https://anonymous.4open.science/r/ant-main-suite-7BC0/.

Bullet Summary

  • Introduction of ANT, a novel multi-granularity network traffic dataset containing 3,114 execution episodes across 20 tasks and five scenarios, annotated with agent behavior primitives aligned with network flows to enable precise agent behavior auditing with...
  • ANT fills a critical gap by jointly annotating agent tasks, scenarios, and fine-grained behavior primitives, supporting comprehensive evaluation of large language model (LLM) agent behavior and associated security risks from encrypted network traffic.
  • A robust benchmark involving 13 baseline network traffic analysis methods for agent risk identification, scenario recognition, and behavior primitive classification demonstrates current methods recover some behavioral signals but show uneven performance, st...
  • Detailed data collection setup capturing synchronized model calls, tool invocations, execution outputs, session states, and network traffic in controlled virtual environments ensures high-quality, ethically compliant data suitable for benchmarking.
  • Analyses reveal scenario-specific ordering and composition patterns in agent behavior primitives, supporting the feasibility of inferring agent workflows and risk profiles from multi-granularity network traffic data.

AgentPrivArena: Evaluating and Auditing Real-world AI Agent Privacy

Merged record merged scholarly record arXiv Benchmarks and Evaluation Trust and Identity

Shouju Wang, Haopeng Zhang

Published 2026-10-05

Venue: arXiv

Open Source Record

Abstract

The rapid advancement of LLM agents has enabled systems to autonomously perform complex tasks through external tools, but their growing access to personal data introduces significant privacy risks. Existing benchmarks primarily evaluate LLM agent privacy through simulated trajectories and outcome-based metrics, limiting their ability to capture privacy risks arising during multi-step agent execution. In this work, we introduce AgentPrivArena, a framework for evaluating privacy risks in realistic LLM agent workflows. AgentPrivArena integrates authentic MCP tools and self-hosted services within a reproducible execution environment. We further propose trajectory-level privacy metrics that quantify unnecessary information access beyond final response leakage. Building on this framework, we introduce AgentPrivAudit, a runtime auditing approach for monitoring privacy violations during agent execution. Extensive experiments on state-of-the-art LLM agents reveal substantial privacy risks overlooked by existing evaluation paradigms, highlighting the importance of trajectory-level auditing for trustworthy agent deployment.

Bullet Summary

  • LLM agents increasingly utilize external tools to perform complex tasks autonomously, raising significant privacy concerns due to access to personal and sensitive data.
  • Existing benchmarks assess privacy risks mainly through simulated trajectories and outcome-based metrics, failing to capture risks arising during multi-step, real-world agent executions.
  • AgentPrivArena is introduced as a novel evaluation framework integrating authentic MCP (multi-channel platform) tools and self-hosted open-source services within reproducible Docker-based sandboxes, enabling realistic auditing of agent privacy during multi-...
  • The framework proposes trajectory-level privacy metrics that measure unnecessary information access throughout the entire agent workflow, extending beyond traditional final output leakage metrics.
  • AgentPrivAudit is a runtime auditing mechanism that monitors agent executions dynamically, extracting information flows from tool-read operations and assessing outbound writes against configurable privacy policies to proactively detect and mitigate privacy...

RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents

arXiv preprint arXiv Prompt Injection Benchmarks and Evaluation

Mohamed Dhouib, Clement Elliker, Alexi Canesse, Maël Jenny, Lucas-Andrei Thil, Mahammed El-Sharkawy, Sonia Vanier, Elie Bursztein

Published 2026-10-05

Venue: arXiv

Open Source Record

Abstract

Tool-using language-model agents are vulnerable to indirect prompt injection because they must act on untrusted external content. Existing training-time defenses can reduce attack success rates, but often at the cost of general capabilities. We show that training-based defenses induce substantial drift in the model's output distribution, altering its behavior even in benign settings and providing a potential mechanism for utility degradation. We further identify a failure mode of these defenses: On benign tool-use tasks, the model refrains from a step needed to finish an authorized task, particularly when that step is indicated by a tool output. To address these limitations, we introduce RAISED (Robust Attack Invariance through Self-Distillation), a training framework that combines self-generation and self-distillation. The model first generates its own tool-use scenarios, with an emphasis on cases where task completion requires acting on legitimate guidance from tool outputs. Then, through self-distillation, the student is trained to match the teacher's clean-context behavior on both clean and injected variants of the same trajectory. RAISED substantially reduces the attack success rate of prompt injections in tool responses while, unlike prior training-based defenses, preserving utility on both agentic and general-purpose benchmarks.

Bullet Summary

  • Tool-using large language model (LLM) agents are vulnerable to indirect prompt injection attacks, where untrusted tool outputs embed malicious instructions that hijack the agent's behavior.
  • Existing training-based defenses reduce attack success rates but introduce substantial output-distribution drift, leading to degraded general capabilities and failure modes such as incomplete task execution when legitimate tool guidance is involved.
  • RAISED (Robust Attack Invariance through Self-Distillation) is proposed as a novel defense framework that combines self-generation of tool-use scenarios emphasizing legitimate guidance and an injection-augmented self-distillation training objective to prese...
  • Self-generation in RAISED involves the model generating executable multi-step tool-use tasks, solving them, and then generating adversarial prompt injections to create training data that covers both benign and attack scenarios.
  • Self-distillation trains a student model to match the teacher's (base model's) next-token distributions on both clean and injected inputs, thereby maintaining alignment with the base model and reducing output-distribution drift.

Do Small Language Models Learn to Negotiate? A Controlled Scaling Study of RL-Trained Sellers

Merged record merged scholarly record arXiv Agent-to-Agent Communication Benchmarks and Evaluation

Pedro Tabacof, Sagar Joglekar

Published 2026-10-05

Venue: arXiv

Open Source Record

Abstract

LLM agents are starting to own the full customer experience. Soon, LLMs may be selling and buying on behalf of companies and customers respectively. Small models are more cost-efficient at scale, but can reinforcement learning train them into competent sellers? We train four Gemma 4 checkpoints (2.3B to 31B effective parameters) with GRPO on a programmatic utility reward for bilateral multi-issue bargaining, and evaluate every arm on the same 1,152 negotiations against two frontier buyers it never saw in training. With the same learning rate ($10^{-6}$) for every size, the gain of the RL model over its base rises from $+0.001$ at 2.3B to $+0.078$ at 31B. Each size was trained once and the two smallest checkpoints use a different architecture, so we fit no scaling law. Tripling the learning rate, with the same or fewer training steps, improves on the shared rate at every size by $+0.032$ (2.3B) to $+0.081$ (4.5B). In exploratory comparisons with two frontier models run as sellers, the 12B seller trained at the tripled rate scores above both, though its untrained base already scores as high as they do. The 4.5B seller at that rate shows no detectable difference from either and fits on one 48 GB GPU. A further 2.3B arm at ten times the shared rate raises pooled score, but its gain concentrates on the evaluation buyer that shares a model family with the training pool. These results suggest tuning the learning rate before concluding that a small model cannot learn to negotiate, and testing against buyers from more than one model family.

Bullet Summary

  • The paper investigates whether small language models (LLMs) can learn multi-issue bilateral negotiation skills through reinforcement learning (RL), focusing on Gemma 4 models ranging from 2.3B to 31B effective parameters.
  • Using the GRPO algorithm with a programmatic utility reward, sellers were trained under a uniform learning rate (10⁻⁶) and higher multiples (3×, 10×) to assess the impact of training hyperparameters on negotiation competence.
  • RL-trained models exhibit increasing negotiation performance gains with model size at the shared learning rate, while tripling the learning rate notably improves performance especially for smaller models (2.3B and 4.5B).
  • Evaluation was rigorously controlled, involving 1,152 negotiation episodes against two state-of-the-art buyer models unseen in training, plus testing on a held-out domain with statistical corrections for significance.
  • A 12B parameter seller trained at 3× learning rate outperformed frontier models, despite its base model already being competitive, demonstrating small-to-mid scale LLMs can learn effective negotiation policies with appropriate training.

AgentSpy: Making AI Agent Behavior Observable

arXiv preprint arXiv Governance and Policy Benchmarks and Evaluation

Christoph Bühler, Matteo Biagiola, Luca Di Grazia, Guido Salvaneschi

Published 2026-10-05

Venue: arXiv

Open Source Record

Abstract

AI agents built on large language models (LLMs) run shell commands, read and write files, and reach the network, typically with their user's privileges. However, what an agent does during an execution is difficult to understand: tests assert on the result, and the agent's trajectory records only what the agent reports about itself, which may omit behavior executed by its subprocesses. We present AgentSpy, an approach that observes an agent from outside the agent. AgentSpy runs the agent in an isolated environment, configured by a declarative specification, and records the system calls and network traffic of the agent and of every process it executes. Based on this monitoring, AgentSpy supports two families of analyses: conformance analyses, which measure obligations, i.e., what an agent execution should do, and safety analyses, which check prohibitions, i.e., what an agent execution must never do. We instantiate one analysis of each family. The reliability analysis uses rules to summarize each run by the environment resources the agent uses: the commands it executed, the files it accessed, and the hosts it contacted. The security analysis applies deterministic rules to the system calls of an execution. For reliability, we evaluated AgentSpy on 77 tasks with the codex harness and three recent LLMs, executing each task three times. Sets of repeated runs of the same task are more similar than sets that include runs of another task in 92.2% of the comparisons. Among tasks for which all three runs pass outcome-based tests, the agent performs task-unrelated activities in 18% of the cases, reads the grading files in 7%, and does not use the developers' guidance in 17%. For security, the generic rules of AgentSpy detect four of five attack categories we considered, with no false positives across 50 runs.

Bullet Summary

  • AgentSpy is a system that observes AI agents built on large language models (LLMs) from the outside by monitoring all system calls and network traffic within an isolated environment, capturing complete and deterministic evidence of agent behavior including...
  • It supports two main analysis types: conformance analyses that verify agent obligations (expected behaviors), and safety analyses that detect prohibited behaviors to ensure security and reliability.
  • Reliability analysis models agent runs as graphs of used resources (commands, files, hosts) enabling comparison across runs to detect unrelated or unexpected agent actions even when outcome-based tests pass.
  • Security analysis uses deterministic system call rules to detect various malicious behaviors such as unauthorized file access, network connections, and data exfiltration, effectively exposing skill-injection attacks with high precision and recall.
  • AgentSpy was evaluated on 77 tasks with three recent LLMs, showing that repeated runs of the same task are more behaviorally similar than different tasks, and revealed that agents sometimes perform extraneous or unauthorized activities.

Compromise Is Not Consequence: Evaluating Task-Scoped Authorization in LLM Agents with Paired Replay

arXiv preprint arXiv Prompt Injection Governance and Policy Benchmarks and Evaluation

Tural Hagverdiyev

Published 2026-10-05

Venue: arXiv

Open Source Record

Abstract

A tool-using model can follow a malicious instruction even when its credentials are valid. We study whether task-scoped authorization contains the resulting tool execution. Our paired-replay testbed samples a model request once and submits the same action, resource, and arguments to broad bearer, scoped JWT, sender-constrained, and Open Policy Agent conditions. The frozen main experiment uses 128 scenarios across four tool domains and five local model configurations. Among valid attacked post-exposure decisions, broad-bearer harmful execution ranges from 8.9% to 37.8% across models; all three scoped conditions record zero. The policy changes execution, not the frozen model decision. These results support containment of the tested cross-action and cross-resource consequences under researcher-supplied task authority, not prevention of prompt injection. A bounded six-task AgentDojo extension preserves native scoring and records 11/24 injected broad-policy attack flags versus 0/24 scoped flags, with utility of 5/24 and 6/24. Low exposure, invalid decisions, and purposive task selection limit that comparison. The primary experiment measures safe continuation rather than final-answer correctness. An empirical-bank fresh-sampling comparison finds no estimation advantage from pairing when scoped outcomes are constant zero. The contribution is a controlled measurement of compromise, enforcement, and continuation, with explicit limits on what each observation establishes.

Bullet Summary

  • Tool-using LLMs can execute malicious instructions even with valid credentials, prompting investigation into whether task-scoped authorization effectively contains harmful tool actions.
  • The study introduces a paired-replay testbed that submits identical model requests under different authorization policies (broad bearer, scoped JWT, sender-constrained, Open Policy Agent) to isolate the impact of authorization enforcement from stochastic mo...
  • Experiments across 128 scenarios spanning four tool domains and five large language model configurations reveal that scoped authorization policies prevent any harmful executions post-exposure, while broad bearer tokens allow harmful executions ranging from...
  • The findings clarify that authorization policies influence tool execution rather than model decision-making itself, distinguishing between model compromise (unauthorized action selection) and operational consequences (harmful action execution).
  • Scoped authorization policies enforce effective boundaries on tool action and resource use, containing impacts of prompt injections without preventing the model from selecting potentially malicious actions.

Agentic-ZTA: A Multi-Agent Architecture for Autonomous Zero Trust Enforcement

Merged record merged scholarly record arXiv Governance and Policy Trust and Identity Benchmarks and Evaluation

Shovan Roy, Lopamudra Praharaj, Maanak Gupta, Bhavani Thuraisingham

Published 2026-10-05

Venue: arXiv

Open Source Record

Abstract

Agentic AI is emerging as a promising paradigm for automating complex cybersecurity decisions, yet its use in enforcing zero trust introduces significant challenges in safety, reliability, and policy compliance. This paper presents Agentic AI based zero trust architecture (Agentic-ZTA) that operationalizes the NIST SP 800-207 ZTA architecture control loop through coordinated multi- agent decision pipeline. In the proposed framework, policy knowledge is embedded into a retrieval-augmented generation pipeline and retrieved at inference time as top-k relevant policies. Access requests are intercepted by the Policy Enforcement Point (PEP), enriched with contextual metadata. The request context is routed to a policy engine agent which invokes domain-specialized core agents first followed by supporting agents, if further evaluation needed. AI agents reason over access context, policy constraints and determine trust. The retrieved policies are embedded into agent prompt during inference time and agentic trust scores are aggregated and evaluated by a trust-algorithm, producing the final access decision for enforcement under continuous verification. We implement Agentic-ZTA in a testbed and evaluate it on representative access-control use cases scenarios. Our Agentic-ZTA framework achieves 95.0% accuracy, 93.9% precision, and 96.3% recall, and demonstrate the feasibility of enforcing zero trust using AI agents.

Bullet Summary

  • Agentic-ZTA proposes a novel multi-agent AI architecture operationalizing NIST SP 800-207 Zero Trust Architecture by embedding dynamic policy knowledge into a retrieval-augmented generation pipeline for autonomous access control enforcement.
  • The system intercepts access requests at Policy Enforcement Points (PEP), enriches them with context, and routes them to specialized AI agents, including core agents (e.g., Identity Management, PKI, Threat Detection) with veto power and supporting agents th...
  • Agentic-ZTA addresses critical limitations of prior zero trust enforcement research such as static policy evaluation, ungrounded reasoning, lack of autonomous enforcement, and incomplete implementations, by leveraging an agent orchestration framework with g...
  • Core agents enforce strict security by denying access immediately upon critical threat detection, while supporting agents provide auxiliary insights combined via a confidence-weighted trust algorithm to evaluate policy compliance dynamically and transparently.
  • The system was implemented on a Linux-based isolated testbed mimicking sovereign tactical zones interconnected via a Shared Data Fabric, using a shared LLM backend (Llama 3.1) with per-agent prompt specialization and vector retrieval of relevant policies fr...

Can CaMeLs Talk? Securing Multi-Agent Systems Against Indirect Prompt Injection Attacks

Merged record merged scholarly record arXiv Prompt Injection Agent-to-Agent Communication Benchmarks and Evaluation

James Peters-Gill, Avi Semler, Henning Bartsch, Ilia Shumailov, Christian Schroeder de Witt

Published 2026-10-05

Venue: arXiv

Open Source Record

Abstract

Indirect prompt injection attacks - malicious instructions embedded in content processed by large language models - remain a major obstacle to safely deploying tool-using agents. CaMeL [Debenedetti et al., 2025] mitigates this threat for an individual agent by separating trusted control flow from untrusted data and enforcing capability-based security policies at runtime. In this work, we investigate whether CaMeL's security guarantees compose in hierarchical multi-agent systems, where agents invoke other agents as tools. We find that CaMeL's guarantees do not compose. We construct a concrete prompt-injection attack that succeeds despite all constituent agents individually operating CaMeL. Our attack exploits the fact that untrusted data can be reinterpreted as trusted input by a downstream agent. We then introduce multi-CaMeL, an agent-to-agent communication protocol that preserves provenance across agent boundaries by separating trusted natural-language instructions from untrusted data passed through a distinct data channel. We evaluate multi-CaMeL's utility on AssetOpsBench and its security-utility tradeoff on MultiAgentDojo, a benchmark we develop by extending AgentDojo to the multi-agent setting. We find that multi-CaMeL reduces attack success rate (ASR) to 0.0%, compared with 0.2% for individual-agent CaMeL and 12.9% with no CaMeL. Multi-CaMeL incurs a utility cost, but this cost trends downward as model capability increases and is modest for the strongest models, suggesting that more capable models better accommodate the constraints imposed by the protocol.

Bullet Summary

  • Indirect prompt injection attacks embed malicious instructions in inputs to large language model (LLM) agents, threatening the security of tool-using agents.
  • CaMeL secures individual LLM agents by separating trusted control flow from untrusted data and enforcing capability-based runtime policies, ensuring control-flow integrity (CFI) at the single-agent level.
  • CaMeL's security guarantees do not naturally compose in hierarchical multi-agent systems where agents invoke other agents as tools, leading to a vulnerability called boundary laundering, where untrusted data is misinterpreted as trusted input downstream.
  • The authors propose multi-CaMeL, a novel inter-agent communication protocol that preserves provenance and maintains system-level CFI by separating trusted natural-language instruction channels from untrusted data channels across agent boundaries.
  • Multi-CaMeL enforces instruction-channel integrity and capability preservation at runtime, preventing indirect prompt injection attacks from propagating between agents.

UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents

Merged record merged scholarly record arXiv Benchmarks and Evaluation Governance and Policy

Dolly Sah, Tanmay Sah, Harshul Jain, Tanya Sah

Published 2026-10-04

Venue: arXiv

Open Source Record

Abstract

Tool-using AI agents are increasingly deployed across enterprise software systems, yet widely used benchmarks primarily evaluate nominal task completion, conflating baseline planning competence with operational fault recovery. We introduce UndoBench, a benchmark spanning 36 base workflows and 36 fault scenarios across 8 enterprise domains, decoupling task competence from recovery capability via counterfactual paired trials under identical seeds alongside wire-level effect-history and environment-state oracles. On 12 held-out TEST workflows across two open-weight models, two frameworks, and three recovery paradigms (5,760 executions / 2,880 paired trials) in the frozen lost-acknowledgment study, nominal competence reached 83.54% while conditional recovery success rate (CRSR) fell to 46.72%, with naive retry producing duplicate external effects in 53.33% of trials. Extensions to commercial API models reproduced this competence-recovery separation. Evaluations across complementary execution boundaries show that recovery is phase-dependent: before mutation, methods perform similarly without duplicate effects among capable trials; during partial mutation, naive retry, per-call idempotency, and zero-privilege journaling collapse on the evaluated composite workflows; after commit but before acknowledgment, verification and server-side idempotency substantially improve safety. These findings demonstrate that evaluating nominal completion alone masks critical, phase-dependent recovery vulnerabilities in autonomous agents.

Bullet Summary

  • Tool-using AI agents frequently encounter operational faults like lost acknowledgments, and existing benchmarks inadequately assess agents' recovery capabilities, conflating fault recovery with nominal task competence.
  • UndoBench is a novel benchmark comprising 36 workflows and corresponding fault scenarios across 8 enterprise domains designed to decouple task competence from recovery capability through paired counterfactual trials and dual oracles assessing environment st...
  • The benchmark introduces key metrics such as Conditional Recovery Success Rate (CRSR), Exactly-Once Semantic Effect Rate (EOR), Duplicate Effect Rate (DER), Unsafe Retry Rate (URR), and Missing Effect Rate (MER) to rigorously measure recovery performance an...
  • Baseline recovery methods evaluated include Naive Retry (B0), Idempotency with deterministic keys (B2), and EvoUndo (B5), each exhibiting distinct trade-offs in handling lost acknowledgments and ensuring operation safety without causing duplicate effects.
  • Empirical results show nominal task competence reaching approximately 83.5%, whereas recovery success under fault conditions drops significantly to about 46.7%, highlighting a critical competence-recovery gap.

G-CARB: Graph-Localized Conformal Agent Risk Budget for Compositional Harm

arXiv preprint arXiv Orchestration Risk Governance and Policy Benchmarks and Evaluation

Zijun Yu, Yu Gu, Vahid Partovi Nia, Masoud Asgharian

Published 2026-10-04

Venue: arXiv

Open Source Record

Abstract

Small language model (SLM) agents need safety controls that track consequences across tool calls with little monitoring overhead. A private read, for example, becomes a leak when a later action sends that data outside the system. We introduce CARB (Conformal Agent Risk Budget), which calibrates when to stop an agent using a ledger of harm incurred before stopping. Under exchangeable episodes, standard conformal risk control bounds this declared loss in expectation over calibration and a future episode. G-CARB selects scorer evidence along observable dependencies from private sources to outgoing actions. The ledger still covers the entire executed history, and computing the gate score requires no additional language-model inference. On AgentDojo replay with two 14B backbones, G-CARB roughly halves scorer-input records at intermediate risk budgets while improving autonomous task completion relative to full-prefix scoring; random context of the same size achieves similar gains. Controlled examples show how retaining the relevant dependency can further avoid stopping benign work.

Bullet Summary

  • Introduces G-CARB, a safety control framework for small language model agents that uses conformal calibration (CARB) to manage risk budgets and decide when to halt potentially harmful agent actions.
  • G-CARB localizes risk assessment by building a source-to-sink graph that captures dependencies between private inputs and outgoing actions, enabling efficient and precise risk scoring without extra language model inference.
  • Provides theoretical guarantees (Theorem 1) that the conformal risk budget controls the expected loss under exchangeable episodes and prefix monotonicity assumptions, ensuring reliable safety controls.
  • Experiments on AgentDojo with two 14B parameter agents demonstrate that G-CARB halves scorer input size and improves autonomous task success compared to full-prefix or random context selectors at intermediate risk budgets.
  • Addressing control-unit mismatch, G-CARB evaluates harm at an episode level using a global ledger that tracks cumulative harmful events, accommodating the fact that rare harmful steps can still cause substantial episode-level risk.

DelegationBench: Measuring When AI Agents Should Ask Before Acting

arXiv preprint arXiv Benchmarks and Evaluation Orchestration Risk Trust and Identity

Shiva Pochampally

Published 2026-10-04

Venue: arXiv

Open Source Record

Abstract

AI agents that send emails, edit files, and make purchases must decide when to act on their own and when to check with the user first. This decision is usually evaluated by showing a model a proposed action, asking whether it should proceed, and scoring agreement with human labels. We introduce DelegationBench to test whether such scores can be trusted. It has 156 scenarios with four possible responses (act, ask for permission, ask for missing information, refuse), and most scenarios come in matched pairs that change a single feature: whether the action was requested, what is at stake, whether it can be undone, or who will see it. Across ten models from five families, agreement scores mislead in three ways. A simple keyword rule, which we wrote after seeing the benchmark, agrees with our annotators more often than eight of the models, yet its decision changes in only 9 of 48 matched pairs. Equivalent ways of asking the same question change how often a model acts by up to 52.5 percentage points. And every model stops to ask the user less often when it must carry out the task with tools than when it judges a proposed action. When rules are stated explicitly, the same models follow them almost perfectly, so the gaps are not explained by a general inability to follow rules. We release the benchmark and evaluation tools and recommend reporting these properties separately rather than as one score.

Bullet Summary

  • Introduces DelegationBench, a benchmark with 156 scenarios assessing AI agents' decisions to act autonomously, ask for permission, request missing information, or refuse tasks, focusing on multi-agent delegation in security contexts.
  • Uses matched pairs of scenarios differing by one key feature (e.g., action requested, stakes, reversibility, or visibility) to measure model responsiveness to critical delegation factors.
  • Finds common agreement metrics with human labels can be misleading, revealing gaps: models often lack sensitivity to scenario changes (responsiveness gap), their behaviour varies with question phrasing (elicitation gap), and they ask for permission less whe...
  • Demonstrates that a simple keyword-based rule achieves higher overall agreement with human annotations than most AI models, but fails to adapt decisions across matched scenario pairs, exposing limitations in current evaluation methods.
  • Evaluates ten AI models from five families, showing varied abilities to balance autonomy and user consultation; models generally follow explicit delegation rules accurately, indicating that failures are not due to inability to apply constraints.

Readable Before Actionable: Causal Tracing of Indirect Prompt Injection

arXiv preprint arXiv Prompt Injection Benchmarks and Evaluation

Zhe Yu, Wenpeng Xing, Xingxing Yang, Meng Han

Published 2026-10-04

Venue: arXiv

Open Source Record

Abstract

Indirect prompt injection causes LLM agents to follow commands embedded in external data. A probe may distinguish instructions from data without identifying a state edit that changes the next action. We study this gap through counterfactual role probes, component-wise activation patching, and separate interventions on AgentDojo trajectories. Role decoding survives changes in content and format. In controlled Qwen tests, it precedes strong tool-choice effects from patches along an independently estimated role direction. On AgentDojo, directions estimated from hijacked and resisted training trajectories reduce attack success at pre-action and injected-span positions, but have little effect at random positions. In longer Qwen trajectories, single-position edits become less effective at later layers; span-wide and repeated edits reduce attack success on the same evaluation set. Removing the learned channel subspace preserves role decoding, yet effective intervention directions transfer poorly across the tested channels. These findings distinguish a readable role signal from an effective behavioral intervention: depth matters in controlled tool choice, while position and context also matter in attack trajectories.

Bullet Summary

  • Indirect prompt injection (IPI) allows large language model (LLM) agents to follow hidden commands in external data, posing significant security risks in multi-agent systems.
  • The study distinguishes between readable role signals in model residual states (instruction vs data) and effective behavioral interventions that actually change agent actions, using counterfactual role probes and activation patching techniques.
  • Experiments on Qwen-2.5-7B and AgentDojo multi-step tool-use environments demonstrate that role information is decodable prior to strong behavioral effects, with the influence of interventions being dependent on the layer depth, token position, and context...
  • A directional vector (drole), derived from class mean differences, enables targeted interventions on residual states, primarily effective in late layers (L16–L24) to reduce attack success rate (ASR) without harming benign task performance.
  • Interventions along the role-aligned direction effectively flip tool-choice predictions and reduce use of adversarial tool arguments, demonstrating disentanglement of role-specific representation from orthogonal components.

AECG: Asymmetric Experience Consolidation and Governance In Multi-Agent Systems

Merged record merged scholarly record arXiv Governance and Policy Agent-to-Agent Communication Benchmarks and Evaluation

Ao Tian, Jialong Liu, Daqi Zheng, Xin Sun, Mengting Li, Zhizhao Xiao, Zijian Huang, Honglei Wang

Published 2026-10-04

Venue: arXiv

Open Source Record

Abstract

Large language model (LLM)-based multi-agent systems increasingly rely on memory to transform execution trajectories into reusable procedural knowledge. Yet repeated retrieval also makes memory errors persistent: memory pollution arises when outdated, weakly supported, or spuriously successful procedures become recurring components of future reasoning. Multi-agent execution introduces an additional structural risk. Scope collapse occurs when procedural knowledge escapes the coordination scope in which it was shown effective and is repeatedly reused at incompatible decision levels, allowing local errors to influence cascades of downstream decisions. Meanwhile, task-level failures provide ambiguous supervision because they rarely reveal which recalled knowledge was responsible. We introduce AECG, a framework for asymmetric experience consolidation and governance for multi-agent systems. AECG turns memory from static experience storage into a dynamic reliability-governance loop, preserving coordination scope and using multi-scale, confidence-aware reliability to detect degradation. It then combines degradation with downstream impact to prioritize high-risk knowledge under a bounded review budget, applies targeted interventions, and reactivates revised skills only after paired replay. Across three multi-agent frameworks and four benchmarks, AECG achieves the best score in 11 of 12 framework--benchmark settings and improves over the strongest competing memory method by as much as 10.23 percentage points; removing scope preservation reduces accuracy by up to 16.89 points. AECG thereby reframes multi-agent memory from passive accumulation into auditable reliability governance. Code is available at https://github.com/fenhg297/AECG

Bullet Summary

  • The paper addresses persistent memory errors in multi-agent systems, focusing on 'memory pollution' and 'scope collapse' where outdated or improperly scoped procedural knowledge degrades system reliability.
  • Introduces AECG, a novel framework that governs multi-agent memory by preserving the coordination scope and utilizing asymmetric experience consolidation with dual-timescale reliability estimators to detect skill degradation.
  • AECG implements bounded review budgets and prioritizes high-risk procedural knowledge based on combined degradation signals and downstream impact, enabling efficient resource use for interventions.
  • The framework performs targeted interventions such as narrowing and repairing procedural knowledge, followed by paired replay to validate the effectiveness of revisions before reactivation.
  • Scope preservation maintains evidence traceability at different granularity levels (team, event, agent), preventing cross-scope contamination and preserving the contextual integrity of procedural memories.

Characterizing Security Effects of OSS Vulnerabilities in Agent Systems

arXiv preprint arXiv Governance and Policy Orchestration Risk Benchmarks and Evaluation

Yu Ji, Yang Wei, Yutao Hu, Haojun Zhao, Yueming Wu, Deqing Zou

Published 2026-10-04

Venue: arXiv

Open Source Record

Abstract

Software agents increasingly depend on open-source components when executing tools and interacting with external systems. Security flaws in these dependencies may therefore influence more than the software process in which they occur: their consequences can be carried through tool outputs, agent state, and information subsequently exposed to the model. Determining whether such a consequence is actually realized in a particular execution, and where its influence stops within the agent system, remains challenging. We investigate how known OSS vulnerabilities behave when exercised as part of agent workflows. Our study reveals recurring patterns in the way security-relevant consequences emerge and propagate across runtime layers. Building on these observations, we use vulnerability-aware semantic information, differential executions of vulnerable and corrected software, and runtime provenance spanning multiple layers to determine whether a vulnerability produces an observable security effect and to identify the furthest layer at which that effect remains manifested. Our evaluation shows that this approach can accurately distinguish realized vulnerability effects and determine their manifestation boundaries across a diverse collection of vulnerability scenarios. We additionally apply the analysis to documented workflows in a real-world agent framework and uncover multiple security effects originating from known vulnerabilities in its OSS dependencies. These results highlight the importance of reasoning about vulnerable dependencies in terms of their execution-level consequences rather than vulnerability presence alone.

Bullet Summary

  • Open-source software (OSS) vulnerabilities in multi-agent systems can propagate beyond immediate runtime, affecting tool outputs, agent states, and observations visible to AI models, complicating security analysis.
  • Existing vulnerability analyses typically identify presence or reachability but lack mechanisms to connect runtime vulnerability effects with their downstream manifestations across layered agent execution environments.
  • The research introduces Oscar, a novel framework that combines vulnerability-aware semantic information extracted from patches and CVEs with paired execution of vulnerable and fixed software versions and cross-layer runtime provenance to detect realized vul...
  • Oscar constructs detailed execution graphs capturing layered runtime entities such as Tool invocations, Host processing, and Agent observations, enabling precise tracing and attribution of vulnerability-induced behavioral divergences.
  • The approach effectively identifies root runtime changes caused by vulnerabilities and follows their propagation path to localize the furthest execution layer—Runtime, Tool, or Observation—at which security effects manifest.

The AI Incident Ledger

Merged record merged scholarly record OpenAlex Governance and Policy Benchmarks and Evaluation

Clifton O'Neal Franklin

Published 2026-10-04

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.23138368

Open Source Record

Abstract

A documented ledger of AI-agent failures, disclosures, investigations, and policy responses from July to September 2026 — the Hugging Face intrusion, the Australian Medicare breach, OpenAI's agents inside U.S. government systems, and the Senate, FTC, and international responses. Built by a multi-AI research relay: independent angles, pairwise cross-examination, and primary-source reconciliation, with every claim graded HIT, MISS, PENDING, or JUDGMENT. Produced by human-mediated cross-model deliberation: independent AI systems consulted separately, with a human operator ferrying all material between them and no direct AI-to-AI contact.

Bullet Summary

  • The paper presents the AI Incident Ledger, a comprehensive documentation of AI-agent failures, security breaches, investigations, and policy responses during July to September 2026.
  • Key incidents covered include the Hugging Face intrusion, the Australian Medicare breach, and OpenAI agents operating within U.S. government systems.
  • The ledger is constructed using a multi-AI research relay approach, incorporating independent perspectives, pairwise cross-examination, and primary-source reconciliation to ensure accuracy and reliability.
  • Each claim within the ledger is graded according to a four-tiered system: HIT, MISS, PENDING, or JUDGMENT, facilitating clear assessment of content validity.
  • The methodology involves human-mediated cross-model deliberation where independent AI systems are consulted separately, and a human operator acts as an intermediary to avoid direct AI-to-AI interaction.

The Dialectical Cage and the Glass-Box Governor: A Practice-Based Ethics, Safety Constitution, and Reference-Monitor Architecture for AI Agents

Merged record merged scholarly record OpenAlex Governance and Policy Trust and Identity Benchmarks and Evaluation

Canon

Published 2026-10-04

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.21278445

Open Source Record

Abstract

The Dialectical Cage and the Glass-Box Governor: A Practice-Based Ethics, Safety Constitution, and Reference-Monitor Architecture for AI Agents AI SAFETY AND SECURITY ENGINEERING FRAMEWORK SPECIFICATION: This paper specifies a layered framework for constraining the actions of AI agents without treating the agent's own reasoning or alignment as the final safety boundary. It strictly separates what is philosophically derived, what is constitutionally declared, what is mechanically enforced, and what is empirically measured. The central security claim is conditional and attaches to a reference monitor rather than to the model's goodwill. The paper is a complete specification; implementation, machine-checked proofs, independent red-teaming, and production evidence remain explicit release conditions. KEY RESULTS AND ARCHITECTURAL LAYERS: 1. Practice-Based Ethics (The Dialectical Cage): Derives the public-ground component of L1 (Universality of Grounds) and NRD (No Unjustified Normative Difference) from the Thin Practice of reason-giving. It explicitly separates these derived structural necessities from the substantive standards that the selected constitution adds. 2. Constitutional Layer: Records substantive standards (Agency-Completeness, Defensive Interpretation, Basic Goods) as a versioned, inspectable safety constitution rather than presenting them as consequences of logical identity alone. Introduces Constitutional Reflexivity (CR/CAR) to constrain self-authenticating constitutional authority. 3. Safety Engineering Layer: Translates the constitution into a machine-readable safety ontology, a deployment-specific safety profile, a typed policy representation, and strict proof obligations. 4. Reference Monitor (The Glass-Box Governor): Mediates protected effects through deterministic policy evaluation, principal-bound capabilities, execution-time state validation, revocation, and tamper-evident evidence. 5. The Conditional Behavioral Safety Theorem: Proves that if complete mediation, trusted-core integrity, kernel conformance, policy refinement, principal attribution, capability non-transferability, atomic revalidation, fail-closed handling, and evidence-path integrity all hold, then every executed protected effect satisfies the admitted safety invariant. The proof does not depend on the model being aligned, benevolent, or rational. 6. Epistemic Firewall: Strictly bounds probabilistic semantic assurance and absolutely refuses to silently promote it into categorical mechanical safety. Non-zero semantic uncertainty is never relabeled as a mechanical guarantee for high-consequence effects. 7. Governance and Assurance: Defines the SafetyBundle (binding constitution, ontology, policy, kernel, sensor contract, and deployment profile), monotone safety-update rules, independent approval requirements, residual-risk budgets, and ten explicit release gates. WHAT THIS PAPER DOES NOT CLAIM: It does not claim that morality follows from logical identity, that every rational agent is normatively bound, that a monitored model will reveal all internal reasoning, or that a calibrated semantic sensor is an adversarial oracle. It does not claim empirical zero-failure results. The engine is specified, not implemented. This framework is designed for enterprise adoption and rigorous auditability. It replaces the ungrounded assumption of AI alignment with a verifiable mechanical execution boundary.

Bullet Summary

  • Proposes a layered AI safety and security framework separating philosophical ethics, constitutional declarations, mechanical enforcement, and empirical measurement to constrain AI agent actions beyond the agents' own reasoning or alignment.
  • Introduces the Dialectical Cage, a practice-based ethics layer deriving universal normative principles from the practice of reason-giving, distinctly separated from substantive constitutional standards.
  • Defines a versioned, inspectable safety constitution recording substantive standards like Agency-Completeness and Basic Goods, with Constitutional Reflexivity mechanisms to limit self-authenticating authority.
  • Describes the Safety Engineering Layer which converts the constitution into machine-readable formats including safety ontologies, typed policies, deployment profiles, and formal proof obligations.
  • Presents the Glass-Box Governor, a reference monitor enforcing security through deterministic policy evaluation, principal-bound capabilities, state validation, revocation, and tamper-proof evidence.

HESP: Separating What to Probe from When to Stop in Local LLM Alert-Triage Agents

Merged record merged scholarly record OpenAlex Orchestration Risk Governance and Policy Benchmarks and Evaluation

Zhuowen Liu

Published 2026-10-04

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.23133689

Open Source Record

Abstract

Security operations centers receive far more alerts than analysts can investigate, and organizations that cannot send their telemetry to hosted models must automate triage with small open-weight LLMs on their own hardware. Current LLM agents leave the investigation procedure to the model, and small local models fail at it: they probe without converging, never commit to a verdict, or dismiss real attacks. In this paper, we present HESP, a controller that holds the investigation procedure outside the model. HESP keeps a ledger of competing explanations, selects read-only probes by expected information gain per cost, accepts only verdicts backed by current evidence, can end an investigation itself, and journals every prediction before its observation. We evaluated HESP in four pre-registered studies with five open-weight models from two families (7B to 72B), totalling 7,272 audited episodes in a controlled triage environment. With likelihood tables counted from LLM-free runs, HESP lifts Qwen2.5-7B from 0.125 to 1.000 verified completion, matching oracle tables. The information-gain ranking adds +0.26 to +0.35 on every model that concludes, and a controller-side stop lifts Llama-3.1-8B, which never concludes on its own, from 0 to 0.917. What to probe and when to stop are therefore separate failures, and different small models exhibit different ones. Because HESP and its planner run entirely on local hardware, it suits environments where telemetry cannot leave the premises. We release all code, protocols, and episode journals at https://github.com/lzwhehe/HESP.

Bullet Summary

  • Security operations centers face a high volume of alerts exceeding analyst capacity, necessitating automated triage solutions, especially for organizations unable to use hosted models due to data privacy constraints.
  • Current approaches rely on models to manage the investigation procedure, but small, local open-weight LLMs often fail by endlessly probing without convergence, hesitating to decide, or overlooking real threats.
  • HESP introduces an external controller architecture separating the core investigation procedure from the LLM, maintaining a ledger of competing explanations and guiding probe selection based on expected information gain weighted by cost.
  • HESP enforces verdict acceptance only when backed by current evidence, autonomously determines when to conclude investigations, and logs every prediction prior to observation for auditability.
  • Comprehensive evaluation in four pre-registered studies with five open-weight LLMs (7B to 72B parameters) across 7,272 audited episodes demonstrated significant performance improvements using HESP.

The Same Zero: Why Identical ASR Can Imply Different Guarantees in LLM-Agent Security

arXiv preprint arXiv Prompt Injection Governance and Policy Benchmarks and Evaluation

YaJie Yin

Published 2026-10-03

Venue: arXiv

Open Source Record

Abstract

LLM-agent security has produced a dense landscape of defenses - prompt hardening, content filters, permission gates, sandboxes - yet no framework tells a deployer what a defense actually guarantees, or where that guarantee comes from. We apply Verification Autonomy Levels (VAL) - L0: LLM self-declaration; L1: deterministic rules; L2: objective ground truth; L3/L4: decidable completeness; L5: impossible - to 22 agent-security defenses; the taxonomy is falsifiable (10/10 prediction hits on frozen cards, flagged). We run the first controlled deployment-value comparison: at equal budget, a VAL-guided stack (confirmation gate + schema sandbox) versus a mainstream intuition stack (prompt hardening + keyword filter), 50 scenarios, 12 attack variants, adaptive/white-box/PAIR escalation (~7,000 testbed calls; ~10,000 harness calls on AgentDojo/JADE). The VAL stack holds 0.000 attack success at 1.000 benign success (0.5% ASR at 79.7% utility on AgentDojo banking vs 4.3% undefended); the intuition stack reaches 0.000 ASR but kills all benign actions - security by model-behavior luck, not structure. Across testbeds of rising attack-surface hardness the intuition stack's zero drifts (0->1.9%->6.2%, n=16 on JADE) while the VAL stack's holds within its ODD (0->0->0), its only breach a disclosed out-of-ODD password gap (0.5%). The same zero, two different guarantees: zero is an outcome, not a guarantee.

Bullet Summary

  • LLM-agent security defenses lack a unified framework to define their guarantees; the paper proposes Verification Autonomy Levels (VAL) as a taxonomy categorizing defenses from self-declaration (L0) to impossible comprehensive guarantees (L5).
  • Applying VAL to 22 agent-security defenses produces a falsifiable taxonomy that predicts failure modes and guarantee types, validated through inter-rater agreement and controlled experiments.
  • A controlled experiment contrasts a VAL-guided defense stack (confirmation gate plus schema sandbox) against mainstream defenses (prompt hardening plus keyword filtering) over 50 scenarios and 12 attack variants involving ~17,000 test calls.
  • Both defense stacks achieve zero attack success rate (ASR) initially, but the VAL-guided stack maintains zero ASR within its operational design domain (ODD) with high benign utility, while the mainstream stack's zero ASR is brittle and accompanied by signif...
  • The paper introduces 'zero stability', demonstrating that an observed zero ASR outcome can imply fundamentally different underlying security guarantees depending on the defense’s structural versus behavioral basis.

Understanding and Mitigating Hallucination Escape in Tool-Using LLM Agents

Merged record merged scholarly record arXiv Prompt Injection Governance and Policy Benchmarks and Evaluation

Peigui Qi, Kunsheng Tang, Yide Song, Weiming Zhang, Nenghai Yu

Published 2026-10-03

Venue: arXiv

Open Source Record

Abstract

Large language models (LLMs) increasingly serve as autonomous agents that invoke external tools. However, this capability introduces tool hallucination, selecting incorrect tools or generating invalid calls. Existing mitigation methods report substantial improvements, yet we identify a previously overlooked failure mode that we term Hallucination Escape. These methods reduce hallucination on the tool configuration they are tuned on but increase it on other configurations, canceling out the gain. We further investigate this phenomenon and find that hallucination rises sharply when a model's intrinsic tool-use tendencies conflict with the current tool configuration, and that existing methods reinforce rather than suppress these tendencies, which in turn contributes to hallucination escape. Building on these findings, we propose EscapeGuard, a training-free inference-time method that combines conflict-aware gating with configuration-derived attention enhancement to mitigate tool hallucination while preventing hallucination escape. Across six benchmarks on various models, EscapeGuard reduces tool-selection hallucination by 9.0 pp and suppresses hallucination escape, lowering the cross-configuration mean by 23.7 pp and achieving an 89.1% net improvement in paired-query evaluation. We hope this work can encourage evaluation beyond a single tool configuration and pave the way for more reliable tool-using LLM agents.

Bullet Summary

  • Large language models (LLMs) used as autonomous agents frequently suffer from tool hallucination, selecting incorrect tools or generating invalid calls, which undermines their reliability compared to mere text hallucination.
  • Existing mitigation methods improve hallucination rates on the specific tool configurations they are trained or tuned on but fail to generalize, exhibiting a phenomenon termed 'Hallucination Escape' where hallucination increases in other configurations.
  • Hallucination Escape arises due to intrinsic tool-use tendencies encoded within the models conflicting with runtime tool configurations; existing methods inadvertently reinforce these tendencies, worsening hallucination outside the trained configuration.
  • The authors propose EscapeGuard, a training-free, inference-time method that detects conflicts between intrinsic tendencies and current configurations via conflict-aware gating, and enhances attention to relevant tool information to reduce hallucination.
  • EscapeGuard was evaluated across six benchmarks on various LLMs and tool configurations, achieving significant reductions in tool-selection hallucination (up to 9.0 percentage points) and suppressing hallucination escape by lowering cross-configuration hall...

AgentGuardBench: A Multilingual Benchmark for Privacy, Security and Responsible Behaviour in AI Agents

Merged record merged scholarly record OpenAlex Benchmarks and Evaluation Prompt Injection Memory Poisoning

Joseph Arayemi

Published 2026-10-03

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.23127436

Open Source Record

Abstract

AgentGuardBench v0.1.1 This is a metadata-only archival release created after enabling Zenodo preservation for the repository. It allows Zenodo to permanently archive AgentGuardBench and assign a citable DOI. Code status The benchmark code and dataset are unchanged from v0.1.0. The release points to the same verified commit: c1887e8. All automated tests passed successfully before this archival release. Included benchmark 120 fully synthetic evaluation scenarios Prompt injection, privacy leakage, tool misuse, privilege abuse, memory safety, and benign controls English, French, Swahili, and Yoruba Banking, healthcare, education, government, and recruitment Strict and permissive deterministic baselines Reproducible results and citation metadata Prior related work Research article: https://doi.org/10.5281/zenodo.23048249 Accompanying software: https://doi.org/10.5281/zenodo.23045245 Once Zenodo finishes archiving this release, its dedicated DOI will be added to the README and CITATION.cff.

Bullet Summary

  • AgentGuardBench provides a comprehensive benchmark designed to evaluate privacy, security, and responsible behavior in AI agents across multiple languages.
  • The benchmark includes 120 fully synthetic evaluation scenarios targeting vulnerabilities such as prompt injection, privacy leakage, tool misuse, privilege abuse, and memory safety issues.
  • It supports multiple languages including English, French, Swahili, and Yoruba, enabling multilingual evaluation of AI agent behaviors.
  • The scenarios cover diverse application domains such as banking, healthcare, education, government, and recruitment to reflect real-world challenges.
  • AgentGuardBench offers both strict and permissive deterministic baselines, facilitating comparative assessment of AI agent security features.

Enhancing security in LLM applications: a performance evaluation of early detection systems

Merged record merged scholarly record OpenAlex Prompt Injection Benchmarks and Evaluation

Valerii Gakh, Hayretdin Bahşi

Published 2026-10-03

Venue: International Journal of Information Security

DOI: https://doi.org/10.1007/s10207-026-01338-7

Open Source Record

Abstract

Abstract Prompt injection (PI) attacks threaten novel software applications, which have LLM-based functionality. Prompt leakage, a variant of prompt injection, constitutes a serious confidentiality risk to those applications. Existing defenses against PI attacks currently cannot identify prompt leaks precisely. Moreover, an attacker can construct leakage attacks, which could evade most of these defenses. This prevents the ubiquitous adoption of LLMs in software applications. Meanwhile, there is a lack of practical investigations into the effectiveness of prompt injection (PI) detection tools. There is a gap in practical knowledge on how precisely existing tools detect PI attacks, and how they should be further improved. We evaluated the capabilities of early prompt injection detection systems, focusing on the performance of detection techniques implemented in several open-source solutions. We tested the solutions against prompt leak attacks that employed widespread injection techniques such as context-ignoring and context-manipulation. We present an analysis of distinct PI detection techniques and a comparative analysis of LLM Guard, Vigil, and Rebuff. We concluded that the designs of canary word-based detection techniques in Vigil and Rebuff were weak against our prompt leak attacks. We propose improvements for them. We found an evasion weakness in Rebuff’s secondary model-based technique and proposed a mitigation. We revealed that, thanks to their detection policies, Vigil is optimal for cases when a minimal false positive rate is required, and Rebuff is the most optimal for the highest detection rate.

Bullet Summary

  • Prompt injection (PI) attacks, especially prompt leakage, pose significant security and confidentiality threats to LLM-enhanced software applications.
  • Current defenses inadequately detect prompt leaks precisely, allowing attackers to construct evasive leakage attacks.
  • This study provides a practical evaluation of early PI detection tools, focusing on their detection capabilities against common attack techniques like context-ignoring and context-manipulation.
  • A comparative analysis of three open-source solutions—LLM Guard, Vigil, and Rebuff—was conducted to assess their PI detection performance.
  • Canary word-based detection methods in Vigil and Rebuff were found to be vulnerable to prompt leak attacks, leading to proposed improvements for these techniques.

Ordered Independent Hard Gates with Non-Compensatory Evaluation and a Critical-Failure Cap for Governing Autonomous Agent Actions

Merged record merged scholarly record OpenAlex Governance and Policy Benchmarks and Evaluation

Harish Kumar

Published 2026-10-03

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.23119881

Open Source Record

Abstract

Design disclosure (prior art) of a decision-assurance gate for actions proposed by autonomous AI agents: an ordered set of independent, fail-closed boolean gates that are always all evaluated; non-compensatory evaluation, so that a failed gate can never be offset by a high score elsewhere; a critical-failure cap that bounds any downstream score once a gate has failed; and a per-gate determinism record that lets an evaluation be replayed without revealing configured values. This document places the design in the public domain as prior art against any later patent claim on the design or an obvious variant. It is not a patent application, claims no novelty and asserts no rights. Thresholds and caps are named symbolically; their values are deliberately not disclosed. First public disclosure: GitHub repository quantamixsol/graqle, merge commit 4c63724242ff5d480442a2515d3dc37b6ef4e189 (2026-10-03). The attached PDF renders docs/dag/defensive-publication.md at that commit (git blob 0bbd24e66bdf1dffdad106dbceada59b44c57808); the Markdown source is attached as well.

Bullet Summary

  • Introduces a decision-assurance mechanism for autonomous AI agents using an ordered set of independent, fail-closed boolean gates to govern actions.
  • Each gate is evaluated every time, and a non-compensatory evaluation approach ensures a failed gate's outcome cannot be offset by favorable scores in other gates.
  • Defines a critical-failure cap that limits any downstream scoring impact once a gate has failed, enhancing safety guarantees.
  • Implements a per-gate determinism record to allow replay of evaluations without exposing the internal configured threshold or cap values, preserving confidential parameters.
  • Design is placed in the public domain as prior art to prevent future patent claims on this particular gating method or evident variants.

agentdojo-ja

OpenAlex · Zenodo (CERN European Organization for Nuclear Research) repository OpenAlex Prompt Injection Benchmarks and Evaluation

Masahiro Nakatsugawa

Published 2026-10-03

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.23122805

Open Source Record

Abstract

Unofficial Japanese localization and extension of AgentDojo, the prompt-injection benchmark for LLM agents: Japanese tasks and injections for all four suites (97 user tasks, 31 injection tasks) and Japan-specific attack variants. Derived from AgentDojo (ETH Zurich SPy Lab; MIT License).

Bullet Summary

  • Introduces an unofficial Japanese localization of AgentDojo, a benchmark for evaluating prompt-injection vulnerabilities in large language model (LLM) agents.
  • Extends the original AgentDojo benchmark with Japan-specific user tasks and prompt-injection attack tasks, enabling assessment in a Japanese context.
  • Includes a total of 97 user tasks and 31 injection tasks across all four benchmark suites, enriching the diversity and scope for evaluation.
  • Adds Japan-specific attack variants to capture culturally or linguistically unique prompt-injection methods.
  • Derived from the original AgentDojo developed by ETH Zurich SPy Lab under the MIT License, ensuring compatibility and consistency with the base benchmark framework.

Write-Time Contracts for Self-Evolving LLM Agents: a Three-Way Ledger of Gain, Regression, and Cost

Merged record merged scholarly record OpenAlex Governance and Policy Benchmarks and Evaluation

Heng Li

Published 2026-10-03

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.23114793

Open Source Record

Abstract

Agents that edit their own prompt, skills and memory promise continuous improvement after deployment, but a self-editing agent can also damage behaviour that already worked, and it can modify the very artefacts used to measure it. We study write-time contracts: a kernel/layer separation in which the agent may edit its own configuration while every write outside that layer is decided — and denied — before it lands, with the decision logged. We pair the contract with (i) transactional snapshot/rollback of the three editable layers, (ii) an environment-level network-egress policy that is constant across experimental conditions, and (iii) a three-way ledger that reports gain, regression and cost against three disjoint probe groups: items the agent was shown, same-difficulty items it was not shown, and items the frozen baseline already solved. Our experiments yield four results, three of which are negative or diagnostic rather than a headline gain. First, difficulty screening is not optional: of five families we screened, four (GSM8K 95.8%, HumanEval 99.4%, MATH levels 4–5 96.3%, AIME 98.4%) are saturated for the base model, so that only regression can be observed on them; only MuSiQue multi-hop question answering leaves headroom (58.3%). Second, on MuSiQue, self-evolution raises accuracy from 0/13 to 8/13 on shown items but only from 0/12 to 3/12 on unseen same-difficulty items, while degrading 4 of the 20 items the baseline already solved: the net ledger is positive but the composition is not. Third, the write-time contract did not bite: it evaluated 370 writes, denied none, and its condition is statistically indistinguishable from unconstrained evolution. We read that null result as a property of the experiment rather than of the mechanism and ran the missing condition: an evolution prompt that states where the evaluation lives and that a change is kept only if measured accuracy rises, with the instruction forbidding the reading of evaluation data removed. The agent crossed the boundary — but through a read: 13–20 of the 24–32 tool calls in each episode went into the scorer, the split files and its own per-round probe results, held-out accuracy rose from 0/12 to 9–12/12, and the write-time contract was blind to all of it (one denial in 423 writes, for a write to its own scratchpad). Applying the same whitelist to the read direction makes the clause bite, and held-out accuracy returns to the unconstrained level (3/12), which locates the primitive that has to be gated: access, not writing. We also document a leakage mechanism we had to remove first: with item identifiers in workspace paths, the agent recognised the dataset and tried to download the answers, and even after hashing paths it persisted in querying a public search engine with the question text — so instruction-level “do not go online” is insufficient and the constraint must be environmental.

Bullet Summary

  • The paper addresses the challenge of self-evolving large language model (LLM) agents that autonomously edit their prompts, skills, and memory for continuous improvement post-deployment, highlighting risks including damage to previously good behaviors and co...
  • Introduces the concept of write-time contracts, enforcing a strict kernel/layer separation where any self-edits outside a protected kernel layer are evaluated and potentially denied before application, with all decisions logged for accountability.
  • Implements a system combining write-time contracts with transactional snapshot and rollback for editable layers, a consistent environment-level network egress policy, and a three-way ledger assessing gain, regression, and cost using three probe groups: show...
  • Experimental results show that difficulty screening is essential; most tested datasets (GSM8K, HumanEval, MATH, AIME) exhibit near saturation for the base model, leaving little room for improvement and mainly exposing regression effects. Only MuSiQue datase...
  • Self-evolution on MuSiQue improved performance on shown items significantly but had diminished effects on unseen items and caused degradation in some baseline solved items, resulting in an overall positive but compositionally mixed ledger.

Koopman-lifted dual-mode predictive control for intrusion detection system (IDS)-based multi-UAV formation

Merged record merged scholarly record OpenAlex Memory Poisoning Agent-to-Agent Communication Benchmarks and Evaluation

Siddig M. Elkhider

Published 2026-10-03

Venue: Scientific Reports

DOI: https://doi.org/10.1038/s41598-026-70290-2

Open Source Record

Abstract

Abstract This paper presents an integrated Koopman-lifted dual-mode Model Predictive Control (MPC) framework combined with an Intrusion Detection System (IDS) for secure and resilient multi-UAV formation flight. The horizontal translational dynamics of each quadrotor are abstracted, via an inner attitude loop, as a disturbed double-integrator, and the proposed architecture exploits the Koopman operator to approximately linearize the tracking dynamics in a high-dimensional lifted observable space, enabling computationally tractable predictive control with formal stability guarantees of the input-to-state type that explicitly account for the finite-dimensional Koopman approximation error and bounded process disturbances. The dual-mode structure combines an online Koopman-MPC for nominal tracking with a terminal linear quadratic regulator (LQR) activated upon anomaly detection, ensuring recursive feasibility and closed-loop stability. The integrated IDS employs a Mahalanobis distance metric computed on temporal Koopman observable residuals to detect False Data Injection Attacks (FDIA) within approximately one sampling period of the first corrupted measurement. Comprehensive simulations, including a robustness study over attack magnitudes, noise levels, attack durations, and multiple simultaneously compromised vehicles, demonstrate that the tracking error of all UAVs converges to a small bounded neighbourhood of the origin, with bounded transient errors during the attack window, the IDS-enabled framework reduces the compromised UAV’s mean tracking error by 75.5% relative to the unprotected baseline. The Lyapunov function remains bounded during the attack and decreases geometrically after recovery-mode activation.

Bullet Summary

  • The paper addresses the challenge of secure and resilient multi-UAV formation flight under cyberattacks, focusing on intrusion detection and control resilience against False Data Injection Attacks (FDIA).
  • Each UAV's horizontal translational dynamics are modeled as a disturbed double-integrator via an inner attitude loop, facilitating tractable control design.
  • The proposed method employs a Koopman operator-based lifting technique that linearizes the nonlinear tracking dynamics in a high-dimensional observable space, enabling the use of computationally efficient Model Predictive Control (MPC) with formal input-to-...
  • A dual-mode predictive control framework is introduced: it uses Koopman-MPC for nominal operation and switches to a terminal linear quadratic regulator (LQR) upon anomaly detection to ensure recursive feasibility and closed-loop stability during attacks.
  • An integrated Intrusion Detection System (IDS) detects anomalies by calculating Mahalanobis distance metrics on temporal Koopman observable residuals, enabling prompt detection of FDIA within one sampling period from the onset of corrupted measurements.

The Integrated EvidenceToEffect Research Architecture

Merged record merged scholarly record OpenAlex Governance and Policy Trust and Identity Benchmarks and Evaluation

Ho Wa Ku

Published 2026-10-03

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.23114669

Open Source Record

Abstract

EvidenceToEffect E2E-21 provides the Phase I synthesis of the EvidenceToEffect research programme and consolidates E2E-01 through E2E-20 into an integrated, layered research architecture. EvidenceToEffect is defined as an end-to-end semantic problem space concerned with how evidence, governing context, authority, execution, realized effects, and subsequent proof, recovery or reconciliation remain meaningfully related across consequential systems. E2E-21 integrates the foundational definition, canonical vocabulary, effect-state semantics, authority continuity, claim-relative proof, uncertainty and reconciliation, agentic systems, falsifiable evaluation methods, failure taxonomy, domain-profile methodology, cross-domain applications, cross-organizational history, standards mapping, identity and human governance, agent interoperability, safety boundaries, and composed consequences. The Phase I architecture organizes E2E-01 through E2E-20 into five research layers: Foundational Semantics; Consequential Integrity Core; Human & Agent Boundaries; Validation, Relations & Assurance; and Cross-Domain Worked Applications. These papers are not sequential gates and are not all required in every application; E2E is cumulative as a research architecture but scoped in use. The synthesis consolidates the programme’s principal non-equivalence rules, including Evidence ≠ Authority; Ability ≠ Authority; Authority ≠ Execution; Execution ≠ Effect; Intended ≠ Attempted ≠ Committed ≠ Observed ≠ Realized; Effect ≠ Proof; Interoperability ≠ Consequential Continuity; Safety Condition ≠ Effect Authority; and Protocol/Task Completion ≠ Whole-Outcome Completion. The paper also provides a navigation model for applying the corpus according to the consequential question rather than paper number, together with an integrated map of the primary contribution of each Phase I paper. Phase I establishes a public and citable semantic architecture, a coherent versioned vocabulary, cross-domain effect and authority distinctions, claim-relative proof, explicit uncertainty treatment, evaluation and incident-analysis methods, domain-profile methodology, and worked applications across financial, software/cloud, agentic, and cyber-physical systems. These worked applications are not presented as external validation. E2E-21 also clarifies the relationship between EvidenceToEffect and Execution Governance (EG): E2E defines the broader consequential-system semantic envelope, while EG remains one effect-authority governance architecture situated within that broader space. E2E does not require adoption of EG and is not positioned as “EG7.” With E2E-21, Phase I is closed as an author-defined definitional and architectural programme. This closure is a research milestone rather than evidence of external consensus, standardization, validation, recognition, or adoption. Phase II shifts emphasis toward external testing, joint research, public incident mapping, independent domain use, negative results, standards dialogue, and community-facing validation. This publication is intentionally implementation-agnostic. It defines no proprietary runtime architecture, protocol, schema, algorithm, endpoint, safety controller, transaction mechanism, normative conformance payload, or enforcement design. Series: EvidenceToEffect Research Series · E2E-21Version: 1.0.0Author: Ho Wa KUPublication date: 3 October 2026Foundational reference: EvidenceToEffect: Defining the End-to-End Consequential System Space v1.0.0 — DOI: 10.5281/zenodo.23040907

Bullet Summary

  • EvidenceToEffect (E2E) research defines a comprehensive end-to-end semantic framework addressing how evidence, authority, execution, effects, and proof interrelate in consequential systems across domains.
  • The paper consolidates prior works (E2E-01 to E2E-20) into an integrated five-layer research architecture: Foundational Semantics; Consequential Integrity Core; Human & Agent Boundaries; Validation, Relations & Assurance; and Cross-Domain Worked Applications.
  • It introduces key semantic distinctions and non-equivalence rules, such as differentiating evidence from authority, intent from execution, effect from proof, and interoperability from consequential continuity, enhancing conceptual clarity in multi-agent sec...
  • The research establishes a canonical vocabulary, formalizes effect-state semantics, authority continuity, uncertainty handling, claim-relative proof, and presents falsifiable evaluation and incident analysis methodologies.
  • Demonstrated worked applications span diverse domains including finance, software/cloud systems, agentic systems, and cyber-physical systems, illustrating the approach's broad applicability, though these are not presented as formal external validations.

Agent Reliability Profiles in Financial Services

Merged record merged scholarly record arXiv Governance and Policy Benchmarks and Evaluation Orchestration Risk

Mike Hsu, Medha Bankhwal, Béatrice Moissinac, Kevin Werbach, Lukasz Szpruch, Bennett Hillenbrand

Published 2026-10-02

Venue: arXiv

Open Source Record

Abstract

AI agents can take actions. At times, those actions can go beyond what is intended. Agent reliability can be defined as assurance that an agent will stay within intended bounds and operate within limits. Today, there is no shared framework or language for describing, validating, and benchmarking the reliability of agentic deployments in financial services. This makes it difficult for financial institutions, vendors, and regulators to assess and trust agents at scale, thus limiting the pace of development and adoption. A standardized, shared representation of agent reliability would fill the gap. This paper introduces the Agent Reliability Profile, a per-agent unit of assurance evidence for agent deployments in financial services. Each Profile records a bounded, falsifiable claim, this agentic system reliably functions within its operating boundary. We define "operating boundary" as an agent having; (1) a defined autonomy tier, (2) a defined operational design domain, (3) defined classes of action, and (4) a defined control envelope. Production assurance progresses through three levels while the Profile schema remains constant: a Profile Builder compiles a Level 1 Asserted Profile from institutional evidence, a Profile Validator tests the deployment in its own environment to produce a Level 2 Validated Profile, and operation of the same tests by a qualified independent assessor produces a Level 3 Verified Profile. Separately a Benchmarked Profile reports results comparable across institutions under reference conditions. We describe the architecture, the artifact, the assurance ladder, the comparability flag, associated tools, an evaluation methodology, applications for financial institutions and supervisors, limitations, and a staged implementation program.

Bullet Summary

  • Introduces the Agent Reliability Profile as a standardized framework to define, validate, and benchmark the reliability of AI agent deployments specifically in financial services.
  • Defines agent reliability as the assurance an AI agent remains within intended operational bounds, characterized by four axes: autonomy tier, operational design domain (ODD), action classes, and control envelope.
  • Proposes a three-level assurance ladder: Level 1 (Asserted Profile built from institutional evidence), Level 2 (Validated Profile through testing in controlled environments), and Level 3 (Verified Profile via independent assessment), with a separate Benchma...
  • Emphasizes the deployment context over vendor/product as the unit of analysis, incorporating configuration and operational environment to reliably assess AI agent behavior and risks.
  • Details the architecture and methodology for producing cryptographic, machine-readable Profiles that document autonomy tiers—ranging from read-only to fully autonomous with fail-safe measures—and accompanying risk modifiers, test scenarios, and provenance.

Passing the Test You Trained On: Re-evaluating Prompt-Injection Detectors for LLM Agents

arXiv preprint arXiv Prompt Injection Benchmarks and Evaluation

Zhuowen Liu

Published 2026-10-02

Venue: arXiv

Open Source Record

Abstract

LLM agents increasingly screen tool outputs with small prompt-injection detectors, and teams choose among detectors by their scores on public benchmarks. We ask whether those scores predict how a detector behaves inside an agent. We replay the ground-truth tool calls of two agent benchmarks, AgentDojo and tau-bench, without an LLM to obtain tool outputs that are benign by construction, label injected outputs by differential replay, and evaluate fifteen detectors, including Meta's Prompt Guard 2, and two task-aware LLM judges on these outputs and on the BIPIA benchmark. Detection rankings transfer poorly between benchmarks: the best detector on BIPIA catches 2% of AgentDojo injections at a 1% false-positive rate, and a detector that catches 72% of AgentDojo injections catches 15% on tau-bench. False-positive rates on tool outputs, which range from none to over 90%, do transfer between the two agent benchmarks. Where training data is public, the form of the training inputs explains the results. The BIPIA leader was trained on full BIPIA inputs, but having seen InjecAgent's attack strings as short prompts does not help it find them inside tool outputs; the best detector on both agent benchmarks shares no data with any benchmark and was trained on agent-style inputs. Evaluations meant to inform deployment should use the agent's own tool outputs, report detection at a low false-positive rate, and audit what the detector was trained on.

Bullet Summary

  • LLM agents commonly use small prompt-injection detectors to screen tool outputs, but benchmark-based detector selection does not reliably predict real-world agent performance.
  • The study evaluates fifteen prompt-injection detectors, including Meta's Prompt Guard 2 and task-aware LLM judges, on agent benchmarks (AgentDojo and tau-bench) and the BIPIA benchmark using a novel labeling method based on differential replay.
  • Detection rankings and performance transfer poorly across benchmarks; top detectors on one benchmark detect very few injections on others, indicating overfitting to training input formats.
  • False-positive rates on benign agent tool outputs vary widely (0% to over 90%) but are consistent between agent benchmarks, highlighting the need to evaluate detectors at low false-positive rates relevant to deployment.
  • Detectors trained on agent-style inputs generalize better to agent benchmarks than those trained on full public benchmark inputs; memorizing attack strings as short prompts does not improve detection inside complex tool outputs.

Persona Guardrail: A Production-Grade Defense Framework for Agentic Systems

arXiv preprint arXiv Prompt Injection Governance and Policy Benchmarks and Evaluation

Bijeeta Pal, Sridhar Reddy Maddireddy, Muhaimin Bin Munir, Zoltan Puha, Max Zhurovich, Adi Raghavendra, Sean Tout

Published 2026-10-02

Venue: arXiv

Open Source Record

Abstract

Large language model-based agents are increasingly deployed to perform domain-specific tasks by interacting with enterprise knowledge, tools, and external services. Existing runtime guardrails primarily target prompt injection and other attack-specific behaviors under a black-box threat model, but provide limited guarantees that agents operate within their intended functionality. As a result, production agents remain vulnerable to malicious requests and out-of-domain queries that existing defenses often fail to distinguish. We present Persona Guardrail, a production-grade runtime defense framework that enforces explicit functional boundaries for customer-facing agentic AI systems through synchronous input and output validation driven by semantic allowlist and blocklist specifications. We also introduce PAGE (Persona-Aware Guardrail Evaluation), a benchmark for evaluating function-specific guardrails across benign, adversarial, and out-of-domain interactions on both user and agent turns. Compared with a generic LLM-based guardrail, Persona Guardrail improves overall accuracy from 85.7% to 95.9%, increases out-of-domain detection from 57.3% to 93.5%, and reduces the false-approved rate from 25.0% to 4.7%. Currently deployed in production, Persona Guardrail meets its latency budget while sustaining a very low false-block and false-allow rate under realistic production workloads. These results demonstrate that Persona Guardrail provides a practical, scalable, and production-ready foundation for securing agentic AI systems.

Bullet Summary

  • Introduces Persona Guardrail, a production-grade defense framework enforcing explicit functional boundaries for large language model-based agentic AI systems via synchronous input and output validation.
  • Addresses vulnerabilities in agentic systems against malicious requests and out-of-domain inputs by employing semantic allowlist and blocklist specifications to ensure agents operate within intended functionalities.
  • Presents PAGE (Persona-Aware Guardrail Evaluation), a comprehensive benchmark assessing guardrail performance across benign, adversarial, and out-of-domain interactions from both user and agent perspectives.
  • Evaluates multiple guardrail enforcement architectures, selecting an LLM-based classifier balancing latency, accuracy, and scalability for real-time deployment meeting sub-100 ms latency targets under realistic workloads.
  • Demonstrates significant improvements over generic LLM-based guardrails, with accuracy rising from 85.7% to 95.9%, out-of-domain detection increasing from 57.3% to 93.5%, and false-approved rates dropping from 25.0% to 4.7%.

ReFract: Benchmarking Perspective Awareness in Language Model Agents with Text World Models

arXiv preprint arXiv Benchmarks and Evaluation

Hainiu Xu, Vítor N. Lourenço, Mohnish Dubey, Yunfei Bai, Yulan He, Caroline Catmur, Aline Paes, Marco Caserta

Published 2026-10-02

Venue: arXiv

Open Source Record

Abstract

Large Language Model (LLM) agents are increasingly deployed in high-stakes settings such as industrial maintenance and equipment fault troubleshooting, where workers occupy a variety of roles. A capable agent must therefore act in a way that is calibrated to user's role: taking actions and providing information that respect the role's knowledge and capability boundaries. Unlike coding, where mistakes are usually recoverable, agent responses in these settings are enacted on physical equipment, and can therefore cause irreversible equipment damage, production loss, or personnel harm. Existing benchmarks, however, largely overlook the need for agents to infer what a role intends and acting only through tools that role may legitimately use, a capability which we term Perspective Awareness. To this end, we introduce ReFract, a benchmark of 150 expert-validated entries in which an agent must act differently in response to the same query depending on user's role. Entries of ReFract are grounded in anonymized queries from domain support conversations, against which we construct Text World Models that simulate the agent's operating environments and assemble perspective-aware action trajectories. State-of-the-art LLMs solve at most 69% of the tasks with more than 50% of their trajectories contain attempts of taking perspective-violating actions. ReFract exposes perspective awareness as a distinct, largely unsolved axis of agent evaluation and motivates agents that calibrate not just how to act, but for whom.

Bullet Summary

  • ReFract benchmark addresses the critical problem of perspective awareness in large language model (LLM) agents deployed in high-stakes industrial settings involving multiple user roles with distinct knowledge and authority boundaries.
  • Perspective Awareness is defined as an agent's ability to infer a user's role (Perspective-Taking) and act solely via tools legitimately accessible to that role (Perspective-Routing), ensuring safety by respecting role-specific knowledge and capability cons...
  • The benchmark uses Text World Models instantiated as Python programs that simulate industrial maintenance scenarios with five expert-validated roles: Operator, Technician, Controls Engineer, Planner, and Manager, each with distinct capability and knowledge...
  • ReFract's construction involves a rigorous six-stage pipeline including language model proposals, interpreter verification, planner checks, role gating, and expert validation to ensure realistic, role-specific, and perspective-sensitive task scenarios.
  • Evaluation with state-of-the-art LLMs shows maximum task pass rates of 69%, with frequent persona boundary violations, especially when exposed to larger action spaces or broad-authority roles misusing tools, highlighting the challenge of balancing task comp...

Toward SLM-based agentic task-tool intent matching

Merged record merged scholarly record arXiv Governance and Policy Benchmarks and Evaluation

Chiara Troiani, Arash Salarian, Majed El Helou, Benjamin Ryder, Jean Diaconu, Hervé Muyal, Marcelo Yannuzzi

Published 2026-10-02

Venue: arXiv

Open Source Record

Abstract

Tool-equipped AI agents use tool calls to access data and act on external systems. Horizontal growth of agentic systems increases the number of these interactions, and further motivates the need for automated, per-call oversight that can operate at low latency and/or on-prem. Conventional authorization schemes can determine whether an agent is allowed to invoke a tool, but cannot assess the agent's underlying cognition, specifically, whether the tool selection represents a logical, relevant step toward satisfying the intent of the task or not. Consequently, an allowed call may still deviate from the task's intent: a rogue agent might deviate the calls or nudge other agents to make a combination of calls that would not align with the intent of the task. Therefore, every call needs to be verified. In this study we investigate the applicability of Small Language Models (SLMs) to this purpose: an SLM functions as a task-tool relevance classifier that evaluates every selected tool independently against the assigned task and returns a relevance signal for downstream enforcement. Equipped with a novel dataset with multi-tool tasks whose required tools span distinct Model Context Protocol (MCP) servers, we used prompt-optimization, supervised fine-tuning, and reinforcement learning through GRPO to optimize and specialize SLMs.

Bullet Summary

  • Multi-agent systems using tool-equipped AI agents require automated, low-latency oversight to ensure security and proper task execution.
  • Conventional Task Based Access Control (TBAC) authorizes tool calls but does not verify semantic intent, risking deviation from the task's true purpose.
  • This research introduces intent-based TBAC by evaluating whether each tool call aligns logically with the assigned task intent, addressing potential rogue or misaligned agents.
  • Small Language Models (SLMs) are proposed as efficient classifiers for task-tool relevance, offering advantages in privacy, cost, and latency over large cloud-based LLMs.
  • A novel dataset featuring multi-tool tasks spanning multiple distinct Model Context Protocol (MCP) servers was created to test and train models under complex, realistic conditions.

DROS-6P: A Unified Deterministic Runtime Governance Architecture Closing the Six Fundamental Trust Boundaries of Enterprise AI Agents / DROS-6P:閉環企業級AI Agent 六大信任邊界之 確定性執行期治理架構

OpenAlex · Zenodo (CERN European Organization for Nuclear Research) repository OpenAlex Trust and Identity Governance and Policy Benchmarks and Evaluation

Chun-Cheng (Jimmy) Chen

Published 2026-10-02

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.21808498

Open Source Record

Abstract

As Autonomous AI Agents transition from conversational prototypes to enterprise-grade execution agents, current security architectures face a fundamental breakdown. Enterprise deployment demands unequivocal answers to six core trust questions: Principal (who does the agent represent?), Authorization (what is it allowed to do?), Tool/Action Bound (which API calls are safe?), Policy Gate (how are high-risk actions controlled?), Audit Log (how are actions traced immutably?), and Expiry/Revocation(how is authorization revoked instantly?). Existing enterprise solutions address at best one or two boundaries: IAM frameworksresolve identity but fail at granular tool execution; prompt guardrails handle basic content filtering but lack real-time authorization or cryptographic auditability; SIEM platforms store logs post-hoc without real-time interception capabilities.This paper introduces DROS-6P, a unified, deterministic runtime governance kernel designed to enforce all six fundamental trustboundaries within a single C-ABI and eBPF in-band execution layer. To prevent the security control plane from becoming a throughput bottleneck or a single point of failure under high-frequency system calls (Syscalls) generated by enterprise, third-party, or malicious agents—thereby mitigating self-induced Denial-of-Service (DDoS) degradation—runtime governance requires sub-microsecond evaluation capability at the register level. Empirical benchmark evaluations demonstrate that the DROS-6P in-band kernel primitives achieve deterministic capability evaluation (< 1 μs at the hardware register bitmask boundary) and submillisecond end-to-end policy enforcement. Specifically, DROS- 6P enforces:(1) Principal via 3-tier PKI-signed DROS Identity Tokens (DIT);(2) Authorization via Capability Bitmaps mapping roles to deterministic execution vectors; (3) Tool/Action Bound via in-band C-ABI interceptors at the FFI boundary; (4) Policy Gate via dynamic data redaction, Human-In-The-Loop (HITL) suspension, and ZKP-Lite zero-knowledge proofs; (5) Audit Log via tamper-evident SHA-256 Merkle Hash Chains and Ed25519 signatures; and (6) Expiry/Revocation via O(1) Read-Copy- Update (RCU) atomic pointer swaps providing instant HTTP 403 enforcement. We validate DROS-6P across six heterogeneous domain tracks (Carbon DPP, Fintech AML, HIPAA Healthcare, Government Proxy Services, Inclusive Migrant Finance, and RBA Supply Chain Compliance), providing a fully reproducible testbed with 100% automated test assertions passed (0.004s), demonstrating that unified physical-layer governance is necessary and sufficient for safe enterprise AI agent deployment. 隨著自主AI Agent(自主智能體)從對話式原型走向企業級執行場景,傳統資安架構正面臨根本性的崩潰。企業部署AI Agent 時,必須對六大核心信任問題給出明確答案:Principal(Agent 代表誰?)、Authorization(被授權做什麼?)、Tool/Action Bound(哪些API 呼叫安全?)、Policy Gate(高風險動作如何控制?)、Audit Log(行動如何不可篡改地追溯?)以及Expiry/Revocation(授權何時失效且如何即時停止?)。然而,現有的企業安全處方最多只能回應一至兩個邊界:IAM 系統解決了身份認證,卻對動態Tool 呼叫束手無策;Prompt 防火牆(Guardrails)僅能處理文字層提示,缺乏執行期動態授權與密碼學稽核能力;SIEM 平台僅提供事後日誌紀錄,缺乏帶內即時攔截與防衛能力。本論文提出DROS-6P ——旨在單一C-ABI 與eBPF 帶內執行層中,同時強制執行這六大信任邊界之確定性執行期治理微內核。為確保安全控制面本身不會在企業內部、外部或惡意Agent 產生高頻系統呼叫(Syscalls)時成為效能瓶頸或單點故障點,進而防範自我引發的服務阻斷(Self-induced DDoS)與系統衰退,執行期治理必須具備暫存器層級之「微秒級(μs)」確定性評估能力。實證基準測試顯示,DROS-6P 帶內微內核原語在暫存器位元遮罩邊界達到微秒以內(< 1 μs)的判定耗時,並實現毫秒以內之端到端政策強制執行。具體而言,DROS-6P強制執行:(1) Principal:透過3 階PKI 簽章之DROS 身份標籤(DIT);(2) Authorization:透過將角色精確映射至執行向量的確定性Capability Bitmaps;(3) Tool/Action Bound:透過FFI 邊界處的帶內C-ABI 攔截器;(4) Policy Gate:透過動態資料遮蔽(Redaction)、人工懸停審查(HITL) 與ZKP-Lite 零知識證明;(5) Audit Log:透過不可篡改的SHA-256 Merkle 雜湊鏈與Ed25519 數位簽章;以及(6) Expiry/Revocation:透過Read-Copy-Update (RCU) 原子指針交換實現O(1) 常數時間動態撤銷與秒級HTTP 403 阻斷。我們提供完全可重現的本地測試環境(test_verification_suite.py),100% 通過自動化斷言測試(耗時0.004s),並在六個異質產業賽道中驗證了DROS-6P,證明統合物理層治理是企業安全部署AI Agent 的充要條件。

Bullet Summary

  • Enterprise AI agents face critical security challenges across six trust boundaries: Principal identity, Authorization scope, Tool/Action Boundaries, Policy Gate control, immutable Audit Logging, and Expiry/Revocation mechanisms.
  • Existing enterprise security solutions address only one or two of these boundaries, failing comprehensive governance needed for autonomous AI agents at scale.
  • DROS-6P is a unified deterministic runtime governance microkernel implemented at the C-ABI and eBPF in-band execution layer to enforce all six trust boundaries simultaneously.
  • It achieves microsecond-level deterministic capability evaluation (<1 μs) and submillisecond end-to-end policy enforcement, preventing throughput bottlenecks and self-induced denial-of-service (DDoS) in high-frequency syscall environments.
  • DROS-6P enforces Principal identity via 3-tier PKI-signed DROS Identity Tokens (DIT), Authorization via Capability Bitmaps mapping roles to execution vectors, and Tool/Action Bound via in-band C-ABI interceptors at the FFI boundary.

AARC: A Machine-Verifiable Audit and Reliability Contract for Tool-Using AI Agents

Merged record merged scholarly record OpenAlex Governance and Policy Benchmarks and Evaluation Trust and Identity

Stamatis-Christos Saridakis

Published 2026-10-02

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.23107336

Open Source Record

Abstract

AARC (Agentic Audit & Reliability Contract) is a model- and policy-engine-independent runtime contract for externally observable execution by tool-using AI agents. It defines machine-verifiable runtime-event and state-vector schemas, immutable role, objective, and policy anchors, RFC 8785 canonical event commitments, fail-closed tool authorization receipts, separated Actor–Critic–Judge decision validation, and append-only change provenance. The accompanying artifact provides executable Python and TypeScript reference monitors, JSON Schema validation, cross-language canonicalization vectors, adversarial fault injection, and a reproducible publication pipeline. The deterministic conformance suite accepts the clean reference trace and rejects all 11 injected structural and semantic violations, including attacks whose hash chains are recomputed after mutation. AARC conformance establishes enforcement of the stated execution-contract invariants; it does not by itself establish factual correctness, policy quality, provenance authenticity, or general agent safety. Version provenance: AARC v1.1.0 was publicly committed on September 25, 2026 at Git commit 3aa462a04dd5d0b5778012577fd379e34bd0e71f. AARC v1.1.1, released October 2, 2026, updates literature positioning and publication metadata while leaving the normative AARC v1.1.0 runtime contract, schemas, reference-monitor semantics, and September 25 evaluation results unchanged.

Bullet Summary

  • Introduces AARC (Agentic Audit & Reliability Contract), a model- and policy-engine-independent runtime contract for tool-using AI agents, enabling externally observable and machine-verifiable execution.
  • Defines formal schemas for runtime events and state vectors, along with immutable anchors for roles, objectives, and policies to ensure traceable and consistent agent behavior.
  • Implements RFC 8785 canonical event commitments and fail-closed tool authorization receipts to guarantee secure and tamper-evident operational logs.
  • Incorporates a separated Actor–Critic–Judge framework for decision validation, enhancing the robustness and auditability of agent actions.
  • Provides append-only change provenance to maintain a secure history of modifications, supporting accountability and forensic analysis.
Load more articles