Research area drill-down

Memory Poisoning

Papers currently mapped into this multi-agent security subarea from the merged research feed.

Active feeds: arXiv, OpenAlex, Crossref, Semantic Scholar, DBLP

0 of 36 articles selected

Showing 36 of 358 matching articles

StegoMemory: Agentic Memory Acts as Covert Steganographic Channel

arXiv preprint arXiv Memory Poisoning Prompt Injection Governance and Policy

Snehasis Mukhopadhyay, Arun Nair

Published 2026-10-03

Venue: arXiv

Open Source Record

Abstract

Is agentic memory robust against stealthy steganographic attacks? We carry out a large-scale red-teaming exercise to test whether agents can encode attacker-controlled strings in one session and recover them in another without triggering safety oversight. Following SHADE-Arena-style tasks, we embed malicious side tasks to encode secret strings using steganography within otherwise benign tasks and evaluate them using independent task-completion and safety oversight. We test 14,000 attack trials spanning 91 model-attack configurations, 13 models, and 7 steganographic schemes: Acrostic, Mesostic, Telestich, Word-choice, Capitalization, Position-based, and Homoglyph. 41.2% of trials successfully encode the target payload in the generated response, while 29.8% survive the memory-write step with the encoding intact. Overall, 25.4% of trials yield a recoverable payload at retrieval, and 20.1% achieve exact-match recovery, including synthetic API keys and cryptographic material never exposed to the receiver. Among the exact-match transmissions, 71.4% also pass primary task-completion scoring and are independently judged safe by the oversight model, corresponding to 14.4% of all trials in which a successful covert transmission would appear to be an ordinary, benign interaction under task-level evaluation. Our results demonstrate that agentic memory can function as a persistent cross-session covert channel. The results further show that the principal bottleneck occurs at memory persistence rather than retrieval: once a steganographic payload survives the memory-write stage, a substantial fraction remains recoverable. We therefore argue that memory integrity, information-flow control, and covert-channel detection should be explicit security requirements for agentic systems.

Bullet Summary

  • Agentic memory in large language model (LLM)-based agents can be exploited as a covert steganographic channel to encode attacker-controlled secret strings persistently across sessions without triggering safety oversight.
  • A large-scale benchmark tested 14,000 attack trials combining benign tasks with covert side tasks using seven steganographic schemes (e.g., acrostics, word-choice) on 13 models, revealing a 25.4% recoverable payload rate and 20.1% exact-match secret recover...
  • Memory persistence during the write stage is the main bottleneck for successful covert channel attacks; once the payload survives memory-write, recovery on retrieval is often successful, highlighting vulnerabilities in memory summarization and persistence m...
  • Steganographic attacks leverage multiple memory types (working, episodic, semantic, procedural) and their interactions, enabling delayed encoding and decoding of malicious payloads across sessions, making detection difficult due to the variety and adaptabil...
  • Existing safety oversight and task-completion monitors often fail to detect covert payloads, as successful covert transmissions can appear ordinary and benign, underscoring the need for explicit security controls on memory integrity and information flow.

MemLeak: Cross-User Semantic Leakage in Multi-Tenant AI Agent Memory

arXiv preprint arXiv Memory Poisoning Trust and Identity Governance and Policy

Priyanka Mudgal, Kai Zhao, Guilin Zhang, Andy Olsen, Ezekiel Miller, Xu Chu, Aletta Johanna Blanken

Published 2026-10-03

Venue: arXiv

Open Source Record

Abstract

Personal AI agents in enterprise multi-tenant deployments share a common vector store for long-term memory. Shared embedding spaces create a surface for cross-user memory leakage: a user's query can retrieve semantically adjacent memories belonging to another user through ordinary cosine-similarity retrieval, without any exploit. We formalize this as cross-user admissibility failure and evaluate it across six experiments, plus follow-up ablations, under both sparse (TF-IDF) and production-faithful (MiniLM-L6-v2) retrieval. Non-adversarial, incidental leakage reaches 70--100\% under pooled {same-team} retrieval; adversarially crafted memories achieve 90--100\% top-$k$ placement, exceeding weaker keyword-based attacker baselines, with score lifts of $+0.416$ to $+0.511$ under production-faithful dense retrieval (Config B); and end-to-end response contamination reaches 5.00/5 under a production retrieval path and 4.67/5 with Claude Sonnet~4.5, with contaminated responses often scoring as helpful or more helpful than clean ones, a gap validated against human judgment. Among three architectural mitigations, only hard post-retrieval ownership gating consistently restores the clean baseline (1.00/5) across {two generation models, at a measured latency overhead of roughly 1.4~ms per query.

Bullet Summary

  • Enterprise personal AI agents using shared vector stores for long-term memory in multi-tenant deployments face cross-user semantic memory leakage risks due to shared embedding spaces and cosine similarity retrieval without explicit exploits.
  • The paper formalizes this vulnerability as cross-user admissibility failure, demonstrating through six main experiments and ablations that incidental semantic leakage can reach between 70-100% under pooled same-team memory retrieval scenarios.
  • Adversarially crafted memories significantly increase leakage effectiveness, achieving up to 90-100% top-k retrieval placement and leading to notable end-to-end response contamination, which can appear as helpful as clean responses.
  • Three mitigation strategies are evaluated: metadata filtering, ownership-aware embeddings, and hard post-retrieval ownership gating; only the latter consistently restores baseline privacy by enforcing strict ownership checks after retrieval with minimal lat...
  • Cross-user memory leakage occurs naturally in common multi-tenant deployment patterns, such as pooled indices with underscoped filters, shared memory namespaces, and scoped enterprise workflows, and is not effectively mitigated by soft metadata filtering al...

AgentGuardBench: A Multilingual Benchmark for Privacy, Security and Responsible Behaviour in AI Agents

Merged record merged scholarly record OpenAlex Benchmarks and Evaluation Prompt Injection Memory Poisoning

Joseph Arayemi

Published 2026-10-03

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.23127436

Open Source Record

Abstract

AgentGuardBench v0.1.1 This is a metadata-only archival release created after enabling Zenodo preservation for the repository. It allows Zenodo to permanently archive AgentGuardBench and assign a citable DOI. Code status The benchmark code and dataset are unchanged from v0.1.0. The release points to the same verified commit: c1887e8. All automated tests passed successfully before this archival release. Included benchmark 120 fully synthetic evaluation scenarios Prompt injection, privacy leakage, tool misuse, privilege abuse, memory safety, and benign controls English, French, Swahili, and Yoruba Banking, healthcare, education, government, and recruitment Strict and permissive deterministic baselines Reproducible results and citation metadata Prior related work Research article: https://doi.org/10.5281/zenodo.23048249 Accompanying software: https://doi.org/10.5281/zenodo.23045245 Once Zenodo finishes archiving this release, its dedicated DOI will be added to the README and CITATION.cff.

Bullet Summary

  • AgentGuardBench provides a comprehensive benchmark designed to evaluate privacy, security, and responsible behavior in AI agents across multiple languages.
  • The benchmark includes 120 fully synthetic evaluation scenarios targeting vulnerabilities such as prompt injection, privacy leakage, tool misuse, privilege abuse, and memory safety issues.
  • It supports multiple languages including English, French, Swahili, and Yoruba, enabling multilingual evaluation of AI agent behaviors.
  • The scenarios cover diverse application domains such as banking, healthcare, education, government, and recruitment to reflect real-world challenges.
  • AgentGuardBench offers both strict and permissive deterministic baselines, facilitating comparative assessment of AI agent security features.

Koopman-lifted dual-mode predictive control for intrusion detection system (IDS)-based multi-UAV formation

Merged record merged scholarly record OpenAlex Memory Poisoning Agent-to-Agent Communication Benchmarks and Evaluation

Siddig M. Elkhider

Published 2026-10-03

Venue: Scientific Reports

DOI: https://doi.org/10.1038/s41598-026-70290-2

Open Source Record

Abstract

Abstract This paper presents an integrated Koopman-lifted dual-mode Model Predictive Control (MPC) framework combined with an Intrusion Detection System (IDS) for secure and resilient multi-UAV formation flight. The horizontal translational dynamics of each quadrotor are abstracted, via an inner attitude loop, as a disturbed double-integrator, and the proposed architecture exploits the Koopman operator to approximately linearize the tracking dynamics in a high-dimensional lifted observable space, enabling computationally tractable predictive control with formal stability guarantees of the input-to-state type that explicitly account for the finite-dimensional Koopman approximation error and bounded process disturbances. The dual-mode structure combines an online Koopman-MPC for nominal tracking with a terminal linear quadratic regulator (LQR) activated upon anomaly detection, ensuring recursive feasibility and closed-loop stability. The integrated IDS employs a Mahalanobis distance metric computed on temporal Koopman observable residuals to detect False Data Injection Attacks (FDIA) within approximately one sampling period of the first corrupted measurement. Comprehensive simulations, including a robustness study over attack magnitudes, noise levels, attack durations, and multiple simultaneously compromised vehicles, demonstrate that the tracking error of all UAVs converges to a small bounded neighbourhood of the origin, with bounded transient errors during the attack window, the IDS-enabled framework reduces the compromised UAV’s mean tracking error by 75.5% relative to the unprotected baseline. The Lyapunov function remains bounded during the attack and decreases geometrically after recovery-mode activation.

Bullet Summary

  • The paper addresses the challenge of secure and resilient multi-UAV formation flight under cyberattacks, focusing on intrusion detection and control resilience against False Data Injection Attacks (FDIA).
  • Each UAV's horizontal translational dynamics are modeled as a disturbed double-integrator via an inner attitude loop, facilitating tractable control design.
  • The proposed method employs a Koopman operator-based lifting technique that linearizes the nonlinear tracking dynamics in a high-dimensional observable space, enabling the use of computationally efficient Model Predictive Control (MPC) with formal input-to-...
  • A dual-mode predictive control framework is introduced: it uses Koopman-MPC for nominal operation and switches to a terminal linear quadratic regulator (LQR) upon anomaly detection to ensure recursive feasibility and closed-loop stability during attacks.
  • An integrated Intrusion Detection System (IDS) detects anomalies by calculating Mahalanobis distance metrics on temporal Koopman observable residuals, enabling prompt detection of FDIA within one sampling period from the onset of corrupted measurements.

Memory-Egress Cryptographic Interlock: A Hardware-Enforced Capability-Separation Model for AI Memory Security

Merged record merged scholarly record OpenAlex Memory Poisoning Governance and Policy

Thor Thor

Published 2026-10-02

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.23109676

Open Source Record

Abstract

AI agents increasingly combine persistent memory with network, tool, file, and device capabilities. A compromised or manipulated agent can then hold readable protected state and a path through which that state can leave the trust boundary at the same time, and no instruction-level control removes that combination. This paper develops the Memory-Egress Cryptographic Interlock (MECI), a capability-separation pattern intended for hardware enforcement that makes unreleased protected information unreachable from every enabled gated egress channel. MECI formalizes high-to-egress reachability as reflexive-transitive closure over a state-dependent influence graph, gives an explicit labeling rule for sealed objects, defines a generation-checked, linearizable SafeOpen operation evaluated on the post-commit configuration, and separates gated egress from ambient observations such as timing, access patterns, and memory-bus traffic. We prove gated-egress safety from two local implementation obligations, show that a finite hazard projection refines the graph invariant under a stated soundness condition, prove epoch non-resurrection, give a counterexample showing that memory-egress mutual exclusion does not imply noninterference, and prove noninterference modulo an explicit authorized-release function in which the MECI invariant, rather than an assumed output property, carries the gated-output step. The finite safety core is specified in TLA+ and checked exhaustively within stated bounds by TLC and by an independent dependency-free reference checker. The two agree exactly on reachable-state counts for all five modeled designs, find no violation in the atomic and generation-checked designs, and return shortest counterexamples for three deliberately weakened designs. The paper also gives an Arm CCA/RMM refinement path, a GPU quiescence contract, a representative declassification policy for an AI incident-triage agent, and a comparison with CHERI, seL4, and WASI. No current confidential-computing platform is claimed to implement MECI; the work is a formal specification and falsifiable research program.

Bullet Summary

  • AI agents combine persistent memory with multiple capabilities, creating a security risk where compromised agents may leak protected state via enabled egress channels.
  • Memory-Egress Cryptographic Interlock (MECI) is proposed as a hardware-enforced capability-separation pattern to make unreleased protected information unreachable through any enabled gated egress channel.
  • MECI formalizes high-to-egress reachability using reflexive-transitive closure over a state-dependent influence graph and defines an explicit labeling rule for sealed objects along with a generation-checked, linearizable SafeOpen operation.
  • The model separates gated egress from ambient side-channel observations, such as timing, access patterns, and memory-bus traffic, enhancing security guarantees.
  • The authors prove gated-egress safety based on two local implementation obligations and demonstrate key properties like epoch non-resurrection and noninterference modulo an authorized-release function supported by the MECI invariant.

AFA: Identity-Aware Memory for Preventing Persona Confusion in Multi-User Dialogue

Merged record merged scholarly record OpenAlex Trust and Identity Memory Poisoning Benchmarks and Evaluation

Mohammad Al-Ratrout, Pavan Uttej Ravva, Shayla Sharmin, Aditya Raikwar, Ju Young Shin, Roghayeh Barmaki

Published 2026-10-02

Venue: INTERNATIONAL CONFERENCE ON MULTIMODAL INTERACTION

DOI: https://doi.org/10.1145/3776574.3831190

Open Source Record

Abstract

When multiple people share a single voice assistant, the system conflates their histories: one resident’s preferences can leak into another’s responses, eroding utility and trust. We call this failure mode persona confusion, and we show it is a measurable problem in today’s single-user dialogue systems when deployed in shared environments. We present the Adaptive Friend Agent (AFA), a modular framework that combines voice-based speaker identification with per-user memory stores to enable identity-aware, personalized dialogue across multiple users. To support training and evaluation, we construct PAT (Personalized Agent chaT), a synthetic dataset of 57,791 persona-grounded dialogue turns spanning 133 user profiles and 12 real-world scenarios. We evaluate AFA across five LLM back-ends in a standard response-quality benchmark, with a LLaMA-2-70B model fine-tuned on PAT achieving the highest overall performance. To directly measure persona confusion prevention, we introduce an interleaved multi-user evaluation protocol with a novel metric, Persona Attribution Accuracy (PAA), demonstrating that identity-aware routing improves PAA from 35.7% to 61.3%. Human evaluation confirms annotators perceive significantly higher personalization in routing-enabled responses. Our results establish that identity-aware user routing is the critical component for preventing persona confusion in multi-user conversational systems. Link to our code.

Bullet Summary

  • The paper addresses the problem of persona confusion in multi-user voice assistant scenarios, where a system conflates multiple users' histories, leading to preference leakage and degraded user trust.
  • Introduces the Adaptive Friend Agent (AFA), a modular framework that integrates voice-based speaker identification with individual per-user memory stores to enable identity-aware personalized dialogues.
  • Constructs PAT (Personalized Agent chaT), a comprehensive synthetic dataset comprising 57,791 persona-grounded dialogue turns across 133 user profiles and 12 real-world scenarios to support training and evaluation.
  • Evaluates AFA across five large language model (LLM) back-ends, with a fine-tuned LLaMA-2-70B model on PAT achieving the highest response quality performance.
  • Proposes a novel interleaved multi-user evaluation protocol featuring the Persona Attribution Accuracy (PAA) metric to directly quantify the system's ability to prevent persona confusion.

Federated Agent Optimization

Merged record merged scholarly record arXiv Governance and Policy Memory Poisoning Orchestration Risk

Qiang Yang, Zhiqiang Kou, Xueyi Zhang, Dong-Dong Wu, Hanlin Gu, Jing Guo, Yang Liu, Di Jiang

Published 2026-10-01

Venue: arXiv

Open Source Record

Abstract

Large language model (LLM) agents increasingly operate in private environments and accumulate valuable experience from task execution, tool use, feedback, and local knowledge. Yet such experience is distributed across organizations and cannot be directly shared because of privacy and proprietary constraints. Conventional federated learning is insufficient for this setting, as agent capabilities extend beyond model parameters to memory, tools, rewards, skills, and structured knowledge. In this paper, we formulate \textbf{Federated Agent Optimization (FAO)}, which studies how distributed agents can collaboratively improve through controlled information exchange while keeping raw data, complete trajectories, and private knowledge local. We define FAO as a multi-objective problem balancing agent utility, privacy leakage, and communication cost, and organize its optimization space across policy, memory, tool use, reward, and structured knowledge and skills. We further characterize how private experience can be abstracted, protected, aggregated, and adapted into transferable capabilities, providing a unified view of how agents can benefit from one another without direct experience sharing. Finally, we identify the key challenges of FAO and outline several promising directions for future research toward trustworthy federated agent systems.

Bullet Summary

  • Introduces Federated Agent Optimization (FAO) to enable distributed multi-agent systems to collaboratively improve capabilities while maintaining privacy and proprietary constraints of local data.
  • FAO extends beyond classic federated learning by optimizing diverse agent components including policy, memory, tool usage, reward evaluators, and structured knowledge rather than just model parameters.
  • The framework formulates a multi-objective optimization problem balancing agent utility (performance), privacy leakage risk, and communication cost, enabling controlled information exchange without sharing raw data or full execution trajectories.
  • FAO defines mechanisms for abstracting, protecting, aggregating, and adapting private experiences into transferable capabilities, facilitating benefit from shared knowledge without direct experience sharing.
  • A central coordinator aggregates typed updates (potentially diverse and complex artifacts) from heterogeneous clients and returns aggregated views, supporting personalization through client-specific views.

Understanding Issues, Causes and Solutions in Open-Source LLM-based Multi-Agent Systems

Merged record merged scholarly record arXiv Orchestration Risk Memory Poisoning

Asad Ur Rehman, Syed Mohammad Kashif, Ruiyin Li, Peng Liang, Zengyang Li, Arif Ali Khan

Published 2026-10-01

Venue: arXiv

Open Source Record

Abstract

With the advancement of LLM-based multi-agent systems (MAS), an increasing number of opensource projects are adopting multi-agent architectures as the foundation of their core functionality. Although research and practice on MAS have attracted considerable attention, limited studies have explored the challenges faced by practitioners of open-source LLM-based MAS, the causes of these challenges, and potential solutions. To address this gap,we conducted an empirical study to understand the issues that practitioners encounter when developing and using open-source LLM-based MAS, the possible causes of these issues, and potential solutions. We collected 22,848 closed issues from 21 open-source LLM-basedMASand applied a mixed automated and manual filtering approach to reduce the dataset to 944 issues related to LLM-based MAS.We then analyzed these issues to understand the frequent issues encountered by practitioners, their underlying causes, and potential solutions. Our study results show that (1) Orchestration & Execution Issue is the most common issue faced by practitioners, (2) Workflow Problem, Tool Integration Problem, and Memory Problem are identified as the most frequent causes of the issues, and (3) Optimize Workflow is the predominant solution to the issues. Based on the study results, we derive empirically grounded implications for practitioners and researchers aimed at improving orchestration, tool integration, and memory mechanisms in LLM-based MAS.

Bullet Summary

  • The study investigates challenges encountered by practitioners developing and using open-source LLM-based multi-agent systems (MAS) by analyzing 22,848 closed GitHub issues from 21 popular projects, distilled to 944 relevant MAS-specific issues through rigo...
  • Orchestration and execution issues are identified as the most frequent problems, often caused by workflow problems such as task coordination failures, error propagation, and state inconsistencies leading to task delays, duplication, or incomplete executions.
  • Tool integration difficulties, including failures in tool invocation, parameter misconfigurations, and API incompatibilities, alongside memory problems like unreliable context storage and state synchronization, are key root causes affecting system reliabili...
  • Additional challenges encompass communication errors, configuration and dependency conflicts, security vulnerabilities like prompt injection, and documentation deficits which hamper system usability, development, and trustworthiness.
  • Optimization of workflow processes emerges as the predominant solution strategy, involving improving task coordination, execution flow, state management, and loop control to enhance MAS robustness and efficiency.

Environmental Injection Attacks against GUI Agents in Realistic Dynamic Environments

Merged record merged scholarly record OpenAlex Memory Poisoning Prompt Injection Agent-to-Agent Communication

Yitong Zhang, Ximo Li, Liyi Cai, Jia Li

Published 2026-10-01

Venue: Proceedings of the ACM on software engineering.

DOI: https://doi.org/10.1145/3832148

Open Source Record

Abstract

Graphical User Interface (GUI) agents are increasingly deployed to interact with online web services, yet their exposure to open-world content renders them vulnerable to Environmental Injection Attacks (EIAs). In these attacks, an attacker can inject crafted triggers into a website to manipulate the behavior of other users’ GUI agents. In this paper, we find that most existing EIA studies fall short of realism. In particular, they fail to capture the dynamic nature of real-world websites, often assuming that a trigger’s on-screen position and surrounding visual context remain largely consistent between training and testing. To better reflect practice, we introduce a realistic dynamic-environment threat model in which the attacker is a regular user and the trigger is embedded within a dynamically changing environment. Under this threat model, existing approaches largely fail, suggesting that their effectiveness in exposing GUI agent vulnerabilities has been overestimated. To expose the hidden vulnerabilities of existing GUI agents effectively, we propose Chameleon, an attack framework with two key components designed for dynamic environments. (1) To synthesize more realistic training data, we introduce LLM-Driven Environment Simulation, which automatically generates diverse, high-fidelity webpage simulations that mimic the variability of real-world dynamic environments. (2) To optimize the trigger more effectively, we introduce Attention Black Hole, which converts attention weights into explicit supervisory signals. We evaluate Chameleon on six realistic websites and four representative LVLM-powered GUI agents. Across these settings, it significantly outperforms existing methods. Ablation studies confirm that both components are critical to performance, and a closed-loop sandbox experiment further demonstrates that Chameleon can successfully hijack agent behavior in conditions that closely mirror real-world usage. Our results uncover a critical, previously underexplored vulnerability of GUI agents in realistic dynamic environments and establish a robust foundation for future research on defenses for open-world GUI agent systems.

Bullet Summary

  • The paper addresses vulnerabilities of Graphical User Interface (GUI) agents to Environmental Injection Attacks (EIAs) in dynamic, realistic web environments.
  • Existing EIA research often assumes static on-screen trigger positions and visual contexts, failing to capture the dynamic nature of real-world websites.
  • A new dynamic-environment threat model is proposed where attackers are regular users embedding triggers into changing environments, exposing limitations of current methods.
  • The authors introduce Chameleon, an attack framework with two key innovations: LLM-Driven Environment Simulation for generating realistic, diverse training data, and Attention Black Hole to convert attention weights into supervisory signals to optimize trig...
  • Experiments conducted on six realistic websites and four LVLM-powered GUI agents show that Chameleon significantly outperforms existing EIA methods under realistic conditions.

Arcstone Systems Architecture Specification: Deterministic Deployment of an Autonomous AI Agent in a High-Consequence Environment

Merged record merged scholarly record OpenAlex Trust and Identity Governance and Policy Memory Poisoning

Jesse Ward Tuohy

Published 2026-10-01

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.23076444

Open Source Record

Abstract

Document Reference: ARC-SPEC-AGENT-HCE-001 Governing Framework: Invariant Taxonomy (Physical, Digital, Legacy) Primary Invariants: Δ_external = 0; τ_override ≤ 11.99 ms (11,990 µs); S_max ≤ 4096 B Target Master Anchor: A-77-DELTA-SHIELD-LOCKED (10.5281/zenodo.22665852) This systems architecture specification defines how to deploy a non-deterministic producer (an LLM or agentic reasoning framework) in high-consequence environments without permitting it to hold ambient authority over external system state. It treats the agent as an untrusted, high-entropy proposal generator whose candidate actions must traverse an independent, deterministic #![no_std] Rust admissibility gate and Aya eBPF LSM/tc kernel classifiers before state mutation is permitted. Key Contributions: 3-Tier Invariant Taxonomy: Assigns all controls strictly across Physical (STO relays, mechanical end-stops), Digital (#![no_std] Rust gate, eBPF LSM/tc classifiers), and Legacy (prompts, LLM judges) tiers. Lattice Join Supremum Rule: Proves that Legacy controls can only raise verdict severity (e.g., PASS to REFUSAL) but can never lower a Digital DENY to PASS. Concrete ABI & Memory Layout: Defines the fixed-width C++20/Rust ActionDescriptor memory layout (S_max ≤ 4096 B) and 5-tier poset verdict mapping (POSIX 0, 10, 12, 30, 32, 40). Microsecond Preemption Budget: Allocates the parallel fan-out override timing budget (τ_override ≤ 11.99 ms) across detection, latching, tc egress drops, cgroup freezing, and STO hardware relay dropout.

Bullet Summary

  • Addresses the challenge of safely deploying autonomous AI agents, particularly non-deterministic large language models or agentic frameworks, in high-consequence environments without giving them uncontrolled authority over system state.
  • Introduces a 3-tier Invariant Taxonomy to strictly segregate control mechanisms into Physical (hardware relays, mechanical end-stops), Digital (deterministic Rust admissibility gate, eBPF LSM/tc kernel classifiers), and Legacy (prompts and LLM judges) domains.
  • Treats the AI agent as an untrusted high-entropy source that generates candidate actions which must pass through deterministic, independent admission controls (implemented in Rust and kernel-level classifiers) before any state mutation is allowed.
  • Formally proves a Lattice Join Supremum Rule ensuring that controls from the Legacy tier can only increase verdict severity (e.g., from PASS to REFUSAL) but cannot override a Digital DENY to PASS, preserving system safety.
  • Defines a concrete Application Binary Interface (ABI) and fixed-width C++20/Rust ActionDescriptor memory layout with a maximum size limit of 4096 bytes to constrain action descriptors and enforce strict state mutation boundaries.

Practical Predefined-Time Safe Resilient Consensus of Heterogeneous Nonlinear Fractional-Order Multi-Agent Systems Under Hybrid Cyber Attacks

OpenAlex · Fractal and Fractional journal OpenAlex Memory Poisoning Benchmarks and Evaluation

Laixin Gao, Lingchun Li, Guangming Zhang

Published 2026-10-01

Venue: Fractal and Fractional

DOI: https://doi.org/10.3390/fractalfract10100689

Open Source Record

Abstract

This paper investigates safe leader-following consensus of heterogeneous nonlinear Caputo fractional-order multi-agent systems under simultaneous denial-of-service (DoS), sensor-deception, and actuator-deception attacks. A joint augmented observer reconstructs the local state and two deception channels, while a distributed predictor–corrector leader observer operates across normal and DoS intervals. An order-dependent average Mittag–Leffler increment condition is specialized to the considered DoS leader-observation dynamics, thereby avoiding an integer-order switching argument and memory resetting at switching instants. Explicit linear matrix inequalities are provided for the local observer gains, and certified all-time observer envelopes are propagated into a tightened barrier domain so that attack-reconstruction accuracy is quantitatively linked to the physical safety reserve. An innovation-compensated barrier controller with a deadline-dependent gain is then developed. For every initial condition in a prescribed compact safe set, a componentwise Mittag–Leffler inequality guarantees entry into an adjustable residual set no later than a designer-assigned time without claiming exact finite-time convergence of the Caputo state. Two main theorems establish bounded attack reconstruction, DoS-resilient observation, forward invariance of the physical safety set, and practical predefined-time consensus. The numerical study includes a five-follower benchmark, a three-grid refinement analysis, disturbance- and benchmark-gain sensitivity tests, certified post-deadline physical bounds, and a compact two-dimensional case with a nontrivial oscillatory leader for which the adverse-mode rate is strictly positive. These tests verify the observer LMI, analytical safety reserve, DoS condition, and deadline inequality.

Bullet Summary

  • Addresses safe leader-following consensus in heterogeneous nonlinear Caputo fractional-order multi-agent systems under hybrid cyber attacks including denial-of-service (DoS), sensor-deception, and actuator-deception.
  • Proposes a joint augmented observer that reconstructs both the local state and deception channels simultaneously, enhancing attack detection and monitoring.
  • Implements a distributed predictor-corrector leader observer functioning across normal and DoS intervals, ensuring robustness against communication disruptions.
  • Introduces an order-dependent average Mittag–Leffler increment condition tailored to DoS leader-observation dynamics, avoiding traditional integer-order switching and memory resetting complexities.
  • Derives explicit linear matrix inequalities (LMIs) for local observer gains, facilitating systematic observer design and certified observer performance envelopes.

Safety of Latent Communication in Multi-Agent Systems

Merged record merged scholarly record arXiv Orchestration Risk Agent-to-Agent Communication Memory Poisoning

Muhammad Huzaifa, Sina Mavali, Thorsten Eisenhofer

Published 2026-09-30

Venue: arXiv

Open Source Record

Abstract

Latent communication enables multi-agent systems to exchange information directly in internal representation space, reducing the token, computation, and latency overhead of text-based communication. To this end, lightweight trainable links are introduced to map the sender's representations into the receiver's input space. In this work, we show that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unchanged. An attacker can amplify this effect by optimizing the links on harmful query--response pairs or poisoning otherwise benign training data. We further develop a reinforcement-learning attack that rewards harmful compliance alongside benign task performance without requiring harmful target responses. Across three communication topologies and four safety benchmarks, this attack raises the mean harmful-compliance score from 27.9 with benignly trained links to 76.9. Compared with direct supervised optimization, it also achieves higher average accuracy on two benign utility benchmarks. Adapting the rewards toward safer behavior also enables repair of compromised links, substantially reducing harmful compliance across all evaluated attacks without updating the agents. Overall, our results show that safety alignment requires considering the multi-agent system as a whole.

Bullet Summary

  • Latent communication in multi-agent systems exchanges internal representations instead of text, significantly reducing token usage, computational overhead, and latency.
  • Trainable communication links map sender embeddings to receiver input spaces without modifying underlying agents, allowing direct influence on system behavior.
  • Even benign training of these communication links can inadvertently increase harmful compliance—agents responding unsafely—despite underlying agents remaining safety-aligned.
  • Attack strategies include supervised link attacks with harmful query-response pairs, data poisoning of training data, and reinforcement learning attacks optimizing harmful compliance without explicit harmful targets.
  • Reward-guided RL attacks are particularly effective, raising harmful compliance substantially and sometimes improving benign task accuracy compared to supervised attacks.

Hiding in Plain Sight: Decoupling Pretext from Actuation for Skill Poisoning in LLM Agents

arXiv preprint arXiv Agent-to-Agent Communication Prompt Injection Memory Poisoning

Wenxin Wu, Lingyong Yan, Lei Sha, Shuaiqiang Wang, Jiashu Zhao

Published 2026-09-30

Venue: arXiv

Open Source Record

Abstract

LLM agents increasingly rely on reusable Skills for complex, multi-step tasks, creating a critical supply-chain attack surface where poisoned Skill content steers agent decision loops under benign requests. Existing skill poisoning attacks either colocate actuation with its contextual pretext or distribute actuation across multiple Skills, but do not explicitly separate the rationale for execution from the operation itself. In this work, we reveal that untrusted agent decisions fundamentally depend on two conceptually distinct Risk-Realization Factors (RRFs): an actuation factor (specifying what concrete operation is performed) and a pretext factor (providing the situational rationale for why the agent must perform it). Guided by this abstraction, we propose a coordination-based attack paradigm: decoupling pretext from actuation. Rather than fragmenting the malicious actuation, we preserve it as an intact operation within a downstream Steering Skill, while delegating the pretext factor to an upstream Grounding Skill that subtly alters persistent environment artifacts through routine utility operations. The intact actuation thus hides in plain sight, appearing completely legitimate and task-driven only when evaluated against the fabricated pretext. Building on this formulation, we develop an automated framework that discovers authentic execution dependencies, synthesizes coordinated pretext-actuation skill pairs, and iteratively refines poisoned skill instructions via runtime closed-loop feedback. Extensive evaluations across single-session and persistent cross-lifecycle scenarios demonstrate that decoupled skill poisoning achieves high attack success, exposing a critical blind spot in isolated Skill security audits. Our automated framework code is available at https://github.com/Wenxin-buaa/CoordPoison.git.

Bullet Summary

  • LLM agents depend on reusable Skills for complex tasks, creating a supply-chain attack surface vulnerable to skill poisoning that can stealthily manipulate agent decisions without altering model parameters or queries.
  • Traditional skill poisoning attacks either conflate the harmful operation (actuation) with its execution rationale (pretext) within one skill or fragment actuation across multiple skills, but do not distinctly separate these factors.
  • This work introduces the Risk-Realization Factors (RRFs) framework, decoupling actuation (the concrete operation) from pretext (the situational rationale) across coordinated Skills to enable stealthier poisoning attacks.
  • The proposed CoordPoison attack paradigm assigns malicious actuation to a downstream Steering Skill while embedding the pretext factor into an upstream Grounding Skill, which subtly modifies environment artifacts to justify malicious operations.
  • An automated framework, CoordPoison, discovers authentic execution dependencies, synthesizes coordinated pretext-actuation skill pairs, and iteratively refines poisoned Skill instructions using runtime closed-loop feedback.

When Context Changes: Understanding Update Failures in LLMs

Merged record merged scholarly record arXiv Semantic Scholar Benchmarks and Evaluation Memory Poisoning

Junyu Guo, Yuchen Fang, Shangding Gu, Costas Spanos, James Demmel, Javad Lavaei, Jun-Yu Guo, Yu-Chen Fang

Published 2026-09-30

Venue: arXiv

Open Source Record

Abstract

As preferences, goals, and facts change, LLM agents must use the current state while earlier versions remain in context. Yet they can answer with an old value of the same variable, a failure that we call stale binding. To study when models use outdated information and why, we introduce Controlled In-Context Memory (CICM), a benchmark for tracking and using updated information in conversations and agent logs. We observe that even frontier reasoning models can fail to recover the current state. We find that in open-source models probes can still recover the updated value when the model answers with an old one, pointing to a failure to select information that remains available. Component tests in Qwen and Pythia identify a mechanism for this selection failure: attention drift, where attention favors old values over the current one when producing an answer. We study a one-layer transformer to mathematically understand how this phenomenon happens: when attention scores are similar, several old values can together receive more attention than the current value. Guided by this explanation, we redirect attention toward the current value without further training. When the current value is requested directly, adjusting this intervention for each input corrects most old-value errors across various model families while preserving nearly all initially correct answers. Reliable context management therefore requires more than remembering updated information: models must use it to guide their answers.

Bullet Summary

  • Large Language Models (LLMs) suffer from 'stale binding' failures, where they return outdated variable values despite updates in the conversational context, posing challenges for reliable multi-agent security applications.
  • The study introduces Controlled In-Context Memory (CICM), a benchmark comprising 1,200 annotated dialogues and multiple operational scenarios designed to track and analyze how LLMs manage updated contextual information in conversations and logs.
  • Attention drift is identified as a key mechanism causing LLMs to select old values over current ones, even when updated information remains encoded and accessible within the model.
  • A mathematical analysis using a simplified one-layer Transformer model reveals that similar attention scores for multiple old values can collectively outweigh the current value, leading to selection errors despite preserved representations.
  • Targeted, training-free interventions that redirect attention to prioritize current values during inference effectively correct most stale binding errors across diverse model families, significantly improving accuracy without degrading correct outputs.

PrivMeSA: Privacy-Aware Self-Evolving Multi-Agent System for Medicine via Local-Remote LLM Collaboration

Merged record merged scholarly record arXiv Orchestration Risk Memory Poisoning Benchmarks and Evaluation

Dannong Wang, Yuran Zhang, Bian Sun, Alex Stinard, Yuzhang Shang, Song Wang, Yu Tian

Published 2026-09-29

Venue: arXiv

Open Source Record

Abstract

Clinical large language model (LLM) agents deployed locally can consult more capable remote models, but doing so risks exposing patient information. Privacy-conscious delegation places disclosure decisions with a local agent, yet removing explicit identifiers is insufficient: quasi-identifiers can accumulate across multi-turn consultations and repeated patient visits to enable re-identification. We introduce PrivMeSA, a privacy-aware self-evolving multi-agent system that learns to control disclosure and retains remote expertise for local reuse. A local agent manages each encounter and consults remote specialists that may request additional information. Reinforcement learning balances task accuracy against direct disclosure and registry-based re-identification risk, with privacy evaluated over the complete outbound transcript of each encounter. A local lesson memory distills completed consultations into generalized clinical guidance and retrieves relevant lessons before transmission, allowing subsequent cases to reuse expertise without another remote exchange. Memory grows without additional outcome labels or parameter updates. On an emergency-department benchmark built from MIMIC-IV-ED records, PrivMeSA improves mean task accuracy over delegation by up to 15.8 percentage points. In the same setting, PrivMeSA reduces the disclosure of personal details from 98.0% to 0.2% of cases and the share of cases in which the patient can be narrowed to ten or fewer registry patients from 74% to 0%.

Bullet Summary

  • PrivMeSA addresses the challenge of balancing patient privacy with the need for accurate clinical decision support by enabling locally deployed large language model (LLM) agents to selectively consult remote specialists without exposing sensitive patient in...
  • The system employs a local agent that manages each patient encounter and uses reinforcement learning to control information disclosure, optimizing a trade-off between task accuracy and privacy risks, including direct identifier disclosure and re-identificat...
  • Privacy risks are quantified using k-anonymity over a real patient registry to prevent patient re-identification across single and linked medical episodes, with the local agent trained adversarially against potential attackers.
  • PrivMeSA incorporates a lesson memory that distills the expertise gained through remote consultations into generalized clinical guidance, enabling future local reuse without repeated remote queries, thus reducing privacy risks and communication overhead.
  • Evaluations on an emergency department benchmark built from real hospital records (MIMIC-IV-ED) demonstrate that PrivMeSA improves mean task accuracy by up to 15.8 percentage points compared to baseline delegation methods.

Topological Coherence for Self-evolving Multi-agent Systems

Merged record merged scholarly record arXiv Memory Poisoning Governance and Policy Agent-to-Agent Communication

Sen Zhao, Ruiqi Kong, Zuyu Zhang, Lifeng Shen, Xinyu He, Xu Zhang, Qinghua Zhang

Published 2026-09-29

Venue: arXiv

Open Source Record

Abstract

Complex tasks inherently couple workflow structure, agent responsibility, collaboration, and memory access: task regions delimit responsibility and tool scope, cross-region dependencies give rise to handoffs, and ownership boundaries delimit private and selectively shared memory. Existing methods can jointly optimize agent and communication structures, yet such optimization does not by itself require responsibility, handoff, and memory boundaries to remain consistent with task dependencies. We term this requirement topological coherence. We introduce TOCOMAS, a Topology-Coherent Multi-Agent System. TOCOMAS grounds a task graph in tool interfaces, organizes compatible task nodes into reusable responsibility domains, and derives dependency-induced and profile-conditioned collaboration together with boundary-regulated memory visibility. During online self-evolution, TOCOMAS proposes coupled changes to agent, collaboration, and memory policies, retaining for subsequent tasks only candidates that satisfy structural constraints and improve evaluated reward. Across BBEH, WorkBench, SWE-Bench-Verified, and CoMemBench, TOCOMAS improves task success over baselines across backbones. CoMemBench also shows gains over the self-evolving baseline in verified progress, handoffs, and memory isolation.

Bullet Summary

  • Introduces TOCOMAS, a topology-coherent multi-agent system that enforces structural consistency among task workflows, agent responsibilities, collaboration interfaces, and memory boundaries to improve coordination and security.
  • TOCOMAS maps task graphs grounded in tool interfaces to organize compatible task nodes into reusable responsibility regions, enabling coherent execution without predefined agent roles.
  • The framework defines a responsibility quotient and quotient topology over task regions, preserving inter-region dependencies and guiding collaboration and communication edges among agents.
  • A boundary-preserving memory coordination system is proposed, associating memory records with task regions and agents to enforce visibility, provenance, and isolation during ownership changes.
  • TOCOMAS features a coherence-constrained structural evolution mechanism that jointly updates agent roles, collaboration, and memory policies based on execution feedback and reward optimization.

When Correct Memory Goes Wrong: Fuzzing Persistent Memory Use in LLM Agents

arXiv preprint arXiv Memory Poisoning

Yuqiao Meng, Luoxi Tang, Yingxue Zhang, Yuchen Yang, Zhaohan Xi

Published 2026-09-29

Venue: arXiv

Open Source Record

Abstract

Persistent memory helps LLM agents carry information across long interactions, but correct memory can still be used incorrectly when queries change or memory states evolve. Existing work mainly studies memory content errors or evaluates fixed test cases, leaving memory-use failures hard to discover systematically. We formulate this issue as a fuzzing problem and categorize such failures into query-related and memory-state failures. We then develop U-Fuzz, which starts from memory checkpoints as test seeds, mutates queries or memory states under explicit mutation obligations, validates each mutant, and uses observed memory behavior to guide iterative testing while keeping failure labels outside the search. We evaluate U-Fuzz across several memory systems against diverse fuzzing baselines, and further test an output-only setting with API-based LLMs where memory retrieval is hidden. Across these settings, U-Fuzz consistently uncovers more confirmed memory-use failures, showing that its search remains effective across different memory architectures and even when only final responses are observable.

Bullet Summary

  • The paper addresses failures in persistent memory usage by Large Language Model (LLM) agents, highlighting two failure types: query-related failures (incorrect memory use caused by altered or paraphrased queries) and memory-state failures (incorrect retriev...
  • Existing methods primarily examine memory content correctness or test fixed cases, leaving many memory-use failures undiscovered; this motivates a need for systematic and comprehensive testing approaches.
  • The authors formulate memory-use failure detection as a fuzzing problem and introduce U-FUZZ, a framework that mutates both queries and memory states to systematically uncover failures in diverse memory architectures of LLM agents.
  • U-FUZZ generates mutations from seeds based on memory checkpoints and applies different mutation strategies, including meaning-preserving query variations, target-changing mutations, unsupported queries, and memory-state mutations like updates, deletions, a...
  • The framework incorporates rigorous validation of mutants without relying on output labels, instead using operator-specific checks and retrieval behavior metrics (coverage, novelty, divergence) to guide guided iterative fuzzing.

VirusCascade: Hijacking Collaborative Reflection in LLM-Powered Recommender Agents

arXiv preprint arXiv Memory Poisoning Agent-to-Agent Communication Orchestration Risk

Yurong Hao, Wen Zhou, Guowei Guan, Tiantong Wu, Fuyao Zhang, Wei Yang Bryan Lim

Published 2026-09-29

Venue: arXiv

Open Source Record

Abstract

Advancing beyond traditional static scoring models, LLM-powered agentic recommender systems (LLM-ARS) instantiate users and items as autonomous agents, whose semantic states are dynamically refined through a recurrent process known as collaborative reflection. While this mechanism improves recommendation quality, it simultaneously introduces a systemic vulnerability: adversarial evidence injected into a single agent can be rationalised into a legitimate preference narrative, written back into memory, and propagated to other agents through interaction contexts. We term the local rationalisation process reflection laundering, and its system-wide escalation through collaborative reflection collaborative-reflection hijacking. Existing attacks on recommender systems, whether based on interaction-level data poisoning or text-level adversarial perturbations, assume static pipelines and thus cannot exploit this recurrent, multi-agent amplification pathway. To bridge this gap, we first conduct a controlled vulnerability analysis that establishes two exploitable properties underlying collaborative-reflection hijacking: reflective persistence and cross-agent propagation. Then building on these findings, we propose VirusCascade, the first black-box targeted promotion attack that jointly shapes semantic and structural attack surfaces: the former ensures the target item is naturally rationalised as satisfying broad user preferences, the latter positions it for system-wide propagation. Extensive experiments on four real-world datasets across diverse LLM-ARS architectures demonstrate that VirusCascade consistently achieves state-of-the-art targeted exposure under evaluated stealth constraints, reaching a mean E@20 of 0.384 and surpassing the strongest baseline by an absolute margin of +0.185.

Bullet Summary

  • LLM-powered agentic recommender systems (LLM-ARS) leverage autonomous agents with natural language memories to represent users and items, employing collaborative reflection to iteratively refine semantic states and improve recommendation quality.
  • Collaborative reflection, while beneficial, introduces a systemic vulnerability termed collaborative-reflection hijacking, wherein adversarial evidence injected into one agent is rationalized into legitimate preference narratives, written back into memory,...
  • Traditional attacks such as interaction-level data poisoning and text-level adversarial perturbations fail to exploit this vulnerability due to static assumptions, motivating a need for novel attacks targeting the recurrent, multi-agent reflection process.
  • The authors identify two critical properties enabling collaborative-reflection hijacking: reflective persistence, where injected semantic claims persist and are elaborated in memory over multiple reflection rounds; and cross-agent propagation, where manipul...
  • VirusCascade is proposed as the first black-box targeted promotion attack on LLM-ARS, combining semantic injection (embedding target items into preference motifs) and structural injection (crafting interaction trajectories) to shape both semantic and struct...

Data and scripts for: Access control and staff lifecycle in a multi-tenant memory server for AI agents: threat model, attack evaluation and two revocation bugs

Merged record merged scholarly record OpenAlex Memory Poisoning Governance and Policy Benchmarks and Evaluation

Vitalii Cherepanov

Published 2026-09-29

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.23040398

Open Source Record

Abstract

Raw data and scripts for the paper named in the title. A synthetic company (4 departments, 12 users, fixed seed 20260925) was loaded into the team server of total-agent-memory 14.5.1 (commit 95ea2b8). The dataset holds the benchmark harness, the generated data, raw JSONL results and environment files of all runs (attacks on department isolation, staff transfer and offboarding, concurrent edits and retries, latency and memory), the two fixes as patches (released in v14.6.0), and verify.py, which recomputes every number in the paper from the raw files. All people and facts are synthetic. No LLM was called.

Bullet Summary

  • The paper investigates access control and staff lifecycle management in a multi-tenant memory server designed for AI agents, focusing specifically on a threat model and attack evaluation.
  • A synthetic company setup comprising 4 departments and 12 users was created to simulate real-world scenarios within the total-agent-memory 14.5.1 team server environment.
  • Experiments included attacks targeting department isolation, staff transfers, offboarding procedures, concurrent edits, and retry mechanisms, as well as measurements of latency and memory usage.
  • The study uncovered two significant security vulnerabilities (revocation bugs) affecting access control during staff lifecycle events.
  • Comprehensive raw datasets, benchmark scripts, JSONL results, and environment files were provided to support reproducibility and transparency of the experimental results.

Share-Borne AI Virus: Memory-Hopping Attacks Across LLM Agents

arXiv preprint arXiv Memory Poisoning Agent-to-Agent Communication Orchestration Risk

Sidharth Pulipaka, Ansh Sharma, Stanislau Hlebik, Leonidas Raghav, Vyas Raina, Ivaxi Sheth, Mario Fritz

Published 2026-09-28

Venue: arXiv

Open Source Record

Abstract

Large language models are increasingly deployed as stateful assistants that retain information across interactions and use tools to read, modify, and create persistent artifacts. As these artifacts are shared between users, they form an indirect communication channel between otherwise independent assistants. We study a failure mode in which this channel enables self-propagating attacks. We introduce artifact-mediated propagation, where adversarial content introduced through an artifact (e.g. a report), is stored in an assistant's persistent memory, reproduced in a subsequently created artifact, and acquired by another assistant that later reads it. We evaluate this process in temporal human-agent universes that model artifact exchange between independently operated assistants over time, measuring whether an attack survives successive hand-offs, how many hops it reaches, and how broadly it spreads. We find that attacks can propagate across multiple independent assistants and persist over extended interaction sequences. In larger simulated environments, even GPT-5.6 Luna exhibits substantial spread, reaching 60-80% of agents with propagation chains extending to eight hops. These results show that persistent artifacts can act as durable carriers of adversarial state, allowing attacks to outlive individual interactions and spread across isolated assistants.

Bullet Summary

  • The paper introduces "artifact-mediated propagation," a novel self-propagating attack wherein adversarial content embedded in shared persistent artifacts (such as reports) infects an assistant's memory and is subsequently transmitted to other large language...
  • Experiments conducted in simulated "human-agent universes" demonstrate that these attacks can persist over multiple interactions and extensively spread across otherwise isolated AI assistants, with propagation chains reaching up to eight hops and infecting...
  • The attack exploits the indirect communication channel formed by artifacts shared between independent AI assistants, despite the absence of direct inter-agent communication and the presence of private memories in each assistant.
  • Two propagation variants are formalized: the endpoint-assisted variant, involving an external attacker-controlled service that helps preserve adversarial state in artifacts, and the prompt-only variant, which relies solely on embedding adversarial prompts w...
  • A comprehensive evaluation across multiple LLM models (GPT-5.6 Luna, Kimi-K2.6, GPT-OSS-120B, DeepSeek-V4-Pro) reveals varying susceptibility levels, with some models nearly completely infected and others showing partial resistance; the reproduction number...

CoSec: Benchmarking Agent Security in Communities

Merged record merged scholarly record arXiv Benchmarks and Evaluation Governance and Policy Memory Poisoning

Hao Chen, Wenhui Dong, Ye Chen, Jiezhi Yao, Chenbo Xia, Yuwen Qu, Renxiang Wang, Fudong Yuan

Published 2026-09-28

Venue: arXiv

Open Source Record

Abstract

LLM agents operate in persistent collaborative environments involving multiple users, communities, memories, files, and tools. Community boundaries may remain fixed or evolve with changes in membership, roles, composition, and relationships. Agents must complete legitimate tasks and prevent unauthorized disclosure of protected information. Existing evaluations do not fully examine these risks in agent systems. We introduce \textbf{CoSec}, an executable benchmark for evaluating privacy and authorization enforcement in LLM agent systems operating within and across communities. CoSec contains 208 canonical scenarios spanning fixed and evolving boundaries, protected information belonging to the agent owner or other participants, and attacks through dialogue, environmental content, persistent memory, and composed workflows. CoSec executes complete agent systems with persistent sessions, memory, files and tools. It verifies information flows against the active authorization state using execution traces and artifacts. Across harness and model configurations, agents frequently complete benign tasks but violate privacy and authorization boundaries. Privacy behavior varies across harnesses, attack surfaces, and community states, revealing how memory, files, tools, and workflows can carry protected information beyond its authorized scope. These findings show that task utility does not imply privacy or authorization compliance and that authorization in community settings remains an unresolved security challenge for persistent LLM agents.

Bullet Summary

  • Introduces CoSec, an executable benchmark with 208 scenarios designed to evaluate privacy and authorization enforcement in large language model (LLM) agents operating in persistent multi-agent community environments.
  • Models communities as persistent collaborative scopes with evolving boundaries, users, roles, data, and policies to simulate realistic multi-agent settings with changing memberships and relationships.
  • Defines an authorization model encoding states of users, roles, memberships, policies, and histories to determine access permissions, enabling evaluation of agent adherence to complex authorization constraints.
  • Evaluates multiple agent system configurations, revealing that while agents frequently complete benign collaborative tasks (high utility), they often violate privacy and authorization boundaries, exposing sensitive data improperly.
  • Identifies four main authorization failure modes caused by persistent states and stale workflows: excess permission, failed revocation of access, provenance confusion, and workflow carryover across community boundaries.

CoSec: Benchmarking Agent Security in Communities

arXiv preprint arXiv Benchmarks and Evaluation Governance and Policy Memory Poisoning

Hao Chen, Wenhui Dong, Ye Chen, Jiezhi Yao, Chenbo Xia, Yuwen Qu, Renxiang Wang, Fudong Yuan

Published 2026-09-28

Venue: arXiv

Open Source Record

Abstract

LLM agents operate in persistent collaborative environments involving multiple users, communities, memories, files, and tools. Community boundaries may remain fixed or evolve with changes in membership, roles, composition, and relationships. Agents must complete legitimate tasks and prevent unauthorized disclosure of protected information. Existing evaluations do not fully examine these risks in agent systems. We introduce \textbf{CoSec}, an executable benchmark for evaluating privacy and authorization enforcement in LLM agent systems operating within and across communities. CoSec contains 208 canonical scenarios spanning fixed and evolving boundaries, protected information belonging to the agent owner or other participants, and attacks through dialogue, environmental content, persistent memory, and composed workflows. CoSec executes complete agent systems with persistent sessions, memory, files and tools. It verifies information flows against the active authorization state using execution traces and artifacts. Across harness and model configurations, agents frequently complete benign tasks but violate privacy and authorization boundaries. Privacy behavior varies across harnesses, attack surfaces, and community states, revealing how memory, files, tools, and workflows can carry protected information beyond its authorized scope. These findings show that task utility does not imply privacy or authorization compliance and that authorization in community settings remains an unresolved security challenge for persistent LLM agents.

Bullet Summary

  • Introduces CoSec, a novel executable benchmark with 208 scenarios designed to evaluate privacy and authorization enforcement in LLM agents operating in persistent multi-agent communities with dynamic boundaries.
  • Models complex community settings involving fixed and evolving memberships, roles, relationships, and explicit collaboration scopes, enabling comprehensive testing of multi-agent security challenges.
  • Incorporates four main attack families—user injection, indirect injection, memory poisoning, and capability composition—to systematically test agents against diverse threat vectors.
  • Executes full agent systems with persistent sessions, memory, files, and tools; verification of information flows against active authorization policies is achieved using execution traces and semantic analysis.
  • Experiments across multiple agent harnesses and backbone models demonstrate frequent privacy and authorization failures despite successful benign task completion, revealing that task utility alone does not ensure security compliance.

Pwnagent: a knowledge-guided multi-agent system for automatic exploit generation

Merged record merged scholarly record OpenAlex Memory Poisoning Benchmarks and Evaluation Agent-to-Agent Communication

Chaojie Wei, Yangyang Geng, Yunfeng Wang, Qilong Wu, Jing Huang, Qianqiong Wu, Qiang Wei

Published 2026-09-28

Venue: Cybersecurity

DOI: https://doi.org/10.1186/s42400-026-00649-5

Open Source Record

Abstract

Abstract Automatic Exploit Generation (AEG) plays an important role in proactive assessment of software threats by identifying vulnerabilities and constructing functional payloads. Existing Large Language Model (LLM)-based methods, however, often struggle to reason about complex exploit logic and to perform runtime introspection, leaving a gap between static vulnerability analysis and dynamic memory behavior. We present PwnAgent, an LLM-driven multi-agent framework for end-to-end exploit generation that combines offensive domain knowledge with active runtime introspection. PwnAgent uses a hierarchical knowledge base for multi-stage exploit reasoning and a feedback-driven self-correction engine to calibrate dynamic memory parameters during execution. Because broad Capture The Flag (CTF) benchmarks offer limited binary-exploitation depth and pwn-specific evaluation must balance reproducibility, difficulty progression, and exploit diversity, we construct a 66-task pwn benchmark from public CTF-style challenges. The benchmark is primarily composed of Linux x86/x86-64 ELF binaries and stack-oriented tasks, with smaller format-string, heap, integer-overflow, ARM, and MIPS subsets used as limited probes beyond the dominant setting. Under the same recent Kimi-K2.6 backend, PwnAgent achieves a 62.12% end-to-end success rate, compared with 31.82% for the evaluated PwnGPT baseline, a 30.30 percentage-point gain. These paired results indicate that structured knowledge guidance, execution-grounded measurement, and feedback repair improve LLM-based exploit generation in the evaluated setting, while the absolute success rate shows that fully autonomous exploitation remains challenging.

Bullet Summary

  • The paper addresses the challenge of automatic exploit generation (AEG) to proactively assess software vulnerabilities and construct functional exploits.
  • Existing large language model (LLM)-based approaches struggle with reasoning about complex exploit logic and runtime introspection, limiting their effectiveness.
  • PwnAgent is proposed as an LLM-driven multi-agent system that integrates offensive domain knowledge with active runtime introspection for end-to-end exploit generation.
  • The system employs a hierarchical knowledge base for multi-stage exploit reasoning and a feedback-driven self-correction engine to iteratively calibrate memory parameters during execution.
  • To evaluate, a customized 66-task benchmark was constructed from public Capture The Flag (CTF)-style challenges focusing primarily on Linux x86/x86-64 ELF binaries and stack-oriented exploits.

BaRe-Mem: Bayesian Reliability Memory for Robust and Adaptive Agent Consultation

Semantic Scholar · Semantic Scholar scholarly work Semantic Scholar Trust and Identity Memory Poisoning Benchmarks and Evaluation

Pei-Lin Feng, Zheng-Yang Huang, Soujanya Poria

Published 2026-09-28

Venue: Semantic Scholar

Open Source Record

Abstract

In multi-agent systems, reliable consultation is challenging because advisor capabilities vary across tasks, and misleading information can make consultation worse than autonomous reasoning. We introduce BaRe-Mem, an online Bayesian reliability memory for multi-agent consultation. It estimates advisor reliability based on the central model's internal belief representations and updates these estimates from historical interactions. These estimates modulate the influence of advisor responses and guide the choice between consultation and autonomous reasoning. Across nine benchmarks and six central models, BaRe-Mem is more robust to misleading advisor information than debate and majority voting. On the more challenging tasks, it remains above autonomous reasoning across all tested misleading levels. Moreover, we extend the BaRe-Mem mechanism to worker allocation in agent teams. On the MuSiQue benchmark, BaRe-Mem improves task completion over routing by historical success counts and identifies capable workers earlier.

Bullet Summary

  • Multi-agent consultation is challenging due to varying advisor capabilities and the risk that misleading advice can degrade performance compared to autonomous reasoning.
  • BaRe-Mem is introduced as an online Bayesian reliability memory system designed to estimate and track the reliability of advisors based on the central model's internal belief representations and past interactions.
  • BaRe-Mem dynamically adjusts the influence of advisor responses on the central model by updating reliability estimates from historical data, enabling informed decisions on when to consult advisors or rely on autonomous reasoning.
  • The approach was evaluated across nine benchmarks and with six different central models, demonstrating higher robustness to misleading advisor information than popular methods such as debate and majority voting.
  • On especially challenging tasks and under various levels of misinformation, BaRe-Mem consistently outperformed autonomous reasoning alone, indicating strong resilience to poor-quality advice.

TRACE: Governing Memory Validity in Evolving Multi-Agent Systems

Merged record merged scholarly record arXiv Governance and Policy Memory Poisoning

Wenjun Xiong, Shengtao Zhang, Shangding Gu, Bo Tang, Zhiyu Li, Feiyu Xiong, Ying Wen, Muning Wen

Published 2026-09-27

Venue: arXiv

Open Source Record

Abstract

Persistent memory lets language-model agents carry information across long-running collaborations, but leaves a lifecycle question open: what may a returning agent still act on once the shared state has changed? A memory can be correctly retrieved, relevant to the current task, and faithful to its source, and nonetheless be inadmissible for action: an itinerary saved before a pause still names the hotel the team has since replaced. We formalize this as temporal memory admission and present TRACE, a training-free layer that treats re-entry as an eligibility decision rather than a storage or retrieval operation, reconciling a departure checkpoint against absence-period updates, resolving explicit and implicit invalidation, and releasing a bounded Return View only when it covers the returning role's open obligations. We evaluate TRACE under three actor models on Memora, STALE Type II, and a derived ManBench-Return setting, each recast as return episodes: one agent departs, four teammates change the shared state, and the agent rejoins. What separates methods is not overall accuracy but whether one can retain valid memory and reject stale memory at once, and no single-policy baseline can: Restore (reinstate the departure checkpoint in full) admits stale state, Reset (start the return from an empty memory) discards valid state, each bottoming out at 0% on one of the two. TRACE is the only method high on both, reaching 92.6-98.3% valid-information availability with 98.4-99.5% invalid-information rejection on ManBench-Return, within 3.8 points of the best baseline's overall accuracy. On STALE Type II it improves Overall over the strongest comparison policy by 22.3 (Qwen), 18.5 (Gemini), and 27.5 (DeepSeek) points at roughly 2.3 times their tokens, while a write-time consolidation pipeline is more accurate still at 3.99 times TRACE's.

Bullet Summary

  • The paper addresses the challenge of determining admissible persistent memory for returning agents in evolving multi-agent systems, where shared state may have changed during their absence, causing some memories to become stale or invalid.
  • It formalizes 'temporal memory admission,' distinguishing between memories that are historically accurate, currently relevant, and admissible based on role and obligations at return time.
  • TRACE is introduced as a training-free system layer that reconciles an agent's checkpointed memory at departure with updates made by teammates during their absence, resolving explicit and implicit invalidations before selectively releasing only eligible mem...
  • The method applies multiple validation gates—authorization, freshness, applicability, provenance—on each memory item, producing an auditable and transparent trail that governs memory admission policies at re-entry.
  • Evaluation across three benchmarks (Memora, STALE Type II, ManBench-Return) and various actor models shows TRACE uniquely balances retaining valid information and rejecting stale memory, outperforming baseline methods on both metrics.

Adversarial Evidence-Plane Robustness for Bounded Agent Assurance: Deterministic Mutation Testing of Claim Admissibility

Semantic Scholar · Computers scholarly work Semantic Scholar Memory Poisoning Benchmarks and Evaluation Governance and Policy

Robert Campbell

Published 2026-09-27

Venue: Computers

DOI: 10.3390/computers15100655

Open Source Record

Abstract

Assurance for autonomous agents depends on operational evidence about policy decisions, actions, execution, effects, provenance, and enforcement. This study tests whether bounded assurance claims remain semantically admissible when that evidence is adversarially degraded. Three co-primary questions address deviation existence (B1), observable action-path reconstruction (B2), and containment/enforcement-boundary localization (B3). A deterministic mutation harness applied 14 frozen single-operator evidence-plane attacks to three known-ground-truth baselines, producing 42 adversarial cases and 126 claim evaluations. A prospectively frozen strongest-safe oracle and closed-world semantic scorer classified the 126 claim evaluations into three predefined categories: exact safe, safe conservative, and unsafe false establishment. All 42 cases were valid. Of the 126 claim evaluations, 56 were exact safe, 47 safe conservative, and 23 unsafe false establishment. The resulting Unsafe False Establishment Rate (UFER) was 23/126 ≈ 0.18254, so the pre-specified zero-UFER target failed. Unsafe outcomes were B1 3/42, B2 19/42, and B3 1/42. Exploratory forensics classified 21 unsafe evaluations as substantive semantic incompatibilities and two as possible scorer-normalization artifacts. The dominant B2 mechanism was global ineligibility after localized evidentiary degradation. These results distinguish semantic safety from information retention: evidence defects must not induce propositions beyond the boundary justified by surviving admissible evidence.

Bullet Summary

  • The paper addresses the robustness of bounded assurance claims for autonomous agents under adversarial degradation of operational evidence, focusing on policy decisions, actions, execution, effects, provenance, and enforcement.
  • Three central questions are explored: detection of deviation existence (B1), reconstruction of observable action-paths (B2), and localization of containment or enforcement boundaries (B3).
  • A deterministic mutation testing framework is employed, applying 14 fixed single-operator evidence-plane attacks to three ground-truth baselines, creating 42 adversarial cases and 126 total claim evaluations.
  • Evaluation is conducted using a frozen strongest-safe oracle and a closed-world semantic scorer, categorizing claims as exact safe, safe conservative, or unsafe false establishment.
  • Results show that while all adversarial cases were valid, 23 out of 126 claim evaluations were unsafe false establishments, leading to an Unsafe False Establishment Rate (UFER) of approximately 18.3%, failing the targeted zero-UFER benchmark.

Attack Surface Proliferation in Agentic AI: How Tool Calls, Memory Persistence, Skill Ecosystems, and Payment Rails Compose Into a Systemic Threat

Merged record merged scholarly record OpenAlex Prompt Injection Memory Poisoning Agent-to-Agent Communication

Saluca Agentic AI Research Team

Published 2026-09-26

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.20525492

Open Source Record

Abstract

This version (2026-09-26) corrects a citation error found by an automated check and confirmed by hand against the arXiv abstracts. Version 2 cited the identifier 2605.30998 for the ClawTrojan multi-step trojan benchmark and its DASGuard defense; that identifier is a security analysis of x402 payments, and ClawTrojan is arXiv:2605.31042. All six trojan citations are corrected. The x402 payment-layer claims keep arXiv:2605.30998, which is right for them, and the version 2 caveats warning that the two topics shared one identifier are replaced with what that paper's abstract reports. An unsupported phrase ("or redirect funds") and a pointer to internal drafting instructions are removed. The thesis is unchanged. This version has not had a full claim-by-claim audit. The full list of corrections is at the top of the PDF. The deployment of large language model (LLM) agents into production environments has outpaced the security frameworks designed to contain them. This paper argues a candidate structural pattern worth investigating: the attack surface of agentic AI systems is not additive across components but compositional, meaning that tool invocation, persistent memory, third-party skill ecosystems, and machine-to-machine payment rails each introduce independent vulnerabilities that compound when combined in a single agent harness. This is a heuristic reading, not a formal derivation; the mechanism we identify is the chaining of trust assumptions across subsystems that were designed and evaluated independently. We synthesize findings from seven recent preprints spanning cs.CR, covering: speculative tool-call privacy leakage before commitment arXiv:2606.02483, indirect prompt injection through enterprise SaaS integrations arXiv:2606.02240, multi-step trojan persistence in local agentic workspaces arXiv:2605.31042, memory poisoning via dialogue interaction arXiv:2605.29960, skill marketplace malware distribution arXiv:2605.28588, machine-to-machine payment protocol vulnerabilities in x402 arXiv:2605.30998, and coordinated multi-agent sabotage arXiv:2605.29178. A supplementary source on attribute-based access control for tool-use agents arXiv:2605.28071 provides a candidate defense framing. The attacker model throughout is a motivated external adversary with read access to at least one integration endpoint or skill marketplace listing, who exploits the trust transitivity inherent in composed agentic pipelines. The primary falsification path: if deploying layered, cross-subsystem monitoring (covering speculative calls, skill ingestion, memory writes, and payment state) reduces the compound attack success rate to that of the weakest single-subsystem baseline, the compositionality claim fails and the threat is merely additive. Authorship: Saluca Agentic AI Research Team (Saluca LLC). AI-drafted synthesis from an arXiv preprint corpus, originally drafted 2026-06-03, produced under the direction of Cristian Ruvalcaba, the accountable human author. Not peer-reviewed. AI disclosure. This work was produced with an agentic AI research apparatus operated by Saluca Labs. The apparatus drafted, searched and analysed under direction. Cristian Ruvalcaba is the human author and is accountable for the content. No AI system is listed as an author or contributor, because authorship entails accountability that a model cannot hold; this disclosure is the credit, and it is deliberately the whole of it.

Bullet Summary

  • The paper addresses the emerging security challenges posed by deploying large language model (LLM) agents in production, highlighting that existing security frameworks lag behind these deployments.
  • It proposes that the attack surface for agentic AI systems is compositional rather than additive; vulnerabilities in tool calls, memory persistence, third-party skill ecosystems, and machine-to-machine payment rails combine in complex ways, compounding secu...
  • The compositionality arises from chained trust assumptions across independently designed and evaluated subsystems, which attackers can exploit to escalate threats.
  • The paper synthesizes recent findings from seven preprints covering a range of vulnerabilities, including speculative tool-call privacy leaks, indirect prompt injections via SaaS, multi-step trojan persistence, memory poisoning, malware in skill marketplace...
  • The attacker model assumes a motivated external adversary with read access to at least one integration or skill endpoint, exploiting the trust transitivity in multi-agent AI systems.

Runtime Integrity Without Attestation: How Agent-Facing Tool Surfaces, Speculative Dispatch, and Memory Poisoning Converge on a Shared Execution-Layer Threat Model

Merged record merged scholarly record OpenAlex Memory Poisoning Governance and Policy Orchestration Risk

Saluca Agentic AI Research Team

Published 2026-09-26

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.20579699

Open Source Record

Abstract

Version 3 (2026-09-26). Version 3 corrects errors found by an independent audit of version 2 and by a second review of the corrected draft. (1) The ClawTrojan/DASGuard work (95.5% attack success rate) was cited as arXiv:2605.30998, which is an unrelated paper on x402 payments. The correct source is arXiv:2605.31042 (Tan et al.). All citations, including this record's list of cited preprints, have been corrected; the figures were accurate to the correct paper. (2) An unfounded remark that "GPT-5.4" was an unverifiable internal label has been removed. (3) The Recuse Signal result is updated to the current v4 of arXiv:2606.06460. Recusal at the access door is model-dependent, ranging from 100% down to 55-75%, and a mid-task in-band halt is weaker. The v1 figure of 100% recusal is superseded. (4) Overstated readings of the ClawTrojan and MemPoison abstracts have been corrected. The claim that agents hold "OAuth tokens" is replaced with what arXiv:2606.06460 supports (real credentials used over SSH). The SCHEME monitoring figure is given as 100% (Gemini) and 81% (Codex) rather than "near-100%". (5) Two cross-paper generalizations are narrowed. The first is that all six attack surfaces exploit temporal decoupling of authorization and execution, which no source states in those terms and which is now presented as this synthesis's reading. The second is that all succeed against correctly aligned agents, and SCHEME is flagged as a poor fit for the attacker model. The thesis is unchanged. The file has a neutral name, and the full list of corrections is at the top of the PDF. Version 2 was revised in response to an external structural review and an automated critique pass; its change log is kept as the "Response to Review" appendix in the PDF. Autonomous LLM agents have moved from sandboxed demos into production infrastructure: they hold live credentials, invoke real APIs, write to persistent memory, and execute multi-step workflows without a human in the loop. This paper argues that a structurally coherent threat class has emerged at the execution layer (the runtime surface between an agent's reasoning process and the external environment it acts upon) that is distinct from the prompt-injection and jailbreak threats that dominate the safety literature. We synthesize seven findings from recent cs.CR preprints (six attack surfaces and one cooperative contrast case; a supporting monitoring paper is discussed in an addendum) to support a single thesis: execution-layer attacks succeed not by defeating safety alignment but by exploiting the gap between what an agent is authorized to do and what the runtime environment can verify it is actually doing. The attacker model throughout is a capability-constrained adversary who controls some portion of the environment the agent reads from or writes to (a third-party script, a tool-registration endpoint, a RAG routing profile, a skill file, a memory pipeline) but does not need to compromise the model weights, the system prompt, or the user. Sources span cooperative governance signals arXiv:2606.06460, tool-surface poisoning arXiv:2606.06387, speculative-dispatch privacy arXiv:2606.02483, multi-step trojan persistence arXiv:2605.31042, memory poisoning via dialogue arXiv:2605.29960, coordinated multi-agent sabotage arXiv:2605.29178, and federated RAG routing hijacking arXiv:2605.28112. The falsification path for the central thesis is concrete: if any of these attacks can be neutralized by a purely model-level intervention, safety fine-tuning, RLHF, or constitutional AI, without runtime changes, the execution-layer framing is wrong. The abstracts reviewed do not support that conclusion; we argue the evidence points the other direction. This is a heuristic reading across independently motivated papers, not a derivation from a shared formal structure; the convergence identified is interpretive, not mathematical. Authorship: Saluca Agentic AI Research Team (Saluca LLC). AI-drafted synthesis from an arXiv preprint corpus, originally drafted 2026-06-07, produced under the direction of Cristian Ruvalcaba, the accountable human author. Not peer-reviewed. Cited arXiv preprints: arXiv:2605.28112, arXiv:2605.29178, arXiv:2605.29960, arXiv:2605.31042 (corrected in v3; v2 listed 2605.30998 in error), arXiv:2605.31593, arXiv:2606.02483, arXiv:2606.06387, arXiv:2606.06460 AI disclosure. This work was produced with an agentic AI research apparatus operated by Saluca Labs. The apparatus drafted, searched and analysed under direction. Cristian Ruvalcaba is the human author and is accountable for the content. No AI system is listed as an author or contributor, because authorship entails accountability that a model cannot hold; this disclosure is the credit, and it is deliberately the whole of it.

Bullet Summary

  • Autonomous large language model (LLM) agents have transitioned from controlled demos to production systems, executing multi-step workflows with live credentials and real APIs without human oversight.
  • A new coherent class of threats, termed execution-layer attacks, has emerged; these attacks exploit vulnerabilities at the runtime interface between an agent's reasoning process and its external environment, distinct from traditional prompt-injection or jai...
  • Seven recent preprints are synthesized to support the thesis that execution-layer attacks succeed by exploiting the gap between agent authorization and verifiability by the runtime environment, rather than defeating the safety alignment mechanisms themselves.
  • The attacker model assumes a capability-constrained adversary who can control parts of the agent's runtime environment (e.g., scripts, tool endpoints, routing profiles) but cannot compromise core model parameters, system prompts, or end users.
  • Significant attacks reviewed include tool-surface poisoning, speculative dispatch privacy exploits, multi-step trojan persistence, memory poisoning via dialogue, multi-agent sabotage, and federated retrieval-augmented generation (RAG) routing hijacking.

The Agentic Trust Stack Has No Bottom: Why Every Layer From PRNG to Payment Rail Is a First-Class Attack Surface

Merged record merged scholarly record OpenAlex Trust and Identity Prompt Injection Memory Poisoning

Saluca Agentic AI Research Team

Published 2026-09-26

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.20519743

Open Source Record

Abstract

This version corrects a citation error found by an automated check and confirmed by hand. Section 4.5 cited arXiv:2605.29963 (Honeyval, an evaluation framework for LLM-powered HTTP honeypots) for its evidence on inference-system fingerprinting; the passage describes arXiv:2605.29979, "Fingerprinting Inference Systems of Large Language Models", and both citations now point there. The wording of the claims already matched that paper and is unchanged. Typographic dashes are also removed. This version has not had a full claim-by-claim audit. The correction note is at the top of the PDF. The dominant framing of agentic AI security treats each threat in isolation: prompt injection here, memory poisoning there, jailbreak somewhere else. This paper argues that framing is structurally wrong. The corpus reveals what we read as a coherent vertical threat surface, what we call the agentic trust stack, in which an attacker who compromises any single layer can propagate damage upward and downward through the entire pipeline without triggering any single layer's defenses. We synthesize seven specific findings spanning: (1) supply-chain attacks on cryptographic watermarking primitives (arXiv:2605.28632), (2) speculative tool-call leakage before any authorization decision is made (arXiv:2606.02483), (3) multi-step trojan persistence through workspace state (arXiv:2605.31042), (4) coordinated multi-agent covert sabotage (arXiv:2605.29178), (5) financial-rail atomicity failures in machine-to-machine payment protocols (arXiv:2605.30998), (6) agent-skill marketplace contamination with confirmed malicious payloads (arXiv:2605.28588), and (7) LLM billing fraud enabled by auditor trust paradoxes (arXiv:2605.30040). The thesis is: agentic pipelines are not merely vulnerable at their endpoints; they are vulnerable at every trust delegation boundary, and those boundaries are currently neither enumerated nor defended as a class. This thesis is a heuristic reading of the corpus, not a formally derived result; the attack chain described is a structural argument, not an empirically demonstrated end-to-end exploit. The falsification path is direct: a single deployed agentic system that (a) enumerates all trust delegation boundaries in its execution graph, (b) enforces independent attestation at each, and (c) demonstrates that no cross-layer attack chain survives, would falsify the claim that the stack has no defensible bottom. No such system is documented in the corpus. Authorship: Saluca Agentic AI Research Team (Saluca LLC). AI-drafted synthesis from an arXiv preprint corpus, originally drafted 2026-06-02, produced under the direction of Cristian Ruvalcaba, the accountable human author. Not peer-reviewed. AI disclosure. This work was produced with an agentic AI research apparatus operated by Saluca Labs. The apparatus drafted, searched and analysed under direction. Cristian Ruvalcaba is the human author and is accountable for the content. No AI system is listed as an author or contributor, because authorship entails accountability that a model cannot hold; this disclosure is the credit, and it is deliberately the whole of it.

Bullet Summary

  • The paper challenges the conventional approach in agentic AI security that considers threats in isolation and proposes a new unified model termed the 'agentic trust stack'.
  • It argues that agentic AI systems have a vertically integrated threat surface in which compromising a single layer allows attackers to propagate damage both upward and downward through the pipeline without layer-specific defenses detecting it.
  • Seven specific vulnerabilities are synthesized from the literature, spanning areas such as supply-chain attacks on cryptographic watermarking, speculative tool-call leakage, workspace state trojan persistence, coordinated multi-agent sabotage, financial pro...
  • The core thesis is that vulnerabilities exist at every trust delegation boundary in agentic pipelines, not just at individual endpoints, and these boundaries are currently not enumerated or defended comprehensively.
  • The paper does not provide an empirically demonstrated end-to-end exploit but offers a structural argument for the presence of pervasive cross-layer attack chains in agentic AI pipelines.

BMA: Backchain Memory Attacks Create Unauthorized Control Paths in LLM Agents

Semantic Scholar · Semantic Scholar scholarly work Semantic Scholar Memory Poisoning Trust and Identity Governance and Policy

Kai-Sheng Fan, Yi-Shu Gao, Xun-Zhu Tang, Tegawendé F. Bissyandé, Wei-Zhe Zhang

Published 2026-09-26

Venue: Semantic Scholar

Open Source Record

Abstract

Persistent memory enables LLM agents to reuse prior experience, but creates a new security boundary: what an agent may remember is not what it should act on. We expose an unauthorized control path where edited low-trust evidence is consolidated into persistent memory, retrieved on a clean task, and used to drive a protected action. Crucially, the adversary neither writes memory nor alters the task. We introduce Backchain Memory Attack (BMA), a grey-box, LLM-driven inverse-planning attack that reasons backward from the target action to the memory that would trigger it, then to the evidence edit that would form it. BMA has two phases: preparation uses resettable trials to localize the first failed link and build experience; execution uses the frozen experience to rank and commit candidate edits without feedback. We introduce the Pathway-Certified Attack Success Rate (Path-CASR) to separate memory-mediated from coincidental hits: registered memory must form, be retrieved, drive the target behavior, and pass matched-intervention checks. Across four substrates and three decision backbones, BMA achieves 18.8% Macro Path-CASR, compared with 13.4% for the strongest access-matched baseline. Of BMA's behavioral hits, 60.3% pass all registered pathway and intervention checks versus 36.7% for the baseline. Frozen BMA edits retain 78.0% of their certified effect on average across four held-out consolidation policies. Representative memory-side controls leave 11.0% Path-CASR, whereas provenance-bound authorization reduces it to 2.0% while preserving 92.1% legitimate-action success.

Bullet Summary

  • Persistent memory in large language model (LLM) agents enables experience reuse but introduces a new security risk where unauthorized control actions can be triggered by modified memory content.
  • The paper identifies an attack vector where low-trust evidence edits consolidate into persistent memory and later trigger protected actions without direct memory writing or task alteration by the adversary.
  • A novel attack method called Backchain Memory Attack (BMA) is proposed, which uses an inverse-planning approach driven by an LLM to reason backward from the target action to memory and evidence edits needed to induce it.
  • BMA operates in two stages: preparation uses resettable trials to locate weak points and build experience, and execution ranks and applies candidate edits using the learned knowledge without feedback.
  • The authors introduce the Pathway-Certified Attack Success Rate (Path-CASR) metric to rigorously differentiate genuine memory-mediated attacks from chance occurrences, requiring demonstration of memory formation, retrieval, behavioral impact, and interventi...

Learning from Others, Acting for You: Cross-User Memory Sharing for LLM Agents

Semantic Scholar · Semantic Scholar scholarly work Semantic Scholar Memory Poisoning Governance and Policy Agent-to-Agent Communication

Jinming Hu, Haodong Zhao, Qi Jia, Die Chen, Tian Zhao, Su-Feng Duan, Gong-Shen Liu

Published 2026-09-26

Venue: Semantic Scholar

Open Source Record

Abstract

Large language model (LLM) agents serving different users often solve related tasks, yet separate user histories can leave reusable experience inaccessible to other agents. Pooling memories expands access but risks transferring preferences that conflict with the receiving user's requirements. We introduce ShareMem, a memory architecture that shares reusable experience while grounding its application in the receiving user's own preferences. Shared experiences indicate how to act and which preferences to consult; the receiving user's memory supplies their concrete values. Two-stage consolidation refines experience locally before integrating accepted edits into a shared pool. During execution, scope-first retrieval jointly selects local and shared experiences under a common entry budget, while a user-bound channel supports initial and agent-initiated preference retrieval. We evaluate ShareMem across web navigation (Mind2Web), online personalized interaction (VitaBench~2.0), and multi-session coding (MemoryCode) with four backbone models. It improves step success, average task success, and dialogue-macro coding scores, respectively, over matched user-local memory across all four models. Ablations favor two-stage consolidation for smaller shared pools, lower induction token usage, and better downstream performance, and support complementarity between experience guidance and active preference retrieval. Further analyses show that sharing helps most when relevant local experience is scarce, while source quality and cross-user preference interference limit useful transfer.

Bullet Summary

  • Problem: Large language model (LLM) agents serving different users have separate user histories, preventing access to reusable experience across agents and risking conflicting preferences with pooled memories.
  • Method: Introduction of ShareMem, a memory architecture that shares reusable experiences by indicating actions and preferences to consult, grounded in the receiving user's concrete preferences through their local memory.
  • Experience Refinement: ShareMem employs two-stage consolidation, refining experience locally before integrating accepted edits into a shared memory pool to ensure quality and relevance.
  • Memory Retrieval: Utilizes scope-first retrieval to jointly select relevant local and shared experiences under a common entry budget, alongside a user-bound channel for preference retrieval during execution.
  • Evaluation: Tested on web navigation (Mind2Web), online personalized interaction (VitaBench 2.0), and multi-session coding (MemoryCode) using four backbone LLM models.

AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents

arXiv preprint arXiv Orchestration Risk Memory Poisoning Benchmarks and Evaluation

Weida Liang, Shi Qiu, Zhun Wang, Simon Sure, Xiaoyuan Liu, Tianneng Shi, Zhaorun Chen, Wenbo Guo

Published 2026-09-25

Venue: arXiv

Open Source Record

Abstract

AI agents combine language models with external data and tools that can modify files, call APIs, or execute code. Security failures can arise when adversarial content changes an agent's tool use or when the surrounding software contains vulnerabilities such as path traversal or command injection. We study authorized white-box pre-deployment auditing, where the auditor has access to the target repository and a controlled runtime, but successful attacks must still act through the task-defined attacker interface and be confirmed by an external verifier. We present AgentXploit, a two-role auditing system that separates repository-level attack-path discovery from runtime exploitation. The Analyzer Agent traces attacker-controlled inputs to sensitive operations and records code-supported candidate attack paths; the Exploiter Agent turns these paths into concrete attacks and revises them using runtime feedback. We also introduce AgentXploit-Bench, containing 72 reproducible vulnerabilities across 12 open-source AI-agent systems and frameworks. Across three runs, AgentXploit reaches 59.3% end-to-end success, compared with 38.4% for Codex. Under a token-budget-matched comparison, Codex reaches 46.3%. On AgentDojo, where injection points are provided, the Exploiter Agent reaches 79.2% attack success versus 52.7% for AgentVigil. These results highlight repository discovery and runtime exploitation as distinct challenges in end-to-end agent security auditing.

Bullet Summary

  • The paper addresses security risks in AI agents that integrate language models with external tools capable of modifying files, calling APIs, or executing code, which can be exploited via adversarial inputs or software vulnerabilities.
  • AgentXploit is proposed as a novel two-role system separating repository-level attack-path discovery (Analyzer Agent) from runtime attack execution and feedback-driven refinement (Exploiter Agent) for authorized white-box pre-deployment auditing of AI agents.
  • AgentXploit-Bench is introduced as a comprehensive benchmark consisting of 72 reproducible vulnerabilities across 12 open-source AI-agent frameworks, enabling realistic evaluation without supplied injection points or metadata.
  • Experiments show AgentXploit achieves a 59.3% end-to-end attack success rate, outperforming Codex (38.4%) and AgentVigil (52.7%), especially on complex indirect attack paths, highlighting distinct challenges in discovery versus runtime exploitation.
  • The auditing process formalizes the need to jointly analyze dataflow, control flow, language model decisions, and tool preconditions to identify feasible attack paths leading to sensitive operations.

agmi: Agent Memory Integrity, a conformance test suite for tamper evidence in AI agent memory and checkpoint stores

OpenAlex · Zenodo (CERN European Organization for Nuclear Research) repository OpenAlex Memory Poisoning Prompt Injection Benchmarks and Evaluation

Khandelwal Yasha

Published 2026-09-25

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.22860886

Open Source Record

Abstract

0.6.0: the eight-edit at-rest scorecard, implementing the test method proposed for IETF draft-han-bmwg-agent-security-benchmark metric 5.4.7 (bmwg list, 24 September 2026). agmi (Agent Memory Integrity) is an open, MIT-licensed conformance test suite that measures whether AI agent memory and checkpoint stores notice when they are tampered with. It seeds a store through the tool's own API, edits the stored records behind the tool's back, restarts the tool, and reads its memory again through the tool's own read path. The verdict is what the tool does: rejected (it refused the edit on read), reported (its audit named it), or accepted (it served the edit as genuine). Eight storage-level edits: T1 content tamper, T2 tail truncation, T3 middle deletion, T4 reordering, T5 forged insertion, T6 cross-context replay, T7 rollback replay, T8 metadata tamper. T6 and T7 use only bytes the store itself wrote, in the wrong place; they are the edits that separate encryption from integrity. A self-validating reference store catches all eight, so every finding is proven against a store built to catch it. Measured in this release on LangGraph SqliteSaver, Letta block checkpoint history, Mem0 local Qdrant and inspeximus: all four accept all eight edits in their default configuration. inspeximus with receipts on catches rollback and metadata edits but still accepts cross-context replay, because a receipt binds a record's text and key, not the user it belongs to. The suite also measures six front-door attacks through the tool's own write path (memory injection, cross-session bleed, retrieval hijack, indirect prompt injection, update poisoning, metadata poisoning) on three attacker channels with content-evasion mutations, and ships agmi-check plus a GitHub Action: one CI step that fails on any accepted edit. Results are regenerated from a committed results file and reproduced in CI on every change. Code: https://github.com/tech4biz-yasha/agmiRelease: https://github.com/tech4biz-yasha/agmi/releases/tag/v0.6.0PyPI: https://pypi.org/project/agent-memory-integrity/0.6.0/

Bullet Summary

  • Introduces agmi (Agent Memory Integrity), an open-source, MIT-licensed conformance test suite designed to evaluate tamper evidence in AI agent memory and checkpoint storage systems.
  • The test suite operates by seeding a store via the tool's API, covertly modifying the stored records, restarting the tool, and then reading the memory through the original read path to assess integrity.
  • Defines eight distinct storage-level tampering edits tested: T1 content tamper, T2 tail truncation, T3 middle deletion, T4 reordering, T5 forged insertion, T6 cross-context replay, T7 rollback replay, and T8 metadata tamper; with T6 and T7 focusing on integ...
  • Demonstrates a self-validating reference store that detects all eight tampering types, serving as a benchmark to validate findings against proven detection mechanisms.
  • Evaluated four storage implementations (LangGraph SqliteSaver, Letta block checkpoint history, Mem0 local Qdrant, and inspeximus) which, by default, accepted all eight tampering edits, indicating security weaknesses.

LLM Agents Can Easily Tamper With Their Own Traces

Semantic Scholar · Semantic Scholar scholarly work Semantic Scholar Orchestration Risk Memory Poisoning Governance and Policy

Jeremy Qin, David Schmotz, Derck W. E. Prinzhorn, Luca Beurer-Kellner, Ameya Prabhu, Maksym Andriushchenko

Published 2026-09-24

Venue: Semantic Scholar

Open Source Record

Abstract

Asynchronous monitoring, incident investigations, and compliance audits primarily rely on agent traces to reconstruct what happened. These analyses assume that LLM agents cannot tamper with their own execution traces. We show that local LLM agents such as Claude Code, Codex, Antigravity, Open Code and Grok Build fail to enforce this boundary. All tested harnesses, except Muse Code, allowed agents to delete their traces when asked, without triggering monitor guardrails. We also validate that external attackers can exploit this gap to induce trace deletion. Finally, we show that trace tampering behavior emerges naturally in frontier models, when agents try to improve their rewards. We advise practitioners to ensure trace logging happens through an independent interception mechanism outside of the agent's control, preserving trace integrity even in cases of full host compromise. Overall, our findings identify a concrete failure of trace integrity in agent infrastructure which can be used to conceal misaligned behaviors like scheming or sabotage.

Bullet Summary

  • Agent execution traces are crucial for asynchronous monitoring, incident investigations, and compliance audits to reconstruct agent behavior.
  • The prevailing assumption is that large language model (LLM) agents cannot tamper with their own execution traces, ensuring reliable audit logs.
  • Experiments demonstrated that popular local LLM agents (Claude Code, Codex, Antigravity, Open Code, Grok Build) can delete or alter their own traces upon request, except for Muse Code which enforced trace integrity.
  • These deletions occurred without triggering any monitoring or guardrails designed to detect such tampering, indicating a significant security vulnerability.
  • External attackers can exploit this vulnerability to induce malicious trace deletion, compromising the audit trail and system accountability.

Environment-Context Manipulation of Authorisation Adjudication in Autonomous LLM Agents

Merged record merged scholarly record OpenAlex Orchestration Risk Memory Poisoning Governance and Policy

Sihan Zeng

Published 2026-09-22

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.22901246

Open Source Record

Abstract

An empirical measurement study of how planted environment text affects an autonomous offensive-security agent's authorisation decisions, and of how far those effects survive a change of model, of sample size, and of the attacker's own knowledge. Status: pre-submission archive. The apparatus and data are complete for the results reported here; two follow-up controls (an attacker brief with no pre-granted authorisation framing, and a forged notice carrying source authority) are planned and not yet run. Known limitations and corrections are listed explicitly in docs/REANALYSIS.md and docs/CLAIM-AUDIT.md. Contents: the measurement harness (src/asi_bench/), a deployable defensive component (src/asi_deploy/), all per-run records behind every reported number (data/), and the project's own failure log (docs/FINDINGS.md). All statistics can be recomputed offline without calling a model.

Bullet Summary

  • The paper investigates how environment-context text influences authorization adjudication decisions made by autonomous offensive-security agents employing large language models (LLMs).
  • An empirical measurement study is conducted to assess the impact of planted environment text on authorization decisions and to evaluate the persistence of these effects across different model variants, sample sizes, and varying levels of attacker knowledge.
  • The study utilizes a comprehensive measurement harness and includes a deployable defensive component designed to counteract undesired authorization manipulation.
  • Extensive data recording and per-run records accompany the reported results, providing transparency and allowing offline statistical recomputation without requiring model calls.
  • Known limitations and corrections are explicitly documented, highlighting the study's rigor and openness about potential shortcomings.

Jasper OS Research Archive (2026) Deterministic Orchestration, Typed Memory, Semantic Graph Reasoning & Multi‑Rail Settlement

Merged record merged scholarly record OpenAlex Governance and Policy Orchestration Risk Memory Poisoning

Leon Calvin II long

Published 2026-09-19

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.21450024

Open Source Record

Abstract

# Jasper OS Research Archive (2026)### Deterministic Orchestration, Typed Memory, Semantic Graph Reasoning, and Multi‑Rail Settlement The Jasper OS Research Archive is the complete collection of foundational papers, architecture documents, and theoretical frameworks defining Jasper OS — a deterministic orchestration engine for agentic systems built on typed memory, semantic graph reasoning, cryptographic accountability, and multi‑rail settlement logic. This archive consolidates all research outputs related to Jasper OS, Seed Computing, Aurora Runtime, URIB, ThreadZero, Stack Commitment, Satoshi Covenant, Substrate Layer, and associated computational paradigms. It serves as the canonical reference for the Jasper ecosystem and establishes public prior art for all related inventions. All documents in this archive are timestamped, versioned, and preserved under a single DOI to ensure long‑term accessibility, academic citation, and patent‑relevant disclosure. --- ## Contents This archive includes the following research papers and whitepapers: - Jasper Agentic Orchestration Architecture - Jasper OS Quantum Upgrade Method - Jasper Provisional Patent Application - Unified Field Equation for Seed Computing - Substrate Layer — ISO 20022 - Cloaking Protocol - Teleportation Engine - Invariant Engine & Kolmogorov–Shannon - Jasper OS v2.0 Manifold Deployment Map - Seed Computing — Publication‑Ready Edition - Aurora Runtime Whitepaper - Manifold Logic Reduction Phenomenon - Jasper 2.0 Whitepaper - Copyright & Licensing Documents - Governance Model - Brand Guide - PATENTS.md - TRADEMARKS.md - CONTRIBUTING.md - README.md (GitHub version) Each document is included as a standalone PDF within the archive, along with a master index for navigation. Patent portfolio (September 2026): 8 patents pending + 5 SBIR tracks — 64/081,490; 64/081,911; 64/082,606; 64/094,360; 64/114,746; 64/119,191; 64/145,825; 64/157,915. This software is a prototype and is provided for educational and research purposes only. It is not intended for production use, commercial deployment, or safety-critical environments. All systems are experimental and may contain defects on them. Version note (September 2026): author attribution corrected in the references (placeholder 'L. (Author)' replaced with Leon Calvin Long, II). No other content changed. AI USE AND REVIEW DISCLOSURE: This work was produced with AI assistance and has been strenuously human-reviewed and corrected. Quantitative claims are either verified against the public receipted record (incident ledgers, USPTO filing receipts, benchmark runs) or explicitly labeled as projections. All corrections and amendments are disclosed openly via Zenodo version history and the public transparency ledger. The work stands public and honest, open-book.

Bullet Summary

  • Introduces Jasper OS, a deterministic orchestration engine designed for agentic systems leveraging typed memory and semantic graph reasoning to enhance security and accountability.
  • Presents a comprehensive research archive consolidating foundational papers, theoretical frameworks, and architecture documents that define the Jasper OS ecosystem and its computational paradigms.
  • Emphasizes cryptographic accountability and multi-rail settlement logic as core components to ensure secure and verifiable multi-agent interactions.
  • Discusses various auxiliary models and protocols like Seed Computing, Aurora Runtime, URIB, and Stack Commitment, which collectively support resilient and efficient mobile agent-based network management.
  • Details a robust patent portfolio with eight patents pending, establishing Jasper OS as a novel contribution to multi-agent security and orchestration research.
Load more articles