Research library

Curate the field by topic, not just recency.

This section is the long-lived knowledge layer for important papers, organizing the multi-agent security literature into durable subareas that can support reading lists, annotated references, and future topic pages.

Stored article buckets

These groups come from the categorized article database and show a preview of the latest papers in each bucket.

2280 papers

Governance and Policy

  • The Pushback Paradox: A Two-Probe Diagnostic for Language Model Compliance
  • HERA: Harness-Environment Co-Evolution for Reliable Agentic Abstention
  • ANT: A Multi-Granularity Network Traffic Dataset and Benchmark for Agents Behavior Auditing

1492 papers

Benchmarks and Evaluation

  • BazaarBench: Delegation Safety in Decentralized C2C Marketplaces Run by LLM Agents
  • The Pushback Paradox: A Two-Probe Diagnostic for Language Model Compliance
  • HERA: Harness-Environment Co-Evolution for Reliable Agentic Abstention

1300 papers

Orchestration Risk

  • BazaarBench: Delegation Safety in Decentralized C2C Marketplaces Run by LLM Agents
  • HERA: Harness-Environment Co-Evolution for Reliable Agentic Abstention
  • AECP: Artifact-Exclusive Communication Protocol for Multi-Agent Code Generation

1244 papers

Trust and Identity

  • BazaarBench: Delegation Safety in Decentralized C2C Marketplaces Run by LLM Agents
  • AgentPrivArena: Evaluating and Auditing Real-world AI Agent Privacy
  • Let the Agent Do It? How Software Practitioners Understand and Make Permission Decisions in Agentic AI Assistants

1162 papers

Agent-to-Agent Communication

  • AECP: Artifact-Exclusive Communication Protocol for Multi-Agent Code Generation
  • Do Small Language Models Learn to Negotiate? A Controlled Scaling Study of RL-Trained Sellers
  • Can CaMeLs Talk? Securing Multi-Agent Systems Against Indirect Prompt Injection Attacks

688 papers

Prompt Injection

  • AECP: Artifact-Exclusive Communication Protocol for Multi-Agent Code Generation
  • RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents
  • Compromise Is Not Consequence: Evaluating Task-Scoped Authorization in LLM Agents with Paired Replay

358 papers

Memory Poisoning

  • StegoMemory: Agentic Memory Acts as Covert Steganographic Channel
  • MemLeak: Cross-User Semantic Leakage in Multi-Tenant AI Agent Memory
  • AgentGuardBench: A Multilingual Benchmark for Privacy, Security and Responsible Behaviour in AI Agents