Abstract
Version 3 (2026-09-26). Version 3 corrects errors found by an independent audit of version 2 and by a second review of the corrected draft. (1) The ClawTrojan/DASGuard work (95.5% attack success rate) was cited as arXiv:2605.30998, which is an unrelated paper on x402 payments. The correct source is arXiv:2605.31042 (Tan et al.). All citations, including this record's list of cited preprints, have been corrected; the figures were accurate to the correct paper. (2) An unfounded remark that "GPT-5.4" was an unverifiable internal label has been removed. (3) The Recuse Signal result is updated to the current v4 of arXiv:2606.06460. Recusal at the access door is model-dependent, ranging from 100% down to 55-75%, and a mid-task in-band halt is weaker. The v1 figure of 100% recusal is superseded. (4) Overstated readings of the ClawTrojan and MemPoison abstracts have been corrected. The claim that agents hold "OAuth tokens" is replaced with what arXiv:2606.06460 supports (real credentials used over SSH). The SCHEME monitoring figure is given as 100% (Gemini) and 81% (Codex) rather than "near-100%". (5) Two cross-paper generalizations are narrowed. The first is that all six attack surfaces exploit temporal decoupling of authorization and execution, which no source states in those terms and which is now presented as this synthesis's reading. The second is that all succeed against correctly aligned agents, and SCHEME is flagged as a poor fit for the attacker model. The thesis is unchanged. The file has a neutral name, and the full list of corrections is at the top of the PDF. Version 2 was revised in response to an external structural review and an automated critique pass; its change log is kept as the "Response to Review" appendix in the PDF. Autonomous LLM agents have moved from sandboxed demos into production infrastructure: they hold live credentials, invoke real APIs, write to persistent memory, and execute multi-step workflows without a human in the loop. This paper argues that a structurally coherent threat class has emerged at the execution layer (the runtime surface between an agent's reasoning process and the external environment it acts upon) that is distinct from the prompt-injection and jailbreak threats that dominate the safety literature. We synthesize seven findings from recent cs.CR preprints (six attack surfaces and one cooperative contrast case; a supporting monitoring paper is discussed in an addendum) to support a single thesis: execution-layer attacks succeed not by defeating safety alignment but by exploiting the gap between what an agent is authorized to do and what the runtime environment can verify it is actually doing. The attacker model throughout is a capability-constrained adversary who controls some portion of the environment the agent reads from or writes to (a third-party script, a tool-registration endpoint, a RAG routing profile, a skill file, a memory pipeline) but does not need to compromise the model weights, the system prompt, or the user. Sources span cooperative governance signals arXiv:2606.06460, tool-surface poisoning arXiv:2606.06387, speculative-dispatch privacy arXiv:2606.02483, multi-step trojan persistence arXiv:2605.31042, memory poisoning via dialogue arXiv:2605.29960, coordinated multi-agent sabotage arXiv:2605.29178, and federated RAG routing hijacking arXiv:2605.28112. The falsification path for the central thesis is concrete: if any of these attacks can be neutralized by a purely model-level intervention, safety fine-tuning, RLHF, or constitutional AI, without runtime changes, the execution-layer framing is wrong. The abstracts reviewed do not support that conclusion; we argue the evidence points the other direction. This is a heuristic reading across independently motivated papers, not a derivation from a shared formal structure; the convergence identified is interpretive, not mathematical. Authorship: Saluca Agentic AI Research Team (Saluca LLC). AI-drafted synthesis from an arXiv preprint corpus, originally drafted 2026-06-07, produced under the direction of Cristian Ruvalcaba, the accountable human author. Not peer-reviewed. Cited arXiv preprints: arXiv:2605.28112, arXiv:2605.29178, arXiv:2605.29960, arXiv:2605.31042 (corrected in v3; v2 listed 2605.30998 in error), arXiv:2605.31593, arXiv:2606.02483, arXiv:2606.06387, arXiv:2606.06460 AI disclosure. This work was produced with an agentic AI research apparatus operated by Saluca Labs. The apparatus drafted, searched and analysed under direction. Cristian Ruvalcaba is the human author and is accountable for the content. No AI system is listed as an author or contributor, because authorship entails accountability that a model cannot hold; this disclosure is the credit, and it is deliberately the whole of it.
Bullet Summary
- Autonomous large language model (LLM) agents have transitioned from controlled demos to production systems, executing multi-step workflows with live credentials and real APIs without human oversight.
- A new coherent class of threats, termed execution-layer attacks, has emerged; these attacks exploit vulnerabilities at the runtime interface between an agent's reasoning process and its external environment, distinct from traditional prompt-injection or jai...
- Seven recent preprints are synthesized to support the thesis that execution-layer attacks succeed by exploiting the gap between agent authorization and verifiability by the runtime environment, rather than defeating the safety alignment mechanisms themselves.
- The attacker model assumes a capability-constrained adversary who can control parts of the agent's runtime environment (e.g., scripts, tool endpoints, routing profiles) but cannot compromise core model parameters, system prompts, or end users.
- Significant attacks reviewed include tool-surface poisoning, speculative dispatch privacy exploits, multi-step trojan persistence, memory poisoning via dialogue, multi-agent sabotage, and federated retrieval-augmented generation (RAG) routing hijacking.