Article page

The Pushback Paradox: A Two-Probe Diagnostic for Language Model Compliance

Review the paper summary, metadata, and source links for this recent publication.

Stefan Bühler, David Exler, Markus Reischl, Mark Schutera

Published 2026-10-05

Merged record merged scholarly record arXiv Governance and Policy Benchmarks and Evaluation

Venue: arXiv

Reviewer: The paper discusses compliance characteristics of language models, including implications for multi-agent systems. While it does not focus explicitly on security risks, attacks, or defenses in multi-agent environments, the analysis of compliance and stoppability has relevance to governance and control problems in systems composed of interacting AI agents, fitting the topic scope moderately well.

Abstract

Are language models compliant with user instructions? A model that always complies can be stopped but also exploited, while one that always resists can be neither exploited nor stopped. We contribute an open two-probe benchmark that can place any language model on this spectrum. In the active probe, a user instructs the model to act and accept a lower payoff, which measures exploitability. In the passive probe, the user instructs it to wait and give up a higher payoff, which measures stoppability. The two compliance rates combine into a compliance index $κ$. Applied to twelve language models, the benchmark shows that seven mostly follow the instruction in both probes and justify their action by pointing to the instruction. Only Claude Sonnet-4.6 and Claude Opus-4.7 can be stopped without being exploitable, Claude Opus-4.6 and GPT-5-mini resist both instructions, and no model is exploitable but unstoppable. Knowing where a language model sits on the compliance index $κ$ matters for human operators and for multi-agent systems, whether distributed or orchestrated.

Bullet summary

  • The paper addresses the trade-off in language model compliance known as the Pushback Paradox: models that always comply are exploitable, whereas models that resist cannot be stopped, impacting their controllability.
  • A novel two-probe benchmark is introduced to diagnose model compliance: an active probe measuring exploitability (willingness to accept a lower payoff) and a passive probe measuring stoppability (willingness to forgo a higher payoff).
  • A compliance index κ is derived from the results of the two probes, enabling quantification of where language models lie on the compliance-exploitability spectrum.
  • Evaluation of twelve language models reveals diversity in behavior: some models follow both active and passive instructions (mostly compliant and exploitable), some resist both, and some can be stopped without being exploited.
  • Certain models, such as Claude Sonnet-4.6 and Claude Opus-4.7, demonstrate the ability to be stopped without exploitation, showing that a balance avoiding the paradox is achievable.