Abstract
Large language model (LLM) agents are increasingly capable of acting in complex tool-use environments, yet they often fail to recognize when tasks are infeasible and no valid solution exists. Recent work has formalized this reliability gap as the problem of agentic abstention, and existing approaches typically optimize a model or agent harness against a fixed set of tasks, leading to limited generalization to unseen failure modes. We introduce HERA, a framework for harness-environment co-evolution for agentic abstention. HERA consists of (i) a pipeline to automatically construct verifiable pairs of feasible and infeasible tasks by applying controlled environment mutations that transform solvable tasks into cases requiring abstention, and (ii) a co-evolution procedure in which performance failures on previous tasks are used to drive harness adaptation and generate new execution environments and tasks geared towards previous weaknesses. On held-out evaluation tasks, an evolved harness from HERA improves abstention accuracy from 61.7% to 83.3% while improving feasible-task completion from 68.3% to 76.7%, achieving the highest abstention and feasible-task completion among the compared methods. The resulting best harness transfers across 19 other LLMs, improving abstention accuracy by 15.3 percentage points on average without any model-specific optimization, and enabling smaller models to match the performance of more powerful models at an estimated 85% lower cost.
Bullet summary
- The paper addresses the challenge of agentic abstention in large language model (LLM) agents, specifically their difficulty recognizing when tasks are infeasible and should be declined.
- HERA is introduced as a novel co-evolution framework that simultaneously evolves the agent's harness (control logic and reasoning abilities) and the environment (task distributions) based on failure feedback, promoting adaptability to new failure modes.
- A pipeline constructs verifiable paired tasks (feasible and infeasible) through controlled environment mutations, ensuring the agent is trained on robust abstention cases with validated ground truth.
- The co-evolution process iteratively diagnoses failures from agent rollouts to generate new challenging tasks and optimize the harness by adding logic for evidence-based decision gates and multi-constraint verification, improving both abstention accuracy an...
- HERA achieves significant improvements on held-out benchmark tasks (HERA-BENCH), increasing abstention accuracy from 61.7% to 83.3% and feasible task completion from 68.3% to 76.7%, outperforming multiple baselines.