CHILL-Harness:用于长期智能体高效推理的反事实控制学习
CHILL-Harness: Counterfactual Harness Learning for Efficient Reasoning in Long-Horizon Agents
浏览论文内容
中文总结 AI 辅助
研究长期智能体控制中因依赖固定策略导致效率低的问题,提出CHILL-Harness方法,通过因果干预效应学习和优势实现因果编排,结合成功保留目标与约束,在多任务实验中减少消耗与时间,保持或提高任务成功率。
中文摘要 AI 辅助
智能体控制已成为现代大语言模型智能体的操作基础设施,协调上下文、工具、验证和执行控制,将潜在模型能力转化为可靠的长期行为。然而,可靠的长期行为需要控制适应任务需求、执行环境和不断变化的执行状态,而当前控制主要依赖手工或全局固定策略,导致计算开销大,甚至任务成功率降低。为解决此问题,我们将在控制中实现自适应编排的任务建模为因果学习问题,并提出用于长期智能体的反事实控制干预学习(CHILL-Harness)。CHILL-Harness在编排层进行干预,以实现优势引导的工作流适应,从而提高推理和执行效率,同时保持任务性能。具体而言,我们开发了因果干预效应学习作为CHILL-Harness的效应估计组件,从置信加权执行证据中估计干预相关的工作流优势,并识别有利的工作流适应。我们进一步引入优势实现因果编排作为其实现组件,以自适应分配反事实推理,并仅实现由足够预期优势支持的工作流调整。最后,我们将成功保留目标和优势边际授权约束纳入CHILL-Harness,以促进可靠的适应。在跨越信息检索、软件工程和终端交互的异构长期任务上进行的广泛实验表明,CHILL-Harness在显著减少令牌消耗和执行时间的同时,始终保持或提高任务成功率。
英文摘要
Agent harnesses have become the operational infrastructure of modern large language model agents, coordinating context, tools, verification, and execution control to translate latent model capability into reliable long-horizon behavior. However, reliable long-horizon behavior requires harness control to adapt to task demands, execution environments, and evolving execution states, whereas current harnesses predominantly rely on hand-crafted or globally fixed policies; this mismatch manifests as unnecessary computational overhead and, in adverse cases, reduced task success. To address this limitation, we formulate the task of enabling adaptive orchestration in harness systems as a causal learning problem and propose Counterfactual Harness Intervention Learning for Long-Horizon Agents (CHILL-Harness). CHILL-Harness intervenes at the orchestration layer to enable advantage-guided workflow adaptation, thereby improving reasoning and execution efficiency while preserving task performance. Specifically, we develop causal intervention effect learning as the effect-estimation component of CHILL-Harness to estimate intervention-relative workflow advantage from confidence-weighted execution evidence and identify advantageous workflow adaptations. We further introduce advantage-realizing causal orchestration as its realization component to adaptively allocate counterfactual reasoning and realize only workflow adjustments supported by sufficient expected advantage. Finally, we incorporate a success-preserving objective and advantage-margin authorization constraints into CHILL-Harness to promote reliable adaptation. Extensive experiments on heterogeneous long-horizon tasks spanning information seeking, software engineering, and terminal interaction show that CHILL-Harness consistently preserves or improves task success while substantially reducing token consumption and execution time.