arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

选择性批判:面向长时程决策中成本感知的LLM智能体

Selective Critique for Cost-Aware LLM Agents in Long-Horizon Decision Making

Heewon Park, Somin Im, Minhae Kwon

arXiv 2610.07335首次发表:更新:

发表机构

Sungkyunkwan University; Soongsil University(成均馆大学; 崇实大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出SAG框架,通过门控机制选择性调用批判,利用动作级模糊信号近似信息价值,在长时程决策中平衡性能与成本,显著提升任务成功率并降低推理开销。

AI 中文摘要

提高大型语言模型(LLM)智能体在长时程决策中的可靠性仍是一个关键挑战。当作为自主智能体部署在与复杂环境交互的场景中时,早期错误可能沿轨迹传播并导致级联失败。近期方法通过引入外部批判或深思来提高可靠性,但在每一步都调用这些机制会大幅增加令牌消耗和延迟,限制了实际部署。我们提出SAG(带门控批判的自我改进智能体),一种成本感知框架,将批判调用形式化为长时程交互中的逐步决策问题。SAG引入了一种轻量级、无需训练的门控机制,利用基于可行动作计算的全局熵和局部前二边际等动作级模糊信号来估计批判的效用。从决策论角度看,该机制近似了批判的信息价值(VoI),使智能体能够仅在预期收益证明成本合理时选择性分配昂贵的反馈。SAG还整合了在线自助式自我改进,使行动者能够内化批判辅助行为,并逐步减少对批判的依赖。在三个长时程交互基准和多个骨干模型上,与无批判和始终开启批判的智能体相比,SAG显著改善了性能-成本权衡。在ALFWorld上,SAG将任务成功率从24.6%提高到78.4%,同时保持与ReAct相当的令牌预算,实现了归一化令牌效率的3.1倍提升。此外,一个7B行动者搭配轻量级3B批判者所达到的性能可与无批判的14B行动者相媲美,表明选择性批判可以恢复深思的大部分可靠性收益,同时大幅降低推理成本。

英文摘要

Improving the reliability of large language model (LLM) agents in long-horizon decision-making remains a key challenge. When deployed as autonomous agents interacting with complex environments, early mistakes can propagate through trajectories and cause cascading failures. Recent approaches improve reliability by incorporating external critique or deliberation, but invoking these mechanisms at every step substantially increases token consumption and latency, limiting practical deployment. We propose SAG (Self-improving Agent with Gated critique), a cost-aware framework that formulates critique invocation as a step-wise decision problem during long-horizon interaction. SAG introduces a lightweight, training-free gating mechanism that estimates the utility of critique using action-level ambiguity signals--global entropy and local top-2 margin--computed over admissible actions. From a decision-theoretic perspective, this mechanism approximates the Value of Information (VoI) of critique, enabling the agent to selectively allocate expensive feedback only when its expected benefit justifies the cost. SAG further incorporates online bootstrapped self-improvement, allowing the actor to internalize critic-assisted behaviors and progressively reduce reliance on critique. Across three long-horizon interactive benchmarks and multiple backbone models, SAG substantially improves the performance-cost trade-off compared with both no-critique and always-on critique agents. On ALFWorld, SAG increases task success from 24.6% to 78.4% while maintaining a token budget comparable to ReAct, yielding a $3.1\times$ improvement in normalized token efficiency. Moreover, a 7B actor with a lightweight 3B critic achieves performance comparable to a 14B actor without critique, showing that selective critique can recover most of the reliability benefits of deliberation while dramatically reducing inference cost.

CommentsAccepted to NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑