AI 中文总结
GLIDE通过逐层残差一致性提取内在证据并校准为悲观奖励,实现无需外部验证器的异构LLM智能体步骤级评估,提升多跳推理等任务性能与效率。
AI 中文摘要
LLM智能体需要可靠的步骤级评估来比较候选分支并有效分配计算资源。然而,轻量级评估仍然具有挑战性。外部验证器会引入额外的推理成本,而智能体自身产生的置信度或自我评估分数可能校准不当,尤其是在候选由异构智能体生成的情况下。我们提出了面向LLM智能体的广义逐层内在分布评估(GLIDE)。GLIDE从逐层残差一致性中推导内在步骤证据,该一致性衡量局部残差更新是否持续支持候选步骤引起的全局残差变化。它根据生成智能体的近期分数分布校准该证据,并将其转换为悲观奖励,该奖励同时考虑绝对残差证据和智能体相对位置。该奖励为MCTS分支选择提供跨智能体价值信号,而归一化预测不确定性指导自适应分支。在多跳推理、顺序决策和符号逻辑上的实验表明,GLIDE在无需外部验证器或任务特定监督的情况下,提升了任务性能、步骤级排序质量和计算效率。
英文摘要
LLM agents require reliable step-level evaluation to compare candidate branches and allocate computation effectively. However, lightweight evaluation remains challenging. External verifiers introduce additional inference cost, while agent-produced confidence or self-evaluation scores can be miscalibrated, especially when candidates are generated by heterogeneous agents. We propose \textbf{G}eneralized \textbf{L}ayer-wise \textbf{I}ntrinsic \textbf{D}istributional \textbf{E}valuation (\textbf{GLIDE}) for LLM agents. \textsc{GLIDE} derives intrinsic step evidence from layer-wise residual coherence, which measures whether local residual updates consistently support the global residual change induced by a candidate step. It calibrates this evidence against the recent score distribution of the generating agent and converts it into a pessimistic reward that jointly accounts for absolute residual evidence and agent-relative standing. The reward provides a cross-agent value signal for MCTS branch selection, while normalized predictive uncertainty guides adaptive branching. Experiments on multi-hop reasoning, sequential decision making, and symbolic logic show that \textsc{GLIDE} improves task performance, step-level ranking quality, and computational efficiency without external verifiers or task-specific supervision.
CommentsAccepted by EMNLP 2026 Findings