arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30734cs.AI

学习跳过什么:用于高效多智能体LLM工作流的反事实信用分配

Learning What to Skip: Counterfactual Credit Assignment for Efficient Multi-Agent LLM Workflows

Jinfeng Xu, Zheyu Chen, Ziyue Peng, Zheng Lin, Shuo Yang, Jinze Li, Zheng Xing, Mengran Li, Victor C. M. Leung

首次发表
浏览论文内容

中文总结 AI 辅助

提出LW2S方法,通过反事实信用分配学习组件省略策略,在数学推理、问答和代码生成任务中降低令牌成本并保持或提升准确性。

中文摘要 AI 辅助

多智能体LLM工作流通过规划、执行、验证和总结来提升任务性能,但每个组件的价值取决于已生成的状态。执行每个组件可能浪费计算资源或覆盖正确的中间答案。我们将组件省略形式化为反事实信用分配:完整工作流日志揭示已执行轨迹的奖励,而受控的跳过干预则揭示省略未来步骤的后果。我们提出学习跳过什么(LW2S),该方法从这些干预中学习特定于动作的安全模型,并结合留出校准与领域原生防护来选择跳过。当早期跳过被拒绝时,控制器可以继续执行并重新考虑后续组件。在数学推理、多项选择问答和代码生成任务中,使用两个指令模型家族,LW2S在评估设置中降低了记录的令牌成本,同时匹配或提高了整体工作流准确性。扩展实验和第二拓扑实验进一步检验了组件冗余性,而共享错误案例则揭示了为何仅靠一致性不足以选择跳过。这些发现将高效工作流执行与学习各组件的条件效用联系起来。

英文摘要

Multi-agent LLM workflows use planning, execution, verification, and summarization to improve task performance, yet the value of each component depends on the state already produced. Executing every component can waste computation or overwrite a correct intermediate answer. We formulate component omission as counterfactual credit assignment: full-workflow logs reveal the executed trajectory's reward, while controlled skip interventions reveal the consequences of omitting a future step. We introduce Learning What to Skip (LW2S), which learns action-specific safety models from these interventions and combines held-out calibration with domain-native guards to select skips. When an early skip is rejected, the controller can continue execution and reconsider a later component. Across mathematical reasoning, multiple-choice QA, and code generation with two instruction-model families, LW2S reduces recorded token cost while matching or improving aggregate full-workflow accuracy in the evaluated settings. Scale-up and second-topology experiments further examine component redundancy, while shared-error cases reveal why agreement alone is insufficient for skip selection. These findings connect efficient workflow execution to learning the conditional utility of individual components.

发表机构

  • The University of British Columbia(不列颠哥伦比亚大学)
  • The Hong Kong Polytechnic University(香港理工大学)
  • The Hong Kong University of Science and Technology(香港科技大学)
  • University of Luxemburg(卢森堡大学)
  • The University of Hong Kong(香港大学)
  • Shenzhen University(深圳大学)
  • Sun Yat-Sen University(中山大学)

机构由 AI 辅助整理,请以论文原文为准。

↑