Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents
谁获得奖励 & 谁受到责备?面向多LLM智能体的评估对齐训练信号
机构 * Argonne National Laboratory(阿贡国家实验室) ; University of Chicago(芝加哥大学)
专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);post-training(abstract)
AI总结 提出一个理论框架,结合合作博弈归因与过程奖励建模,将系统级评估转化为智能体信用和消息级信号,用于多LLM智能体训练。
Comments Accepted at the NeurIPS 2025 Workshop on Bridging Language, Agent, and World Models for Reasoning and Planning (LAW 2025)