Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents
谁获得奖励 & 谁受到责备?面向多LLM智能体的评估对齐训练信号
机构 * Argonne National Laboratory(阿贡国家实验室) ; University of Chicago(芝加哥大学)
AI总结 提出一个理论框架,结合合作博弈归因与过程奖励建模,将系统级评估转化为智能体信用和消息级信号,用于多LLM智能体训练。
Comments Accepted at the NeurIPS 2025 Workshop on Bridging Language, Agent, and World Models for Reasoning and Planning (LAW 2025)