arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02291cs.AI

共享前缀,更优信用:面向多智能体推理的自适应路由

Shared Prefixes, Better Credit: Adaptive Routing for Multi-Agent Reasoning

Yiqing Liu, Zihao Wang, Hantao Yao, Wu Liu, Yongdong Zhang

AI总结:

本研究针对现有多智能体推理自适应路由方法的粗粒度监督缺陷,提出TreeCredit共享前缀信用分配框架,通过状态匹配下游比较估计算子效用,在六个推理基准上实现了准确率与推理成本的更优权衡。

AI中文摘要:

多智能体推理(MAR)通过迭代的解决方案交换与优化提升推理可靠性。现有自适应MAR方法通常从查询级标签或轨迹级回报学习路由决策,但这类粗粒度监督无法准确估计多步协作中单个算子的状态条件效用。我们提出TreeCredit,一种用于高效自适应MAR的共享前缀信用分配框架。其核心见解是通过与状态匹配的下游比较估计算子效用,而非直接将轨迹级结果归因于先前决策。TreeCredit通过从同一中间状态扩展候选算子构建共享前缀协作树,并基于完整延续的最终正确性与累积额外成本,为每个状态-算子对分配正确性优先的后缀信用。这些结构化信用被转换为状态局部算子偏好,以训练轻量级成对状态路由器,该路由器在推理过程中动态选择下一个可允许的算子。在六个推理基准上的实验表明,TreeCredit适度提升了准确率,同时大幅降低了推理成本,相较于代表性MAR方法实现了更优的准确率-成本权衡。

英文摘要:

Multi-agent reasoning (MAR) improves reasoning reliability through iterative solution exchange and refinement. Existing adaptive MAR methods typically learn routing decisions from query-level labels or trajectory-level returns, but such coarse supervision cannot accurately estimate the state-conditioned utility of individual operators in multi-step collaboration. We propose TreeCredit, a shared-prefix credit assignment framework for efficient adaptive MAR. Its core insight is to estimate operator utility through state-matched downstream comparisons, rather than directly attributing trajectory-level outcomes to preceding decisions. TreeCredit constructs shared-prefix collaboration trees by expanding candidate operators from the same intermediate state and assigns each state--operator pair a correctness-prioritized suffix credit based on the terminal correctness and cumulative additional cost of its complete continuation. These structured credits are converted into state-local operator preferences to train a lightweight pairwise state router, which dynamically selects the next admissible operator during inference. Experiments on six reasoning benchmarks show that TreeCredit modestly improves accuracy while substantially reducing inference cost, achieving a better accuracy--cost trade-off than representative MAR methods.

↑