覆盖率,而非靶向:多轮智能体信用分配中的一种结构 regime
Coverage, Not Targeting: A Structural Regime in Multi-Turn Agent Credit Assignment
浏览论文内容
中文总结 AI 辅助
该研究发现多轮智能体信用分配中,终端状态验证器处于低信息密度 regime,均匀分布奖励(覆盖率)优于靶向分配,在多个基准及模型家族上可复现,提出需用匹配浓度打乱对照检验靶向主张。
中文摘要 AI 辅助
多轮智能体强化学习(RL)日益将信用分配视为一个靶向问题:给定一个可验证的终端奖励,逐轮方法会将信用定位到重要的轮次。我们确定了一个预测何时该方法适用的结构量,即验证器信息密度V_d = k/C(其中C是智能体因果链的步数,k是验证器所揭示的逐轮正确性的比例),并表明终端状态验证器处于低V_d regime,此时靶向并非正确方向。在tau^2-bench上的受控共享rollout比较中,该基准将奖励密度与信用几何分离,均匀分布的连续密集奖励优于稀疏二元结果奖励(在5个随机种子中,有4个产生净危害),而将相同优势集中在进展轮次或随机轮次同样有害:靶向是二阶效应。其机制是覆盖率:终端状态验证器将可观测信号坍缩为单个最终写入轮次(在98%的rollouts中k=1),而成功需要5-8步的先决工具调用链。一个合成相界将交叉点置于V_d*≈0.8,而在tau^2-bench上测得的V_d约为0.15,在BFCL V3上约为0.4;均匀分布在BFCL上同样占优,其中匹配浓度的打乱对照在8个随机种子中全部为负。该效应在ToolACE-2-8B的不同模型家族中可复现(在32个预注册种子上的差值为-0.048;独立的20个种子复现本身具有显著性),且一项预注册的匹配预算广度扫描呈现出单调剂量反应,其缺陷仅在完整链覆盖率时消失,而奖励到走(reward-to-go)分支达到了全覆盖率的对等性。均匀再分配是零信息覆盖率的默认方案,是逐轮方案必须超越的基准;我们提供了匹配浓度的打乱对照,任何靶向主张都应通过该对照的检验。
英文摘要
Multi-turn agentic RL increasingly treats credit assignment as a targeting problem: given a terminal verifiable reward, per-turn methods localize credit onto the turns that mattered. We identify the structural quantity that predicts when this is the right move, the verifier information density V_d = k/C (the fraction of an agent's C-step causal chain whose per-turn correctness the verifier exposes), and show that terminal-state verifiers sit deep in a low-V_d regime where targeting is the wrong axis. In controlled shared-rollout comparisons on tau^2-bench that separate reward density from credit geometry, a continuous dense reward spread uniformly beats the sparse binary outcome reward (net-harmful on 4/5 seeds), while concentrating the same advantage on progress turns or on random turns is equally harmful: targeting is second-order. The mechanism is coverage: terminal-state verification collapses the observable signal to a single final-write turn (k=1 in 98% of rollouts) while success requires a 5-8 step chain of prerequisite tool calls. A synthetic phase boundary places the crossover at V_d* ~ 0.8, whereas measured V_d is ~0.15 on tau^2-bench and ~0.4 on BFCL V3; uniform also wins on BFCL, where a matched-concentration shuffled control is negative on 8/8 seeds. The effect reproduces across model families on ToolACE-2-8B (Delta = -0.048 over 32 pre-registered seeds; an independent 20-seed replication is itself significant), and a pre-registered matched-budget breadth sweep traces a monotone dose-response whose deficit vanishes only at full chain coverage, with a reward-to-go arm reaching full-coverage parity. Uniform redistribution is the zero-information coverage default that per-turn schemes must beat; we contribute the matched-concentration shuffled control that any targeting claim should clear.
发表机构
- School of Engineering, Institute of Science Tokyo(东京科学技术学院工学院)
- College of Control Science and Engineering, Zhejiang University(浙江大学控制科学与工程学院)
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。