无真实值的信用:针对执行重放的大语言模型智能体的步骤级信用分配审计
Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay
浏览论文内容
中文总结 AI 辅助
该研究针对LLM智能体,在ALFWorld环境中审计步骤级信用分配,发现现有信用信号无法识别关键步骤,提出需匹配有效样本量比较信用规则。
中文摘要 AI 辅助
在单智能体工具环境ALFWorld中,基于执行重放的因果真实值进行审计,用于训练大语言模型(LLM)智能体的所有步骤级信用信号——LLM评判分数、结果条件对数概率比或策略自身的置信度,均未比随机猜测更好地识别哪些步骤具有因果重要性。现有评估将这些信号与标注的步骤正确性挂钩,而我们则针对步骤贡献进行审计,即对每个决策点重新采样策略自身的替代方案并向前推进,实际会对结果产生何种改变,这两者存在差异。真实值本身具有结构:因果贡献是稀疏的(在定义了真实值的决策点中,30.5%的点具有可测量的影响),且可测量性依赖于模型——两个规模相似的策略之间,无策略支持的反事实的点的比例相差两倍(13.1% vs. 26.8%)。失败模式可识别:隐式信用与策略的流畅性相关(中位秩相关系数为+0.75,在修正工具下,第二组的复制结果为+0.70),而对结果进行条件处理未增加任何因果信息(部分相关系数为-0.004,Qwen)。仅基于置信度的路由器以随机水平恢复关键步骤,但每回合可降低13.1%的评判成本(每条轨迹降低14.0%)。在七臂预注册训练实验中,没有任何一个臂可靠地优于未训练的策略,且检查点的明显工具特征完全由训练剂量解释——更稀疏的信用保留更少的示例,优化器步骤存在一个数量级的差异,而非信用内容。因此,信用规则的比较必须匹配有效样本量,否则测量的是剂量而非信用。
英文摘要
Audited against policy-conditional ground truth from executed replay in a single-agent tool environment (ALFWorld), none of the step-level credit signals we audit -- LLM-judge scores, outcome-conditioned logprob ratios, or the policy's own confidence -- shows reliable incremental fidelity beyond its own marginal-matched shuffled control. Correcting for replay-target reliability leaves implicit fidelity bounded near zero and judge fidelity inconclusive at the achieved target reliability. Existing evaluations grade these signals against annotated step *correctness*; we audit them against step *contribution* -- what re-sampling the policy's own alternatives at each decision point and rolling forward changes about the outcome -- and they come apart. The ground truth is structured: 30.5% of decision points where it is defined exhibit a nonzero replay contrast at the achieved sampling resolution, and measurability is model-dependent -- the fraction of points with no policy-supported counterfactual differs twofold (13.1% vs. 26.8%) between two similar-scale policies. The failure mode is identifiable: implicit credit echoes the policy's fluency (median rank correlation +0.75, replicating at +0.70 in a second family under a corrected instrument), while outcome conditioning adds no causal information (partial correlation -0.004, Qwen). A confidence-only router recovers pivotal steps at chance level, but cuts judge cost by 13.1% per turn (14.0% per trajectory). In a seven-arm pre-registered training experiment, no arm reliably outperforms the untrained policy, and the checkpoints' apparent instrument signature is statistically consistent with mediation by effective training dose in this design -- sparser credit retains fewer examples, an order-of-magnitude spread in optimizer steps -- not credit content. Comparisons of credit rules must match effective sample size, or they measure dose, not credit.
发表机构
- University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。