跨视图对应是一种测量干预:智能体评估与信用分配的双面验证
Cross-View Correspondence Is a Measurement Intervention: Two-Sided Validation for Agent Evaluation and Credit Assignment
- School of Computation, Information and Technology(计算、信息与技术学院)
- Technical University of Munich(慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
该研究指出跨视图对应是测量干预,开发含双面验证等的有效性理论,发现对应分歧会影响智能体评估,需声明验证跨视图对应以支持可靠结论。
中文摘要 AI 辅助
智能体评估与基于轨迹的学习通常会在响应后通过被视为中性预处理的变换视图间比较输出。我们表明,这种对应是一种测量干预:省略它会产生敏感性,过于激进的映射会产生不变性,而多个最优对应会导致机制标签与带符号学习信用无法识别。我们开发了包含三个部分的有效性理论与审计:对干扰项移除与响应保留的双面验证、对下游结论的全最优识别、以及有效性确立后的不确定性传播。我们刻画了响应保留型干扰项移除的线性可行性边界,计算了精确最优对应集上的精确范围,并提供了无分布证明:仅当所有精确最优对应就其非零符号达成一致时,才保留信用坐标。在公开代码与SQL流水线中,1586个非零轨迹对里有55.9%的两个确定性最优回溯在时间定位上存在分歧;两个冻结的800次工具使用审计(含任务与种子不相交的复现)揭示了预期回合级信用的精确最优反转,而干净的公开快速入门子集则无此现象。一个预注册的传输门在自然响应上失效;冻结的修正与保留对照组随后显示,仅在良性示例上校准的映射会抹去所有保留的有害响应,而双面验证会选择响应保留型替代方案。因此,在智能体评估或信用分配支持单点结论前,必须声明、验证跨视图对应并将其纳入不确定性传播。
英文摘要
Agent evaluations and trace-based learning often compare outputs across transformed views through a post-response correspondence treated as neutral preprocessing. We show that this correspondence is a measurement intervention: omitting it can manufacture sensitivity, an over-aggressive map can manufacture invariance, and multiple optimal correspondences can leave mechanism labels and signed learning credit unidentified. We develop a validity theory and audit with three components: two-sided validation of nuisance removal and response preservation, all-optima identification of downstream conclusions, and uncertainty propagation after validity is established. We characterize the linear feasibility boundary for response-preserving nuisance removal, compute sharp ranges over exact-optimum correspondence sets, and give a distribution-free certificate that retains a credit coordinate only when all exact optima agree on its nonzero sign. Across public code and SQL pipelines, two deterministic optimal tracebacks disagree on temporal localization for 55.9% of 1,586 nonzero trajectory pairs; two frozen 800-rollout tool-use audits, including a task-and-seed-disjoint replication, expose exact-optimum reversals of intended turn-level credit, although a clean public quick-start subset shows none. A pre-registered transport gate failed on natural responses; frozen corrected and held-out controls then show that a map calibrated only on benign examples erases every retained harmful response, while two-sided validation selects response-preserving alternatives. Cross-view correspondence must therefore be declared, validated, and propagated into uncertainty before agent evaluation or credit assignment supports a point conclusion.