AI 中文总结
研究智能体评估中成功溯源问题,通过AcquaBench的匹配值替换和聚类分析审核,在不同场景测试,发现成功与正确值相关,行为依赖可能超出预期,模型比较有分数差距变化,建议基准报告成功及信息状态支持情况。
AI 中文摘要
一个正确答案可能掩盖智能体成功的原因。当智能体在评估过程中改变其信息状态时,正确性无法区分预期推理和答案获取。结果证据和暴露检测无法确定成功是否依赖于获取的目标,我们将这种缺失的评估对象称为成功溯源。AcquaBench通过在四个标准化表面上进行匹配的CLEAN、GOLD和SHAM值替换以及联合qid聚类分析来审核它。CLEAN保留基准授权信息。GOLD提供正确目标。SHAM保留源结构和暴露机会,但替换匹配的错误值。GOLD减去CLEAN衡量对正确目标可用性的总分响应;GOLD减去SHAM测试该响应是否跟踪目标正确性超过匹配的源暴露。在D0中,GOLD比SHAM高出19.1到25.9个百分点,表明成功遵循正确值。在D2中,在分布式充足的情况下GOLD仍超过SHAM,而coloc不再作为高分标记转移,AUROC为0.376和0.142。行为依赖可能在该探测的预期观察单元之外持续存在。在模型比较中,支持的5.0分CLEAN分数差距压缩为原始GOLD差异-0.6分,而未建立排名反转。智能体基准应报告成功以及评估的信息状态是否支持它。
英文摘要
A correct answer can conceal why an agent succeeded. Once agents change their information state during evaluation, correctness no longer distinguishes intended reasoning from answer acquisition. Outcome evidence and exposure detection do not establish whether success depended on an acquired target; we call this missing evaluation object success provenance. AcquaBench audits it through matched CLEAN, GOLD, and SHAM value substitution on four standardized surfaces with joint qid-clustered analysis. CLEAN retains benchmark-authorized information. GOLD makes the correct target available. SHAM preserves source structure and exposure opportunity but substitutes a matched incorrect value. GOLD minus CLEAN measures the total score response to correct-target availability; GOLD minus SHAM tests whether that response tracks target correctness beyond matched source exposure. In D0, GOLD exceeds SHAM by 19.1 to 25.9 percentage points, showing that success follows the correct value. In D2, GOLD still exceeds SHAM under distributed sufficiency while coloc no longer transfers as a high-score marker, with AUROC 0.376 and 0.142. Behavioral dependence can thus persist beyond this probe's intended observation unit. In model comparison, a supported 5.0-point CLEAN score gap compresses to a raw GOLD difference of -0.6 points without establishing rank inversion. Agent benchmarks should report success together with whether the evaluated information state supported it.
Comments17 pages, 3 figures, including supplementary material. Code: https://github.com/luojingkun22/acquabench