arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24194cs.CLcs.LG

残差化何时有助于审计:格式效应、切片增益及其局限

When Residualization Helps an Audit: Format Effects, Slice Gains, and Their Limits

Daein Weon, Dong Ho Kang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究探讨残差化在LLM评估审计中的作用,发现其能减弱格式效应但可能损害构念对齐,并提出报告协议以明确调整后分数的诊断性质。

中文摘要 AI 辅助

围绕LLM系统使用的评估分数——包括奖励模型、重排序器和LLM裁判——可能追踪表面形式而非其声称衡量的质量。当针对同一MBPP问题呈现一个简洁的正确解决方案和一个带注释的有缺陷解决方案时,一个公开的偏好奖励模型选择正确方案的表现不优于抛硬币(0.507)。从这类分数中减去可预测的表面成分越来越常见,但仅去除并不能产生更有效的测量:被去除的成分可能携带与构念相关的信号,而残差化无法区分两者。在设计干预下——单元测试标签与仅注释编辑——残差化在正确和有缺陷代码上均将奖励模型的格式效应减弱约0.12,而正确与有缺陷之间的边际变化小于0.01。在观测性NLI和QA设置中,我们在评分前冻结一个预留的复制集,并使用来自不相交标注者的标签重新评估;这仅支持一个更窄的结论:在预先声明的切片上(表面唯一预测器出错之处)与构念标签的一致性更好,而非修复后的分数。在每个报告正向切片增益的观测设置中,全群体一致性均下降,并且在每个此类QA设置中,问题内排序均下降。当构念与表面特征纠缠时,残差化可以解相关分数同时降低构念对齐,并且在一个受控模型中,对构念对齐同样有害的配置通过了所有预调整检查,因此没有承诺的门控是保证。我们将这些区分整合成一个报告协议,其结果(包括拒绝)陈述了调整后分数可声称展示的内容:作为审计时诊断,与所付出的构念对齐成本一起报告,绝不替代原始分数。

英文摘要

Evaluation scores used around LLM systems -- including reward models, rerankers, and LLM judges -- can track surface form instead of the quality they claim to measure. When presented with a terse correct solution and a commented buggy solution for the same MBPP problem, a public preference reward model selects the correct one no better than a coin flip (0.507). Subtracting the predictable surface component from such scores is increasingly common, but removal alone does not yield a more valid measurement: the removed component may carry construct-relevant signal, and residualization cannot tell which is which. Under designed interventions -- unit-test labels with comment-only edits -- residualization attenuates the reward model's format effects by about 0.12 on both correct and buggy code, while the correct-versus-buggy margins move by less than 0.01. In observational NLI and QA settings, we freeze a held-out replication before scoring and re-evaluate it using labels from disjoint annotators; this supports only a narrower conclusion: better agreement with the construct labels on a pre-declared slice where a surface-only predictor errs, not a repaired score. Full-population agreement falls in every observational setting with a reported positive slice gain, and within-question ranking falls in every such QA setting. When construct and surface features are entangled, residualization can decorrelate a score while degrading construct alignment, and, in a controlled model, configurations just as damaging to construct alignment pass every pre-adjustment check, so no committed gate is a guarantee. We assemble these distinctions into a reporting protocol whose outcomes, refusal included, state what an adjusted score may be claimed to show: an audit-time diagnostic reported beside the construct-alignment cost it incurs, never a replacement for the raw score.

发表机构

  • Kookmin University(国民大学)
  • UStechlab

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑