发表机构
University of Utah; Thoughtworks; University of Massachusetts Amherst; Martian AI(犹他大学; Thoughtworks; 马萨诸塞大学阿默斯特分校; Martian AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文揭示LLM真实性基准存在表面特征泄露,使模型无需推理即可答对,并提出清理机制Audit-Prune及清理后的TruthfulQA版本以消除该问题。
AI 中文摘要
二元选择真实性基准要求模型在正确与错误答案之间做出选择,但如果两个答案在表面特征上存在系统性差异,模型无需执行预期推理即可超越随机水平。我们证明,这种失败模式是可检测的,并且可被下游分类器利用。在TruthfulQA中,一个简单的六特征逻辑分类器在区分正确与错误答案方面达到了可观的准确率。我们进一步表明,类似的表面伪影也存在于其他基准中。为应对这一问题,我们开发了一种通用机制,通过移除最强化泄露的答案对来清理这些基准。我们发布了一个表面特征泄露降至接近随机水平的TruthfulQA版本,并提供了一种机制Audit-Prune,以便数据集在发布前可被清理。
英文摘要
Binary-choice truth benchmarks ask models to choose between a correct and an incorrect answer, but if the two answers differ systematically in surface-level features, models can exceed chance without performing the intended reasoning. We show that this failure mode is detectable and can be exploited by downstream classifiers. In TruthfulQA, a simple six-feature logistic classifier achieves substantial accuracy in separating correct from incorrect answers. We further show that similar surface-level artifacts are present in additional benchmarks. To counteract this, we developed a general mechanism to clean them by removing the most leakage-reinforcing pairs. We release a version of TruthfulQA with surface-feature leakage reduced close to chance and provide a mechanism, Audit-Prune, so that the datasets can be cleaned before release.
Comments31 pages, 4 figures. Code and data: https://github.com/foadnamjoo/audit-prune and https://huggingface.co/datasets/foadnamjoo/audit-prune