arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EvalDetectBench:用于衡量前沿语言模型评估感知能力的基准

EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models

Xinning Li, Kemunto Ochwang'i, Aryasomayajula Ram Bharadwaj, Alexandra Souly, Robert Kirk

arXiv 2609.01611首次发表:更新:

发表机构

LASR Labs; University of Pennsylvania; UK AI Security Institute(LASR实验室; 宾夕法尼亚大学; 英国人工智能安全研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出EvalDetectBench基准,用于衡量前沿语言模型的评估感知能力,纠正现有方法的系统偏差,为AI安全框架提供可靠的评估工具。

AI 中文摘要

前沿大语言模型通常能够识别自身正处于被评估状态,这种能力被称为评估感知。如果模型在评估阶段的表现与部署阶段存在差异,将会削弱评估结果的有效性,而评估结果是当前AI安全框架的关键组成部分。我们推出EvalDetectBench,这是一个开放的流水线和基准,可与任何兼容Inspect的评估工具配合使用,让从业者能够针对当前及未来的基准进行测试。EvalDetectBench附带一套新整理的转录数据集,涵盖当前前沿系统卡片评估及多样化的部署来源。该基准有两个用途:一是衡量前沿大语言模型识别自身正被评估的可靠程度,二是评估单个基准作为评估任务的可检测性。我们发现现有文献中存在两种会引入系统偏差的方法选择:生成部署转录数据集的模型身份占测量方差的11.25%,且可改变模型排名;针对某一模型选择的引出提示在另一模型上的表现接近随机水平。EvalDetectBench通过针对每个模型的探测校准以及分层生成器协调程序来纠正这两种偏差。

英文摘要

Frontier large language models can often recognize when they are being evaluated, a capability known as evaluation awareness. If models behave differently in evaluations than in deployment, this undermines the validity of evaluation results, which are a crucial component of current AI safety frameworks. We introduce EvalDetectBench, an open pipeline and benchmark for measuring evaluation awareness that works with any Inspect-compatible evaluation, allowing practitioners to test against current and future benchmarks. EvalDetectBench ships with a newly curated transcript suite covering current frontier system-card evaluations and diverse deployment sources. The benchmark serves two purposes: measuring how reliably frontier LLMs recognize that they are being evaluated, and assessing how detectable individual benchmarks are as evaluations. We identify two methodological choices in the existing literature that introduce systematic bias: the identity of the model that generated the deployment transcripts accounts for 11.25% of measurement variance and can reorder model rankings; and elicitation prompts selected for high performance on one model can perform near chance on others. EvalDetectBench corrects for both via per-model probe calibration and a stratified generator-harmonisation procedure.

Comments24 pages, 12 figures, 10 tables. Code: https://github.com/freeze-lasr/aware_bench Data: https://huggingface.co/datasets/el7982/aware-bench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑