arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当前AI基准测试未测量的内容:模态、搜索、引用及其对安全评估的启示

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)

Ro Encarnación, Tina Behzad, Emma Lurie, Danaé Metaxa

arXiv 2608.06202首次发表:更新:

发表机构

University of Pennsylvania; Stony Brook University(宾夕法尼亚大学; 石溪大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现AI基准测试未涵盖模态、搜索等关键因素,以ChatGPT为例对比两种模态及搜索条件,发现这些因素会影响模型准确率、一致性等,主张AI安全评估需纳入这些维度。

AI 中文摘要

大型语言模型(LLM)的基准测试评估常被用于支持关于模型安全性、可靠性及部署就绪性的论断。然而,大多数评估依赖单一访问模态(模型API),对每个提示仅执行一次运行,并将准确率作为主要结果指标,未考虑诸如网页搜索这类可能影响模型在部署中行为的条件。我们针对最广泛使用的LLM之一验证这些假设,比较两种模态:ChatGPT的聊天用户界面(UI)与OpenAI的API,分别在启用和未启用网页搜索的情况下进行。我们使用来自两个流行基准BBQ和SafetyBench的分层总样本401个提示,每个提示收集三次重复运行的共4812条响应。除标准性能指标外,我们还评估模型输出维度,包括响应一致性、响应文本相似度、引用依据及弃权(不执行)行为。例如,在未启用搜索时,两种基准上聊天UI的响应准确率均低于API响应;启用网页搜索使准确率降低最多8个百分点,甚至反转了其中一个基准的模态性能趋势方向;相同提示的重复运行在多达21%的提示中产生不一致的响应;两种模态还基于不同的引用为答案提供依据,且弃权(不执行)行为在两种模态间也不一致。这些结果表明,即使在同一模型家族内,仅报告简单的准确率指标也会掩盖与AI安全评估相关的重要模型行为变异形式。我们认为,AI安全评估应系统考虑模态、多运行一致性、搜索条件及响应级行为,以更好反映部署后的AI系统在实践中的表现。

英文摘要

Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness. Yet most evaluations rely on a single access modality (model APIs), perform a single run per prompt, and report accuracy as the primary outcome metric, without accounting for conditions such as web search that may have effects on model behavior in deployment. We audit these assumptions for one of the most widely-used LLMs, comparing two modalities, ChatGPT's chat UI and OpenAI's API, with and without web search enabled. We use a stratified total sample of 401 prompts from two popular benchmarks, BBQ and SafetyBench, collecting 4,812 total responses across three repeated runs per prompt. Beyond standard performance measures, we evaluate model output dimensions including response consistency, response text similarity, citation grounding, and abstention behavior. For instance, chat UI responses were less accurate than API responses on both benchmarks with search disabled. Enabling web search reduced accuracy by up to 8 percentage points, and even reversed the direction of modality performance trends for one benchmark. Repeated runs of the same prompt produced inconsistent responses in up to 21\% of prompts. The two modalities also grounded answers in different citations, and abstention behavior was also inconsistent across both modalities. These results illustrate that, even within a model family, reporting only simple accuracy metrics can obscure important forms of model behavioral variation relevant to AI safety assessments. We argue that AI safety evaluations should systematically account for modality, multi-run consistency, search conditions, and response-level behaviors to better reflect how deployed AI systems behave in practice.

Comments18 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑