arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

提示框架扭曲了基于计数的LLM错误检测评估:来自数字锚定的证据

Prompt Framing Distorts Count Based Evaluation of LLM Error Detection: Evidence from Numeric Anchoring

Dekun Yang

arXiv 2607.01240首次发表:更新:

发表机构

Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文发现基于计数的F1分数可能因提示中的数字锚定而虚高,与跨度定位能力脱节,提出ErrorBench压力测试协议,并建议评估时应避免预置错误计数并报告跨度感知指标。

AI 中文摘要

基于计数的F1被广泛用作LLM错误检测质量的代理指标,但本文表明,它可能在跨度定位没有相应改进的情况下急剧上升,这种差距被称为F1膨胀。本文引入了ErrorBench,一个用于提示引起的计数失真的受控压力测试协议。ErrorBench在143个CoNLL-2014段落的4290个响应上,评估了五种提示条件下的六个当代LLM。在CoNLL-2014 M2风格评分下,锚定提示产生高达0.79点的F1膨胀,在严格匹配下高达0.96。使用官方ERRANT 3.0.0流水线和多参考评分进行的100段落复制重现了该模式:在六个模型上平均,从盲到锚定的提示转换使Count-F1提高了+0.21,而多参考ERRANT F0.5仅提高了+0.04。研究发现,在此压力测试协议下,高度指令遵循的GPT/Claude系统产生更大的计数响应,而Gemini系列产生更小的计数响应。研究结果表明,LLM校对和文档审阅评估应避免预置错误计数,并应在基于计数的指标之外报告跨度感知指标。

英文摘要

Count-based F1 is widely used as a proxy for LLM error-detection quality, but this paper shows that it can rise dramatically without a corresponding improvement in span localization, a gap termed F1 Inflation. The paper introduces ErrorBench, a controlled stress-test protocol for prompt-induced count distortion. ErrorBench evaluates six contemporary LLMs under five prompt conditions over 4,290 responses from 143 CoNLL-2014 passages. Under CoNLL-2014 M2-style scoring, anchored prompts produce up to 0.79 points of F1 Inflation, and up to 0.96 under strict matching. A 100-passage replication using the official ERRANT 3.0.0 pipeline and multi-reference scoring reproduces the pattern: averaged over six models, the Blind-to-Anchored prompt shift raises Count-F1 by +0.21 while raising multi-reference ERRANT F0.5 by only +0.04. The study finds larger count responses in highly instruction-compliant GPT/Claude systems and smaller responses in the Gemini family under this stress-test protocol. The findings suggest that LLM proofreading and document-review evaluations should avoid pre-populated error counts and should report span-aware metrics alongside count-based metrics.

Comments15 pages, 6 figures, 12 tables. Preprint under review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑