arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00494cs.CL

基于人类锚定的事实性评估与策略性标注

Human-Anchored Factuality Evaluation with Strategic Annotation

Yu Wang, Craig Erickson, Kevin Small

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对LLM事实性评估的系统性偏差问题,提出基于失败空间分析(FSA)的标注策略,在AutoFA和RAGTruth数据集上分别实现40.3%、27.1%的有效样本量提升,提升了有限预算下的人类锚定事实性评估效率。

中文摘要 AI 辅助

基于大语言模型(LLM)的事实性评估提供了可扩展的评估信号,但其指标相对于人类判断常存在系统性偏差。本文研究有限标注预算下的人类锚定事实性评估,将模型对全数据集的预测与对小部分选择性采样子集的人类标签结合,以获得统计上有效的估计。该方法的效率关键取决于哪些样本获得人类标注:在事实性评估中,模型与人类判断的不一致不仅由低置信度导致,还由结构化失败模式驱动,如证据不完整、时间不匹配、不可验证的主张及评估标准不一致。为利用该结构,本文提出了事实性特定的标注策略设计流程,使用失败空间分析(FSA)推导多种预测信号以建模人类与模型的不一致。在内部基于参考的事实性评估系统AutoFA和RAGTruth上,模型预测的估计值大幅低估了人类标注的事实准确性;本文的FSA引导策略相比均匀采样和不确定性驱动基线,提升了标注效率,在AutoFA上实现了有效样本量提升40.3%,在RAGTruth上提升了27.1%。

英文摘要

LLM-based factuality judges provide scalable evaluation signals, but their metrics are often systematically biased relative to human judgments. We study human-anchored factuality evaluation under limited annotation budgets, where judge predictions on the full dataset are combined with human labels on a small selectively sampled subset to obtain statistically valid estimates. The efficiency of this approach depends critically on which examples receive human annotation: in factuality evaluation, judge-human misalignment is not driven solely by low confidence, but also by structured failure modes such as incomplete evidence, temporal mismatch, unverifiable claims, and rubric misalignment. To exploit this structure, we introduce a factuality-specific annotation policy design pipeline that uses failure-space analysis (FSA) to derive diverse predictive signals for modeling human-judge misalignment. On an internal reference-based factuality evaluation system (AutoFA) and RAGTruth, where judge-predicted estimates substantially underestimate human-annotated factual accuracy, our FSA-guided policy improves annotation efficiency over uniform sampling and uncertainty-driven baselines, achieving effective-sample-size gains of 40.3% on AutoFA and 27.1% on RAGTruth.

发表机构

  • Amazon AGI(亚马逊AGI)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑