arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

信息满意度:面向读者的摘要评估轴

Information Satisfaction: A Reader-Centered Axis for Summarization Evaluation

Isabel Cachola, William Walden, Reno Kriz, Mark Dredze

arXiv 2608.14457首次发表:更新:

发表机构

St. Edward’s University; Johns Hopkins University(圣爱德华大学; 约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出以读者的信息满意度作为摘要评估的面向读者的轴,发现主流摘要指标无法衡量该满意度,且与人工判断一致性差。

AI 中文摘要

当前摘要评估的多数研究聚焦于通用摘要质量(如ROUGE、BERTScore)或特定期望属性(如可读性、事实性),但这些指标无法衡量摘要对单个用户的效用。例如,了解最新疫苗研究的生物医学研究者与家庭医生的信息需求存在差异。面向查询的摘要捕捉了部分此类需求,但实际中用户很少在查询中陈述所有相关内容:单个简短查询可能不足以区分研究者与医生的需求。相比之下,读者的背景或角色(其职责与专业能力)在不同查询中相对稳定,能恢复大部分缺失的上下文,因此是评估摘要是否满足该读者需求的实用信号。本研究评估了主流摘要指标对信息和角色差异的敏感性,发现包括强大的LLM-as-judge指标在内的许多主流指标,无法通过信息内容的基础扰动测试。此外,我们开展了专家人工评估,基于特定人的背景和使用场景,测量对满足信息满意度的摘要偏好,结果发现传统指标和基于LLM的指标均无法充分衡量信息满意度,且与人工判断的一致性较差。

英文摘要

The majority of work on summarization evaluation focuses on general summary quality (e.g., ROUGE, BERTScore) or specific desired properties (e.g., readability, factuality). However, these metrics fail to measure the utility of a summary to an individual user. For example, a biomedical researcher learning about the latest vaccine research will have different informational needs from a family doctor. Query-focused summarization captures part of this need, but in practice, users rarely state everything relevant in a query: a single short query is likely inadequate to distinguish the needs of a researcher from those of a physician. By contrast, a reader's background or persona (their role and expertise) is comparatively stable across queries and recovers much of this missing context, which makes it a practical signal for assessing whether a summary satisfies that reader's needs. In this work, we assess how sensitive popular summarization metrics are to both informational and persona differences, and find that many popular metrics, including strong LLM-as-judge metrics, fail basic perturbation tests of informational content. We additionally conduct an expert human evaluation, measuring summary preferences based on information satisfaction given a specific person's background and use case. We find that both traditional and LLM-based metrics are insufficient measures of information satisfaction and agree poorly with human judgment.

DOI:10.1145/3834580.3838749

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑