保留重要内容:摘要评估中超越饱和度的语义脚手架
Preserving What Matters: Semantic Scaffolds Beyond Saturation in Summarization Evaluation
- ServiceNow Canada(ServiceNow加拿大)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对摘要评估中ROUGE和LLM-as-judge评分饱和失效的问题,提出语义脚手架框架,通过层次化提取事实、问题和实体属性并生成三个诊断指标,在传统指标失效时仍能有效区分模型性能。
AI中文摘要:
摘要生成已广泛应用于众多生产系统,使得模型选择成为一项常规决策,而这一决策依赖于对摘要质量的衡量。现有指标难以支撑这一需求:ROUGE仅捕捉表面重叠,而LLM-as-judge评分则饱和至近乎相同的数值,无法有效对模型进行排序。我们在三个公开数据集、两个专有数据集以及多语言设置中观察到了这种饱和现象。受此启发,我们提出了语义脚手架(Semantic Scaffold),一个评估框架,它从源文本中提取事实、问题和实体属性的层次化表示,将每个元素标记为主要观点或支撑细节,并将该结构作为固定参考来对摘要进行评分。基于这一表示,我们推导出三个诊断性指标:事实保留分数(FPS)、问题保留分数(QPS)和实体保留分数(EPS),旨在奖励对关键信息的保留,同时惩罚细节过载,并将其定位为可解释的诊断工具,在整体评估维度失效时仍能提供有效信息。最后,我们分析了ROUGE和LLM-as-judge评分的四种常见失败模式,证明基于脚手架的评估在传统指标失效的情况下仍能保持信息量。
英文摘要:
Summarization ships in countless production systems, making model selection a routine decision that depends on measuring summary quality. Existing metrics struggle to support this: ROUGE captures only surface overlap, while LLM-as-judge scores saturate to near-identical values that fail to rank models effectively. We observe this saturation across three public datasets, two proprietary datasets, and multilingual settings. Motivated by this, we introduce Semantic Scaffold, an evaluation framework that extracts a hierarchical representation of facts, questions, and entity attributes from a source text, labeling each as a main point or supporting detail, and reusing this structure as a fixed reference for scoring summaries. From this representation, we derive three diagnostic metrics: Fact Preservation Score (FPS), Question Preservation Score (QPS), and Entity Preservation Score (EPS), designed to reward the preservation of essential information while penalizing detail overload, and position them as interpretable diagnostics that remain informative where holistic axes collapse. Finally, we analyze four recurring failure modes of ROUGE and LLM-as-judge scores, demonstrating that scaffold-based evaluation remains informative where conventional metrics collapse.