arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10806cs.CLcs.AI

用于评估文本摘要的抽象性度量:一种经过实证验证的精确公式

Abstractiveness Metrics for Evaluating Text Summarization: A Refined Formulation with Empirical Validation

  • International Institute of Information Technology, Bhubaneswar(布巴内斯瓦尔国际信息技术学院)
  • Salesforce India Pvt Ltd(Salesforce印度私人有限公司)

机构由 AI 辅助整理,请以论文原文为准。

Praveenkumar Katwe, Rakesh Chandra Balabantaray, Kali Prasad Vittala

AI总结:

研究旨在量化文本摘要抽象性,引入RA、SA和AR度量,利用文档长度调和平均及非重叠因子公式,在四个模型上评估100个XSUM文档,成功区分提取式与抽象式模型,抽象率可识别需人工评估的摘要。

AI中文摘要:

量化生成摘要中的抽象性对于评估超越像ROUGE这样的表面级度量的摘要模型至关重要。我们引入了参考抽象(RA)、摘要抽象(SA)和抽象率(AR)——一组有原则的启发式度量,用于衡量摘要与源文本的提取式复制的差异程度。该公式使用由三次非重叠因子调制的文档长度的调和平均值,产生维度一致、有界的输出,对提取式 - 抽象式边界具有非线性敏感性。在四个摘要模型(BART - large - cnn、Pegasus - xsum、DistilBart、MT5 - small)上对100个XSUM文档进行的评估表明,这些度量成功地区分了提取式模型(SA约为0.12 - 0.26)和抽象式模型(SA约为0.96 - 1.77),并且抽象率识别出需要人工评估是否存在潜在幻觉的摘要。代码和结果可在这个https网址获取。

英文摘要:

Quantifying abstractiveness in generated summaries is essential for evaluating summarization models beyond surface-level metrics like ROUGE. We introduce Reference Abstraction (RA), Summary Abstraction (SA), and Abstraction Ratio (AR) -- a set of principled heuristic metrics that measure how much a summary diverges from extractive copying of the source text. The formulation uses the harmonic mean of document lengths modulated by a cubic non-overlap factor, yielding dimensionally consistent, bounded output with non-linear sensitivity to the extractive-abstractive boundary. Evaluation on 100 XSUM documents across four summarization models (BART-large-cnn, Pegasus-xsum, DistilBart, MT5-small) demonstrates that the metrics successfully discriminate between extractive models (SA ~ 0.12-0.26) and abstractive models (SA ~ 0.96-1.77), and that the Abstraction Ratio identifies summaries requiring manual evaluation for potential hallucination. Code and results are available at https://github.com/katweNLP/AbstractionStudy.

补充信息

↑