arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向预算约束的忠实摘要生成:解耦生成与选择

Decoupling Generation and Selection for Budget-Constrained Faithful Summarization

Zeyu Wang, Guanghua Wang, Meng Xu

arXiv 2608.03655首次发表:更新:

AI 中文总结

该研究针对摘要式模型的事实不一致等问题,提出解耦生成与选择的模块化框架,在多数据集上提升了事实性,仅参考重叠分数略有降低。

AI 中文摘要

摘要式摘要模型仍存在事实不一致、冗余性和长度控制弱的问题。我们提出一种面向句子预算约束的模块化生成-选择框架。预训练生成器生成多个候选摘要,将其分解为句子级候选;组合选择器则在明确预算约束下,通过平衡相关性、事实性和冗余性构建最终摘要。该框架支持MMR、ILP及受DPP启发的对数行列式目标,无需对生成器进行再训练。在CNN/DailyMail、Multi-News、FaithBench和TofuEval上的实验表明,该框架在事实性和源接地指标上实现了持续提升,尤其在多文档摘要任务中表现突出,仅参考重叠分数有所降低。人工评估进一步显示,该框架在感知一致性、相关性、清晰度和简洁性上表现更优,仅连贯性略有下降。这些结果表明,将生成与选择解耦,提供了一种与模型无关的机制,可提升事实接地性。代码可在此URL获取。

英文摘要

Abstractive summarization models remain vulnerable to factual inconsistency, redundancy, and weak length control. We propose a modular generation-and-selection framework for sentence-budget-constrained summarization. A pretrained generator produces multiple candidate summaries, which are decomposed into sentence-level candidates. A combinatorial selector then constructs the final summary by balancing relevance, factuality, and redundancy under an explicit budget. The framework supports MMR, ILP, and a DPP-inspired log-determinant objective without retraining the generator. Experiments on CNN/DailyMail, Multi-News, FaithBench, and TofuEval show consistent improvements in factuality and source-grounding metrics, especially for multi-document summarization, at the cost of lower reference-overlap scores. Human evaluation further indicates higher perceived consistency, relevance, clarity, and conciseness, with a small reduction in coherence. These results show that decoupling generation from selection provides a model-agnostic mechanism for improving factual grounding. Code is available at https://anonymous.4open.science/r/bcfs-D05E/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑