AI 中文总结
该研究针对摘要式模型的事实不一致等问题,提出解耦生成与选择的模块化框架,在多数据集上提升了事实性,仅参考重叠分数略有降低。
AI 中文摘要
摘要式摘要模型仍存在事实不一致、冗余性和长度控制弱的问题。我们提出一种面向句子预算约束的模块化生成-选择框架。预训练生成器生成多个候选摘要,将其分解为句子级候选;组合选择器则在明确预算约束下,通过平衡相关性、事实性和冗余性构建最终摘要。该框架支持MMR、ILP及受DPP启发的对数行列式目标,无需对生成器进行再训练。在CNN/DailyMail、Multi-News、FaithBench和TofuEval上的实验表明,该框架在事实性和源接地指标上实现了持续提升,尤其在多文档摘要任务中表现突出,仅参考重叠分数有所降低。人工评估进一步显示,该框架在感知一致性、相关性、清晰度和简洁性上表现更优,仅连贯性略有下降。这些结果表明,将生成与选择解耦,提供了一种与模型无关的机制,可提升事实接地性。代码可在此URL获取。
英文摘要
Abstractive summarization models remain vulnerable to factual inconsistency, redundancy, and weak length control. We propose a modular generation-and-selection framework for sentence-budget-constrained summarization. A pretrained generator produces multiple candidate summaries, which are decomposed into sentence-level candidates. A combinatorial selector then constructs the final summary by balancing relevance, factuality, and redundancy under an explicit budget. The framework supports MMR, ILP, and a DPP-inspired log-determinant objective without retraining the generator. Experiments on CNN/DailyMail, Multi-News, FaithBench, and TofuEval show consistent improvements in factuality and source-grounding metrics, especially for multi-document summarization, at the cost of lower reference-overlap scores. Human evaluation further indicates higher perceived consistency, relevance, clarity, and conciseness, with a small reduction in coherence. These results show that decoupling generation from selection provides a model-agnostic mechanism for improving factual grounding. Code is available at https://anonymous.4open.science/r/bcfs-D05E/.