合并还是不合并?评估多参考训练中合并对同行评审的影响
To Consolidate or not to Consolidate? Evaluating the Impact of Consolidation in Multi-Reference Training using Peer Reviews
浏览论文内容
中文总结 AI 辅助
本研究针对自然语言生成中间任务,提出将多样化参考合并为统一训练信号的方法,并构建MERC-36K语料库验证其有效性,实验表明合并参考训练显著优于未合并参考。
中文摘要 AI 辅助
自然语言生成(NLG)任务涵盖了条件熵的整个谱系,从高度受限的机器翻译到开放式的对话生成。结构化任务如自动化同行评审生成则处于中间区域,其中单个输入可对应多个有效且部分重叠的输出。在本工作中,我们证明了传统的单参考和多参考训练范式对于这些中间任务并非最优。我们提供了实证证据,表明将多样化的参考合并为统一的训练信号对于开发有效系统至关重要。为此,我们引入了MERC-36K,一个包含超过36,000篇论文及其原始和合并后同行评审的大规模语料库。利用该数据集,我们训练特定架构以隔离不同参考范式的影响,并与现有最先进系统进行基准比较。通过广泛的自动化和人工评估,我们证明了基于合并参考训练的模型显著优于基于未合并参考训练的模型。数据集和代码将在论文被接收后发布。
英文摘要
Natural language generation (NLG) tasks span the spectrum of conditional entropy, ranging from highly constrained machine translation to open-ended dialogue generation. Structured tasks like automated peer-review generation occupy the intermediate region, where a single input admits multiple valid, overlapping outputs. In this work, we demonstrate that traditional single- and multi-reference training paradigms are suboptimal for these intermediary tasks. We provide empirical evidence that consolidating diverse references into a unified training signal is crucial for developing effective systems. To facilitate this, we introduce MERC-36K, a large-scale corpus of over 36,000 papers paired with original and consolidated peer reviews. Using this dataset, we train specific architectures to isolate the impact of different reference paradigms and benchmark against existing state-of-the-art systems. Through extensive automatic and human evaluation, we demonstrate that models trained on consolidated references significantly outperform those trained on unconsolidated references. Dataset and code will be released upon acceptance.
发表机构
- IIIT Hyderabad(海得拉巴国际信息技术学院)
- Microsoft, India(微软印度)
- Sentisum
机构由 AI 辅助整理,请以论文原文为准。