经典语义抽取式摘要能否在印地语中评估?一项复制研究
Can Classical Semantic-Extractive Summarization Be Evaluated in Hindi? A Replication Study
- University of Kashmir(克什米尔大学)
- IBM Research(IBM研究院)
- Manipal University Jaipur(马尼帕尔大学斋浦尔分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究复制并适配了经典语义抽取式摘要方法至印地语,发现其性能显著低于领先基线,特征消融表明仅句子位置有效,印地语基准无法奖励非领先内容,需构建专用评估资源。
AI中文摘要:
我们复制了Mohd、Jan和Shah(2020)的分布语义抽取式摘要方法,并将其适配到印地语,在每个语言特定步骤中替换为适合天城文的组件。该系统在两个独立语料库上进行了评估——XL-Sum的印地语部分和FIRE ILSUM 2.0印地语——使用一种针对天城文感知的ROUGE实现,该实现已针对XL-Sum作者自己的多语言评分器进行了验证,所有比较均采用1000次重采样的配对自助法。在其已发布的等权配置中,复制系统在两个语料库上均显著差于三句领先基线,在XL-Sum上ROUGE-1 F值落后Lead-3 0.042,在ILSUM上落后0.265。特征消融显示,句子位置是唯一有贡献的特征:仅位置特征就能精确复现领先基线,移除位置则产生最弱配置,而验证调优的权重最多只能与Lead-3持平,从未超过。TextRank同样失败,这使得该结果成为类级而非实现级的结果。选择分析表明,其余特征将提取导向长句、实体密集的正文句子,而参考摘要则重用文章的开头。因此,印地语基准无法奖励非领先内容的选择,这促使需要专门构建的评估资源。
英文摘要:
We replicate the distributional-semantics extractive summarisation method of Mohd, Jan and Shah (2020) and adapt it to Hindi, substituting a Devanagari-appropriate component at every language-specific step. The system is evaluated on two independent corpora --- the Hindi portion of XL-Sum and FIRE ILSUM 2.0 Hindi --- under a Devanagari-aware ROUGE implementation validated against the XL-Sum authors' own multilingual scorer, with all comparisons drawn as 1000-resample paired bootstraps. In its published equal-weight configuration the replicated system is significantly worse than a three-sentence lead baseline on both corpora, trailing Lead-3 by 0.042 ROUGE-1 Fon XL-Sum and by 0.265 on ILSUM. A feature ablation shows that sentenceposition is the only feature that contributes: position alone reproduces the lead baseline exactly, removing position gives the weakest configuration,and a validation-tuned weighting can at best equal Lead-3 and never exceed it. TextRank fails identically, making this a class-level rather than an implementation-level result. A selection analysis shows the remaining features steer extraction towards long, entity-dense body sentences while the references reuse the article lead.Current Hindi benchmarks therefore cannot reward non-lead content selection, motivating purpose-built evaluation resources.