arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07204cs.CL

HNR-DAC:用于科学主张验证的难负样本重排序与分布对齐分类

HNR-DAC: Hard-Negative Reranking and Distribution-Aligned Classification for Scientific Claim Verification

发表机构南方科技大学 · 深圳先进技术研究院人工智能研究所 · 中国科学院深圳先进技术研究院
查看机构详情
  • Southern University of Science and Technology(南方科技大学)
  • Institute of Artificial Intelligence, Shenzhen University of Advanced Technology(深圳先进技术研究院人工智能研究所)
  • Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)

机构由 AI 辅助整理,请以论文原文为准。

Zhenchao Wang, Xin Chen, Luoxi Zhang, Min Yang, Shiwen Ni

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出HNR-DAC两阶段框架,解决科学主张验证中证据混淆与训练推理分布不匹配问题,在NLPCC 2026任务10赛道2取得优异成绩,相关提交获该赛道第三及最高整体Macro-F1。

中文摘要 AI 辅助

针对引用论文的科学主张验证任务,需要预测主张与论文的关系并识别支持该预测的段落。该场景存在两个关联挑战:论文内的干扰项常与真实证据相似,且在推理阶段,基于黄金证据训练的分类器需对检索到的证据进行操作。本文提出HNR-DAC,这是一个两阶段框架,每个阶段均基于其实际遇到的案例进行训练。难负样本重排序(HNR)利用基础重排序器在非黄金段落上的得分量化证据混淆度,将黄金证据与最易混淆的候选进行对比。分布对齐分类(DAC)使用与构建推理输入时相同的冻结HNR生成的Top-1段落进行训练,同时HNR的Top-3段落标识符作为证据输出。在NLPCC 2026任务10赛道2中,最终配置取得97.21%的Hit@3、95.79%的Macro-F1、94.47%的Joint@3及平均得分95.13%;对应提交在官方赛道2排行榜中排名第三,同时实现最高的整体Macro-F1(93.05%),以及70.16%的Joint@3和平均得分81.61%。

英文摘要

Scientific claim verification over a cited paper requires predicting the claim--paper relation and identifying the paragraphs that justify that prediction. This setting poses two linked challenges: within-paper distractors often resemble genuine evidence, while a classifier trained on gold evidence must operate on retrieved evidence at inference. We present HNR-DAC, a two-stage framework that trains each stage on the cases it will actually encounter. Hard-Negative Reranking (HNR) quantifies evidence confusability using a base reranker's scores on non-gold paragraphs and contrasts gold evidence against the most confusable candidates. Distribution-Aligned Classification (DAC) trains on the Top-1 paragraph produced by the same frozen HNR used to construct inference inputs, while HNR's Top-3 paragraph identifiers provide the evidence output. On the NLPCC 2026 Task 10 Track 2, the final configuration obtains 97.21% Hit@3, 95.79% Macro-F1, 94.47% Joint@3, and an average score of 95.13%. The corresponding submission ranks third on the official Track 2 leaderboard while achieving the highest overall Macro-F1 of 93.05%, alongside 70.16% Joint@3 and an average score of 81.61%.

补充信息

↑