发表机构
The Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对NLPCC 2026任务10引文级忠实性验证,提出集成DeBERTa与类别校准的离线系统,在赛道2取得综合82.99分并排名第二。
AI 中文摘要
本文介绍了我们针对NLPCC 2026共享任务10赛道2的引文级忠实性系统,该任务聚焦于AI辅助科学报告中的引文级忠实性。给定一个原子科学声明及其引用论文的结构化全文,任务要求输出一个四元关系标签以及最多三个证据段落标识符。标签头集成了段落感知的交叉编码器与文档级DeBERTa-large分类器,并随后进行类别级决策校准。概率级融合的动机源于出折(out-of-fold)趋势中过度预测主题匹配(Topical Match)的现象。证据头结合了来自top-20和top-30联合模型的段落分数与BM25分数。系统完全离线运行,不依赖外部检索或LLM提示。在最终排行榜上,我们的系统取得了82.9898的综合得分(Macro-F1为89.5491,Joint@3为76.4305),在赛道2中排名第二。消融实验和错误分析表明,模型互补性和校准推动了标签性能的提升。金标准证据推断对标签Macro-F1的影响可忽略不计,而证据排序对Joint@3仍然重要。
英文摘要
This paper presents our system for Track 2 of the NLPCC 2026 Shared Task 10 on citation-level faithfulness in AI-assisted scientific reporting. Given an atomic scientific claim and the structured full text of its cited paper, the task requires both a four-way relation label and up to three evidence paragraph identifiers. The label head ensembles a paragraph-aware cross-encoder with a document-level DeBERTa-large classifier, followed by class-wise decision calibration. Probability-level fusion is motivated by an out-of-fold tendency to over-predict Topical Match. The evidence head combines paragraph scores from top-20 and top-30 joint models with BM25 scores. The system runs fully offline without external retrieval or LLM prompting. On the final leaderboard, our system achieved 82.9898 overall (89.5491 Macro-F1 and 76.4305 Joint@3), ranking second in Track 2. Ablations and error analysis show that model complementarity and calibration drive the label gains. Gold-evidence inference changes label Macro-F1 negligibly, whereas evidence ranking remains important for Joint@3.
CommentsAccepted to the NLPCC 2026 Shared Task 10 system paper track. Ranked 2nd in Track 2 on the final leaderboard