arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SciTrue:基于前沿开源语言模型的可靠科学声明验证——NTCIR SciClaimEval任务实践

SciTrue: Reliable Scientific Claim Validation with Frontier and Open Language Models at the NTCIR SciClaimEval Task

Qiming Bao, Neşet Özkan Tan, Siyuan Wang, Mark Gahegan

arXiv 2609.00654首次发表:更新:

发表机构

University of Auckland; The Chinese University of Hong Kong(奥克兰大学; 香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文介绍SciTrue团队参与NTCIR-19 SciClaimEval任务,通过基准测试前沿开源多模态模型、采用无泄露配对先验等方法,在官方盲测中取得优异成绩,同时发现了测量泄露等关键问题。

AI 中文摘要

本文介绍了SciTrue团队参与NTCIR-19 SciClaimEval任务两个子任务的情况,该任务要求系统依据论文的表格和图表验证科学声明。我们未对单一模型进行调优,而是采用统一的、逐样本协议对11个前沿开源多模态模型进行基准测试,并将其与轻量透明的后处理相结合。在官方盲测排行榜(结果部分)中,SciTrue在四个证据类别/子任务组合中的三个以明显优势排名第一,在第四个组合的主要指标上并列第一。该结果源于三项发现:第一,强大的指令调优模型已具备竞争力,Claude Opus~4.8和Gemma-4-31B均超过最强的公开基线o4-mini,GPT-5.5和Claude Fable~5在两个子任务中领先(子任务2得分为97.7);第二,任务的配对结构是最大的关键因素,一种“无泄露配对先验”可仅从声明文本(可见字段)中恢复“支持/反驳”配对,并将“支持”分配给置信度更高的证据,这使子任务1的配对准确率从72.2提升至93.5,效果远超任何模型交换或集成加权;第三,逐案审计发现,大部分剩余错误是视觉无法检测的标签映射交换或数据集标签噪声,因此测得的准确率低估了真实能力,可通过建模修复的空间很小。受控微调、蒸馏和智能体一致性检查支持相同结论,我们还全程记录了一种测量泄露——标签信息通过数据的包装而非内容传递给系统,其中发布的文件顺序编码了标签,包括一个曾短暂误导我们管道的实例。

英文摘要

We describe the SciTrue team's participation in both subtasks of the NTCIR-19 SciClaimEval task~\cite{sciclaimeval}, which asks systems to verify scientific claims against the tables and figures of a paper. Rather than tuning a single model, we benchmark eleven frontier and open multimodal models under one honest, per-sample protocol and combine them with light, transparent post-processing. On the official, blind test leaderboard (Section~\ref{sec:results}), SciTrue placed first by a clear margin in three of the four evidence-category/subtask combinations, and tied for first on the primary metric in the fourth. Three findings explain the result. First, strong instruction-tuned models are already competitive: Claude Opus~4.8 and Gemma-4-31B each exceed the strongest public baseline (o4-mini), and GPT-5.5 and Claude Fable~5 lead both subtasks (97.7 on Subtask~2). Second, the task's pairing structure is the largest lever: a \emph{leak-free pair prior} that recovers the Supported/Refuted pairing from the claim text alone (a visible field) and assigns Supported to the higher-confidence evidence raises Subtask-1 pair-accuracy from 72.2 to 93.5, far more than any model swap or ensemble weighting. Third, a case-by-case audit finds that most residual errors are visually-undetectable label-mapping swaps or dataset label noise, so measured accuracy understates the true ability and the fixable-by-modeling headroom is small. Controlled fine-tuning, distillation, and agentic consistency-checking support the same conclusions, and we document throughout a measurement leak---label information reaching a system through the packaging of the data rather than its content---in which the released file ordering encodes the label, including one instance that briefly misled our own pipeline.

CommentsTo appear in the Proceedings of the 19th NTCIR Conference (NTCIR-19)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑