DS@GT ARC在CheckThat! 2026中的应用:基于大语言模型的多语言数值声明验证的推理轨迹排序和分组奖励建模
DS@GT ARC at CheckThat! 2026: LLM-Based Trace Ranking and Grouped Reward Modeling for Multilingual Numerical Claim Verification
浏览论文内容
中文总结 AI 辅助
针对多语言数值声明验证,探索基于大语言模型(LLM)的推理轨迹排序和分组奖励建模两种方法。LLM方法用LoRA微调验证器评分,还尝试子声明分解;奖励模型用TF-IDF及特征评分。结果显示LLM方法多数指标更优,阿拉伯语中AraBERT表现更佳且子声明分解未提升性能。
中文摘要 AI 辅助
数值声明的自动验证是一个具有挑战性的问题,因为它需要语言理解和定量推理。本文描述了我们针对CLEF 2026 CheckThat!任务2的系统,该任务专注于对大语言模型(LLM)生成的推理轨迹进行排序,并预测英文和阿拉伯文数值声明的最终判定。我们探索了两种方法。第一种方法使用LoRA对基于LLM的验证器进行微调,将每个推理轨迹作为二元分类问题独立评分,并使用Best-of-N选择来选择最终判定。我们还对自适应子声明分解进行了实验,以便在验证前将复杂声明分解为更简单的部分。第二种方法使用具有手工制作的数值和时间重叠特征的轻量级TF-IDF奖励模型对轨迹进行评分,并按判定组聚合分数以确定最终预测。对于阿拉伯语,我们将通用多语言模型与在阿拉伯语文本上预训练的特定语言模型AraBERT进行了比较。我们的结果表明,基于LLM的方法在大多数指标上优于轻量级奖励模型,特别是在Recall@5方面,而基于奖励的方法在冲突类上表现更强。子声明分解并没有提高性能,这表明声明拆分引入了噪声而不是有助于推理。对于阿拉伯语,AraBERT在大多数指标上优于多语言基线。
英文摘要
Automated verification of numerical claims is a challenging problem, as it requires both language understanding and quantitative reasoning. This paper describes our system for CLEF 2026 CheckThat! Task 2, which focuses on ranking reasoning traces generated by large language models (LLMs) and predicting a final verdict for numerical claims in English and Arabic. We explore two approaches. The first approach fine-tunes an LLM-based verifier using LoRA to score each reasoning trace independently as a binary classification problem, and selects the final verdict using Best-of-N selection. We further experiment with adaptive sub-claim decomposition to break complex claims into simpler parts before verification. The second approach uses a lightweight TF-IDF reward model with handcrafted numeric and temporal overlap features to score traces, and aggregates scores by verdict group to determine the final prediction. For Arabic, we compare a general multilingual model against AraBERT, a language-specific model pretrained on Arabic text. Our results show that the LLM-based approach outperforms the lightweight reward model on most metrics, particularly Recall@5, while the reward-based approach shows stronger performance on the Conflicting class. Sub-claim decomposition did not improve performance, suggesting that claim splitting introduces noise rather than aiding reasoning. For Arabic, AraBERT outperforms the multilingual baseline across most metrics.
发表机构
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。