迈向稳健的数值声明验证
Towards Robust Numerical Claim Verification
浏览论文内容
中文总结 AI 辅助
针对LLM数值声明验证的脆弱性,通过对抗性微调小型Qwen3模型,在标签翻转扰动上达到98.7%准确率,并泛化至未见扰动和跨语言场景。
中文摘要 AI 辅助
大型语言模型(LLMs)广泛用于声明验证,但在数值推理方面仍然脆弱:即使数值的微小变化也会急剧降低准确率。我们表明,这种脆弱性在前沿LLMs中依然存在,但可以通过对数值扰动示例进行对抗性微调来缓解。使用参数高效微调,小型Qwen3模型(0.6B–8B)在标签翻转扰动上达到98.7%的准确率,优于更大的零样本模型和前沿系统(GPT-5.4 Pro(74.0%)和Gemini 2.5 Flash(73.9%))。这些增益泛化到未见过的扰动类型,表明模型学到了稳健的数值决策边界而非记忆编辑。鲁棒性还能在没有目标域数据的情况下迁移,显著提高了西班牙语的跨语言性能。我们进一步表明,相同的微调方法也能赋予对证据侧扰动的鲁棒性,使用VitaminC数据集。
英文摘要
Large language models (LLMs) are widely used for claim verification, yet remain brittle for numerical reasoning: even small changes in value can sharply degrade accuracy. We show that this brittleness persists in frontier LLMs, but can be mitigated through adversarial fine-tuning on numerically perturbed examples. Using parameter-efficient fine-tuning, small Qwen3 models (0.6B$\unicode{x2013}$8B) reach 98.7% accuracy on label-flipping perturbations, outperforming larger zero-shot models and frontier systems (GPT-5.4 Pro (74.0%) and Gemini 2.5 Flash (73.9%)). The gains generalise to unseen perturbation types, indicating robust numerical decision boundaries rather than memorised edits. Robustness also transfers without target-domain data, significantly improving cross-lingual performance in Spanish. We further show that the same fine-tuning recipe confers robustness to evidence-side perturbations, using the VitaminC dataset.
发表机构
- University of Stavanger(斯塔万格大学)
- Factiverse AI(Factiverse AI公司)
机构由 AI 辅助整理,请以论文原文为准。