发表机构
University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型总结科学文本时数字和单位幻觉问题,提出将错误检测重定义为类型验证,引入分类法和基准。用符号增强训练框架,提升规范等价鲁棒性,提高准确率,还明确了符号特征及银标签在特定情况下的作用,确定训练时增强是关键集成点。
AI 中文摘要
大语言模型在总结科学文本时会出现数字和单位幻觉,这可能会使科学主张悄然反转。我们将此类错误检测重新定义为类型验证:引入了一个五类类型化数量错误分类法和一个1500项基准,该基准从PMC和arXiv来源重写,并由两名独立的大语言模型注释者标注(Krippendorff's alpha = 0.882)。在该基准上微调的ModernBERT编码器达到了宏观F1 = 0.899,但四个探测器暴露了一个明显的结构盲点:对于物理等价量的规范等价重写,其准确率降至36.5%。我们提出了符号增强,这是一个训练时框架,它反向运行符号验证器的模块以生成保留标签的增强训练数据。增强将规范等价鲁棒性提高到98.2%,同时略微提高了分布内准确率(宏观F1:从0.899提高到0.902);增强后的编码器在无推理成本的情况下与前沿大语言模型匹配,并转移到外部基准(SciFact-Open二进制宏观F1:从0.791提高到0.828)。两个负面结果强化了这一观点:符号特征作为辅助编码器输入没有作用,并且符号银标签在教师噪声下呈负比例缩放。这些结果共同确定训练时增强是符号和学习组件之间正确的集成点。
英文摘要
Large language models hallucinate numbers and units when summarizing scientific text, a failure mode that can silently invert a scientific claim. We recast the detection of such errors as typed verification: we introduce a five-class typed-quantity error taxonomy and a 1500-item benchmark, rewritten from PMC and arXiv sources and labeled by two independent LLM annotators with adjudication (Krippendorff's alpha = 0.882). A ModernBERT encoder fine-tuned on this benchmark reaches macro-F1 = 0.899, far above any off-the-shelf neural fact-checker, yet four probes expose a sharp structural blind spot: on canonical-equivalent rewrites of physically equivalent quantities (e.g., 95°C and 368.15 K) its accuracy collapses to 36.5%. We propose Symbolic Augmentation, a training-time framework that runs the modules of a symbolic verifier in reverse to generate label-preserving augmented training data. The augmentation lifts canonical-equivalence robustness to 98.2% while slightly improving in-distribution accuracy (macro-F1: 0.899 to 0.902); the augmented encoder matches a closed-frontier LLM at no inference cost and transfers to an external benchmark (SciFact-Open binary macro-F1: 0.791 to 0.828). Two negative results sharpen the claim: symbolic features as auxiliary encoder inputs add nothing, and symbolic silver labels scale negatively under teacher noise. Together these results identify training-time augmentation as the right integration point between symbolic and learned components.
Comments18 pages, 3 figures, 6 tables