arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20116cs.CL

当文本与数字不一致时:大型语言模型中的证据仲裁

When Text and Numbers Disagree: Evidence Arbitration in Large Language Models

发表机构牛津大学 · 葛兰素史克 · 牛津大学苏州高等研究院
查看机构详情
  • University of Oxford(牛津大学)
  • GlaxoSmithKline(葛兰素史克)
  • Oxford Suzhou Centre for Advanced Research(牛津大学苏州高等研究院)

机构由 AI 辅助整理,请以论文原文为准。

Mattia Carletti, Edward Phillips, Fredrik K. Gustafsson, Patitapaban Palo, Lei Clifton, Danielle Belgrave, Xiao Gu, David A. Clifton

首次发表
浏览论文内容

中文总结 AI 辅助

该研究构建合成基准探究LLMs在文本与数字证据冲突时的仲裁行为,发现其存在系统性偏好,常依赖启发式策略,凸显工具增强决策系统的失效模式。

中文摘要 AI 辅助

大型语言模型(LLMs)越来越多地应用于文本摘要、数值观测和外部工具输出可能提供冲突证据的场景。我们研究当这些证据来源支持对立决策时,LLMs如何在它们之间进行仲裁。为此,我们引入了一个受控的合成基准,其中潜在风险轨迹同时生成数值时间序列和自然语言摘要,使我们能够构建冲突,其中恰好一个证据来源与真实标签对齐。该设计允许我们独立操纵模态、时间新近度、来源可靠性和证据来源。在开源权重的指令调优模型中,我们发现仲裁行为是系统性的而非随机的:模型表现出明显的文本-数字偏好,比显式可靠性线索更一致地遵循时间新近度,甚至在与直接上下文证据冲突时也可能过度依赖外部预测。这些结果表明,当前LLMs在整合异构证据时通常依赖启发式仲裁策略,凸显了工具增强决策系统的一种失效模式。

英文摘要

Large language models (LLMs) are increasingly used in settings where textual summaries, numerical observations, and external tool outputs may provide conflicting evidence. We study how LLMs arbitrate between such sources when they support opposing decisions. To do so, we introduce a controlled synthetic benchmark in which latent risk trajectories generate both numerical time series and natural language summaries, allowing us to construct conflicts where exactly one evidence source is aligned with the ground-truth label. This design lets us independently manipulate modality, temporal recency, source reliability, and evidence provenance. Across open-weight instruction-tuned models, we find that arbitration behaviour is systematic rather than random: models exhibit distinct text-versus-number preferences, follow temporal recency more consistently than explicit reliability cues, and can over-rely on external forecasts even when they conflict with direct contextual evidence. These results suggest that current LLMs often rely on heuristic arbitration strategies when integrating heterogeneous evidence, highlighting a failure mode for tool-augmented decision systems.

↑