发表机构
Wuhan University; Nanjing Audit University; The University of Manchester; MBZUAI; McGill University; Shanghai Jiao Tong University(武汉大学; 南京审计大学; 曼彻斯特大学; 穆罕默德·本·扎耶德人工智能大学; 麦吉尔大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TimeLitmus是一个诊断基准,通过4,856条记录和反事实干预测试,揭示LLM在事件条件时间序列预测中跨模态理解与解释忠实性的不足,并证明人类优于模型。
AI 中文摘要
大型语言模型(LLMs)越来越多地被用于根据数值时间序列历史和文本事件进行预测。然而,仅凭准确性无法揭示正确答案是反映了对两种输入的有效整合,还是源于事件极性、单模态先验或表面线索。同样,看似合理的解释可能使预测合理化,却并未忠实反映驱动模型行为的证据。我们引入了TimeLitmus,一个用于事件条件时间序列预测中跨模态理解与解释忠实性的诊断基准。TimeLitmus包含金融和交通领域的4,856条评估记录,将自然预测与受控反事实和对比干预、针对解释的忠实性测试以及系统性捷径控制相结合。在十个代表性LLM中,标准预测准确性显著高估了可靠的跨模态理解:硬配对对比(HPC)配对正确率在金融领域最高仅为19.2%,在交通领域为11.7%,且所有十个模型在金融序列侧控制上均表现出低于预期的一致性。模型通常能明确识别情景关系,却在独立预测时未能应用它们。解释忠实性显示出类似的差距:在交通领域,大多数模型在超过90%的案例中引用了被操纵的时间因素,而行为支持率仍低于22%。人类标注者在匹配的受控和硬配对诊断上优于LLM,证实这些区别可从输入中恢复。仅自然适应在证据选择和信息敏感性上带来选择性提升,但在受控或硬配对行为上未带来一致提升。该基准、评估套件和监督适应数据将公开发布。
英文摘要
Large language models (LLMs) are increasingly used to make predictions from numerical time-series histories and textual events. Yet accuracy alone cannot reveal whether correct answers reflect effective integration of the two inputs or instead arise from event polarity, unimodal priors, or superficial cues. Likewise, plausible explanations may rationalize predictions without faithfully reflecting the evidence that drives model behavior. We introduce TimeLitmus, a diagnostic benchmark for cross-modal understanding and explanation faithfulness in event-conditioned time-series prediction. TimeLitmus contains 4,856 evaluation records across Finance and Traffic, combining natural prediction with controlled counterfactual and contrastive interventions, explanation-targeted faithfulness tests, and systematic shortcut controls. Across ten representative LLMs, standard prediction accuracy substantially overstates reliable cross-modal understanding: Hard Paired Contrast (HPC) pair correctness peaks at only 19.2% in Finance and 11.7% in Traffic, and all ten models show lower-than-expected consistency on Finance series-side controls. Models often recognize scenario relations explicitly yet fail to apply them during independent prediction. Explanation faithfulness shows a similar gap: in Traffic, most models cite the manipulated temporal factor in over 90% of cases, while behavioral support remains below 22%. Human annotators outperform LLMs on matched controlled and hard-pair diagnostics, confirming that these distinctions are recoverable from the inputs. Natural-only adaptation yields selective gains in evidence selection and input sensitivity, but not consistent gains in controlled or hard-pair behavior. The benchmark, evaluation suite, and supervised adaptation data will be released publicly.