arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22321cs.CL

语义还是结构?多模态时间序列预测中的文本敏感性审计

Semantics or Structure? Auditing Text Sensitivity in Multimodal Time-Series Forecasting

Karthik Sridhar, Atharva Gupta, Nishant Pradhan, Murari Mandal, Dhruv Kumar, Saurabh Deshpande

首次发表
浏览论文内容

中文总结 AI 辅助

本研究审计多模态时间序列预测模型的文本敏感性,发现文本内容并非其在Time-MMD基准上性能提升的有效信号,并发布了相关诊断工具。

中文摘要 AI 辅助

多模态时间序列预测已成为一种颇具前景的范式,人们期望自然语言上下文能提升预测性能。近期的多模态基础模型包括Aurora,以及早融合、晚融合方法如MM-TSFlib和TaTS,在Time-MMD基准上相比单模态基线取得了显著提升,并将这些改进归因于文本信息。然而,这些模型是否真的对文本的语义内容敏感仍未得到验证。我们通过受控文本扰动、归因分析和对Aurora文本通路的探测来解决该问题。在Time-MMD上,将每一行的文本替换为其他任意真实文本(空文本、常量文本、域内打乱文本或跨域文本),三种架构的平均均方误差(MSE)变化均小于0.5%。当移除附带的数值列而不触及文本时,文献中报告的改进得以恢复。我们得出结论:在该基准及这类冻结编码器架构中,文本内容并非报告的改进背后的有效信号。为支持未来结构化数据多模态基础模型中文本整合的研究,我们发布了扰动协议和评估工具包作为可复用的诊断工具。

英文摘要

Multimodal time-series forecasting has emerged as a promising paradigm in which natural-language context is expected to improve predictive performance. Recent multimodal foundation models, including Aurora, as well as early- and late-fusion approaches such as MM-TSFlib and TaTS, report substantial gains over unimodal baselines on the Time-MMD benchmark, attributing these improvements to textual information. However, whether these models are actually sensitive to the semantic content of the text remains unverified. We address this question through controlled text perturbations, attribution analyses, and probes of Aurora's text pathway. On Time-MMD, swapping each row's text for any other real text (empty, constant, within-domain shuffled, or cross-domain) moves mean MSE by less than $0.5\%$ on all three architectures. The improvement reported in the literature is recovered when a co-shipped numeric column is removed without touching text. We conclude that, on this benchmark and within this family of frozen-encoder architectures, text content is not the operative signal behind the reported gains. To support future work on text integration in multimodal foundation models for structured data, we release our perturbation protocol and evaluation harness as a reusable diagnostic toolkit.

发表机构

  • Birla AI Labs(贝拉人工智能实验室)
  • BITS Pilani(皮拉尼比特斯学院)
  • KIIT(基伊特学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑