arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

文本何时提供信息?多模态时间序列预测的信息论指标基准测试

When Does Text Inform? Benchmarking Information-Theoretic Metrics for Multimodal Time-Series Forecasting

Emma Andrews, Gianmarco Mengaldo

arXiv 2609.11282首次发表:更新:

发表机构

National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究创建合成基准,评估六种信息论指标在多模态时间序列预测中衡量文本注释预测价值的能力,并验证其可审计注释质量、选择最优注释,无需模型训练。

AI 中文摘要

将时间序列与文本注释相结合的多模态预测模型有望通过文本上下文实现更丰富的预测,但我们如何知道文本注释是否对预测器的预测有实质性贡献?这是一个信息论问题,但要评估信息论指标能否可靠地衡量注释所提供的预测价值,需要一个真实基准,而目前尚不存在。我们创建了一个合成时间序列信号,其注释分为三类:语义正确、语义错误和无关。由于数据生成过程完全受控,真实信息内容可精确得知,从而能够对六种互补的互信息估计器(KSG、MINE、InfoNCE、CCA、PID 和 V-information)进行原则性评估。我们表明,所有六种估计器都能将正确注释识别为信息量最大的注释,并且能够审计混合文本语料库的质量,选择能带来最佳下游预测结果的注释,而无需进行模型训练。我们的基准揭示了每种估计器的局限性,并在七个真实世界数据集上进行了验证,这些数据集展示了估计器性能在弱信号上的差异。最后,我们为将这些指标用于注释审计和融合选择制定了实用规则。

英文摘要

Multimodal forecasting models that combine time series with text annotations promise richer prediction through textual context, but how do we know whether a text annotation meaningfully contributes to the forecasters prediction? This is an information-theoretic question, but to evaluate whether information-theoretic metrics can reliably measure the predictive value an annotation provides, a ground truth benchmark is needed, and none currently exist. We create a synthetic time series signal with annotations in three categories: semantically correct, incorrect, and irrelevant. Because the data generation process is fully controlled, ground-truth information content is known exactly, enabling principled evaluation of six complementary mutual information estimators (KSG, MINE, InfoNCE, CCA, PID and V-information). We show that all six estimators identify correct annotations as most informative, and are able to audit the quality of mixed text corpora, choosing the annotations that result in the best downstream forecasting results without the need for model training. Our benchmark identifies limitations of each estimator, and these are validated on seven real-world datasets, which show how estimator performance differs on weak signals. Finally, we establish practical rules for implementing these metrics for annotation auditing and fusion selection.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑