模态应相互协作:用于长期流感预测的双流多模态学习
Modalities Should Talk to Each Other: Dual-Stream Multimodal Learning for Long-Horizon Influenza Forecasting
浏览论文内容
中文总结 AI 辅助
该研究针对长期流感预测问题,提出双流注意力(DSA)多模态框架,通过双向跨模态注意力耦合数值与文本流,在多个数据集上实现优于iTransformer等基线的预测性能,且方法具有鲁棒性。
中文摘要 AI 辅助
预测长期流感样疾病(ILI)对公共卫生准备工作至关重要。公开可用的监测数据集通常将数值流行病学信号与文本信息配对,这些文本信息存在噪声大、结构松散、仅与近期趋势间接相关且往往滞后于数值信号等问题,因此将二者融合需要精心设计。我们提出双流注意力(Dual-Stream Attention, DSA),这是一种多模态深度学习框架,可通过让数值流和文本流相互条件化,基于36周的多模态历史数据预测12周后的ILI活动。使用Time-MMD健康领域数据集,DSA分别基于Transformer的数值编码器和领域适配的标题编码器对两种模态进行编码,随后通过双向跨模态注意力(Cross-Modal Attention, CMA)机制将二者耦合:文本(新闻标题)对数值信号的解释进行条件化,反之亦然。CMA的输出随后传递给因果时间模型进行预测。在10个随机种子上评估,DSA的中位数测试均方误差(MSE)为0.416,而iTransformer、TaTS和GPT4MTS的中位数测试MSE分别为0.668、0.607和0.851,对应平均误差分别降低54.95%、37.29%和67.23%,配对Cohen's d分别为0.555、0.337和0.345,且在100%的自助抽样抽取中排名第一。它的最差窗口误差也显著低于所有基线模型。在外部地理数据集上,DSA再次在9个评估基线中排名第一。消融实验表明,其优势不依赖于文本编码器的选择或语言模型的微调,且双向注意力优于任一单向注意力。最后,基于扰动的忠实性分析显示,学习到的CMA在针对性掩码下具有功能信息,在文本到数值方向上效果更强。
英文摘要
Forecasting long-range influenza-like illness (ILI) matters for public health readiness. Publicly available surveillance datasets typically pair numeric epidemiological signals with textual information that is noisy, loosely structured, only indirectly related to near-term trends, and often lagged relative to the numeric signal. Fusing the two therefore requires careful design. We propose Dual-Stream Attention (DSA), a multimodal deep learning framework that forecasts 12-week-ahead ILI activity from a 36-week multimodal history by letting the numerical and textual streams condition each other. Using the Time-MMD health-domain dataset, DSA separately encodes the two modalities with a Transformer-based numerical encoder and a domain-adapted headline encoder, then couples them through a bidirectional Cross-Modal Attention (CMA) mechanism: the text (news headlines) conditions the interpretation of the numeric signal and vice versa. The CMA output then passes to a causal temporal model for forecasting. Evaluated across ten random seeds, DSA achieves a median test MSE of 0.416, versus 0.668, 0.607, and 0.851 for iTransformer, TaTS, and GPT4MTS, corresponding to mean-error reductions of 54.95%, 37.29%, and 67.23%, with paired Cohen's d of 0.555, 0.337, and 0.345, respectively, and ranks first in 100% of bootstrap draws. It also has substantially lower worst-window error than all baselines. On an external-geography dataset, DSA again ranks first among nine evaluated baselines. Ablations show the advantage does not depend on text-encoder choice or language-model fine-tuning, and that bidirectional attention outperforms either direction alone. Finally, perturbation-based faithfulness analysis shows the learned CMA is functionally informative under targeted masking, with a stronger effect in the text-to-numerical direction.
发表机构
- Institute for Advanced Studies in Basic Sciences (IASBS)(基础科学高级研究所)
机构由 AI 辅助整理,请以论文原文为准。