arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33685cs.LG

T-MoXAI:面向时间多模态数据的层级可解释性框架

T-MoXAI: A Hierarchical Explainability Framework for Temporal Multimodal Data

Ali Inha, Mo Vali, Saaliha Vali, Pietro Liò, Meen-Yau Thum

首次发表
浏览论文内容

中文总结 AI 辅助

T-MoXAI提出层级可解释框架,通过时间Shapley值、注意力分析和梯度归因,在秒级内解释时间多模态数据预测,并在IVF结果预测和小麦产量预测中验证了有效性。

中文摘要 AI 辅助

面向时间多模态数据的人工智能(AI)模型在医疗和农业领域具有应用潜力,但其不透明性可能限制信任与采用。我们提出T-MoXAI(时间多模态可解释人工智能),这是一个层级框架,用于解释:(1)时间点何时影响预测,采用时间Shapley值;(2)在那些时刻哪些模态有贡献,采用注意力分析;(3)哪些特征或图像区域驱动决策,采用基于梯度的归因方法。基于Transformer的架构处理不规则时间序列和异构数据,在不到一秒内生成全部三个解释层级,以支持交互式决策。我们在两个真实世界任务上评估该框架:从超声序列和临床测量预测IVF治疗结果(尽管存在显著类别不平衡,AUC为0.660),以及从时间RGB图像和表型性状预测小麦产量(在大量环境变异性中R²为0.265)。消融研究表明,时间建模在两个领域均至关重要:移除时间建模会将性能降低至相当于随机猜测的水平。时间ROAR实验提供证据表明,解释反映了模型的推理过程。凭借统一的、领域无关的架构和开源实现,T-MoXAI为时间多模态XAI提供了基线,解决了该领域的碎片化问题,并支持那些理解决策与预测准确性同等重要的应用。

英文摘要

Artificial Intelligence (AI) models for temporal multimodal data have potential in healthcare and agriculture, but their opacity can limit trust and adoption. We introduce T-MoXAI (Temporal Multimodal eXplainable AI), a hierarchical framework explaining (1) when timepoints influence predictions, using temporal Shapley values; (2) which modalities contribute at those moments, using attention analysis; and (3) what features or image regions drive decisions, using gradient based attribution. A transformer based architecture handles irregular temporal sequences and heterogeneous data, generating all three explanation levels in under one second for interactive decision support. We evaluate the framework on two real world tasks: predicting IVF treatment outcomes from ultrasound sequences and clinical measurements (AUC 0.660 despite significant class imbalance), and forecasting wheat yield from temporal RGB imagery and phenotypic traits ($R^2$ 0.265 amid substantial environmental variability). Ablation studies indicate that temporal modelling is critical in both domains: removing it reduces performance to the equivalent of random guessing. Temporal ROAR experiments provide evidence that the explanations reflect the model's reasoning process. With a unified, domain agnostic architecture and open source implementation, T-MoXAI provides a baseline for temporal multimodal XAI, addressing fragmentation in the field and supporting applications where understanding decisions is as important as predictive accuracy.

发表机构

  • University of Cambridge(剑桥大学)
  • Imperial College Healthcare NHS Trust(帝国理工学院医疗NHS信托基金)
  • HCA UK(英国HCA医疗集团)

机构由 AI 辅助整理,请以论文原文为准。

↑