发表机构
School of Computer Science, University College Dublin(都柏林大学学院计算机科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对时间序列分类中可扩展性难题,提出drXAI方法,利用XAI归因方法进行数据约简,通过GPU加速分类器生成特征重要性分数并选择显著特征,在多数据集上评估,实现数据约简同时保持准确率,助力资源密集型模型处理更大数据集。
AI 中文摘要
可解释人工智能(XAI)在时间序列方面算法有显著发展,但其为下游任务带来可衡量性能提升的效用仍未充分探索。本文引入drXAI弥合这一差距,它将XAI归因方法用于时间序列分类中的有效数据约简。现代时间序列分类的核心挑战是可扩展性,drXAI利用快速的GPU加速分类器生成局部归因,聚合为全局特征重要性分数并采用自动肘部切割启发式选择最显著特征。在合成和真实世界数据集上评估,drXAI在合成基准上成功恢复传统基线失败的真实特征,在真实数据上实现80%至90%的数据约简且保持分类准确率,还使ConvTran等资源密集型模型能扩展到因内存限制之前无法处理的数据集,表明XAI不仅用于可解释性,还可作为时间序列分析中特征选择和可扩展性的强大工具。
英文摘要
Explainable AI (XAI) for time series has seen significant algorithmic growth, but its utility in providing measurable performance gains for downstream tasks remains under-explored. This paper bridges this gap by introducing drXAI, a novel methodology that repurposes XAI attribution methods for effective data reduction in Time Series Classification (TSC). The core challenge in modern TSC is scalability; state-of-the-art models, such as Transformers, exhibit quadratic complexity relative to sequence length and linear complexity relative to the number of channels. This renders them computationally prohibitive for massive datasets. drXAI addresses this by using a fast, GPU-accelerated classifier (Hydra) to generate local attributions. We aggregate these into global feature importance scores and employ an automated elbow-cut heuristic to select the most salient features without requiring manual thresholds. We evaluate our approach on both synthetic and real-world univariate and multivariate datasets. On synthetic benchmarks, drXAI successfully recovers ground-truth features where traditional baselines fail. On real-world data, drXAI achieves between 80% and 90% data reduction while maintaining classification accuracy comparable to models trained on the full dataset. Most importantly, we show that drXAI allows resource-intensive models like ConvTran to scale to datasets that were previously inaccessible due to memory constraints. Our results show the benefits of using XAI not just for interpretability, but as a robust tool for feature selection and scalability in time series analysis. All our code and data are openly available.
CommentsAccepted for AALTD workshop at ECML-PKDD 2026