大规模标签高效时间序列分类:具有反事实归因的双流OSSE-LSTM
Label-Efficient Time Series Classification at Scale: A Dual-Stream OSSE-LSTM with Counterfactual Attribution
浏览论文内容
中文总结 AI 辅助
针对大规模时间序列分类中标签稀缺问题,提出双流OSSE-LSTM框架,结合全尺度CNN与双向LSTM,并引入反事实积分梯度归因,在19个UCR数据集上以每类少量样本实现最高准确率。
中文摘要 AI 辅助
工业设备、可穿戴设备、电网和临床监测器以巨大规模持续产生时间序列数据,然而标注仍然依赖人工、成本高昂且需要专家参与。因此,大规模时间序列分析中的制约因素不是数据量而是标签量,实践者面临的具体问题是:每个类别需要标注多少样本才能使分类器变得可用?我们直接研究这个问题,在标签空间固定且预先已知、决策规则必须仅由每个类别的K个标注样本构建的设定下进行。我们提出双流OSSE-LSTM,一种情景度量学习框架,将全尺度CNN与压缩-激励重校准配对,用于无需逐数据集内核调优的多尺度模式提取,并与双向LSTM结合以获取全局时间上下文。两个流分别归一化并融合为原型导向的嵌入。由于基于少量标签做出的决策也必须可解释,我们引入反事实积分梯度(C-IG),其归因于目标类与对立类之间的原型间隔而非孤立的分类器logit,并将所得映射重新用作测试时原型细化的软掩码,而无需更新编码器。在19个单变量UCR数据集上,OSSE-LSTM在每个支持大小下均达到最高平均准确率和逐数据集胜场数,且其准确率在该范围内保持在0.36个百分点的带宽内(96.36%-96.72%)。其最弱配置仍超过任何比较基线在任意K下取得的最佳结果(93.99%)。
英文摘要
Time series are produced continuously at enormous scale by industrial equipment, wearables, power grids, and clinical monitors, yet annotation remains manual, expensive, and expert-dependent. The binding constraint in large-scale time series analytics is therefore not data volume but label volume, and the question facing a practitioner is concrete: how many examples per class must be labeled before a classifier becomes usable? We study this question directly, in a regime where the label space is fixed and known in advance and the decision rule must be constructed from only K labeled examples per class. We propose Dual-Stream OSSE-LSTM, an episodic metric-learning framework that pairs an Omni-Scale CNN with Squeeze-and-Excitation recalibration, for multi-scale motif extraction without per-dataset kernel tuning, with a Bidirectional LSTM for global temporal context. The two streams are independently normalized and fused into a prototype-oriented embedding. Because decisions taken from a few labels must also be explainable, we introduce Counterfactual Integrated Gradients (C-IG), which attributes the prototype margin between target and opposing classes rather than an isolated classifier logit, and reuses the resulting maps as soft masks for test-time prototype refinement without updating the encoder. On 19 univariate UCR datasets, OSSE-LSTM attains the highest average accuracy and per-dataset win count at every support size, and its accuracy remains within a 0.36-point band (96.36-96.72%) across that range. Its weakest configuration still exceeding the best result any compared baseline achieves at any K (93.99%).
发表机构
- Loyola University Maryland(洛约拉大学马里兰分校)
- International University, Vietnam(越南国际大学)
- Wuhan University of Technology(武汉理工大学)
机构由 AI 辅助整理,请以论文原文为准。