READ-Bench:面向时间序列诊断的历史实例检索基准
READ-Bench: Benchmarking Historical Instance Retrieval for Time-Series Diagnosis
浏览论文内容
中文总结 AI 辅助
提出READ-Bench基准,评估12个数据集上的时间序列历史案例检索,发现高斯过程重排序器利用少量监督显著提升性能,融合方法使NDCG@10提升0.16。
中文摘要 AI 辅助
时间序列诊断系统很少依赖检索相关的历史案例,即便进行检索,也仅通过下游预测间接评估检索效果。我们提出READ-Bench,一个涵盖12个诊断数据集的历史案例检索基准,聚焦多变量时间序列,以共享故障或事件类型而非信号形状来定义相关性,因此同一故障的视觉上不同的轨迹被视为相关,而不同故障的相似轨迹则不然。将检索视为基础检索器后接重排序器,我们在统一协议下评估经典距离、符号检索器、自监督和基础模型嵌入器及其融合,以及标签感知和语言模型重排序器,该协议变化监督、污染和语料库规模并进行显著性检验。在通用通道独立接口下,预训练表示在仅搜索任务上相比强经典和符号基线无统计学显著优势。决定性因素在于重排序阶段少量已解决案例的监督,即高斯过程重排序器在嵌入空间中传播少量邻居标签,其效果远超更复杂的表示或语言模型推理,且在污染和全语料库规模下保持稳定。基于这些发现,我们通过倒数排名融合将正常残差评分嵌入器与动态时间规整分支融合,再用高斯过程重排序器重排序,在所有12个数据集上相比其自身搜索阶段提升NDCG@10,重排序带来+0.11提升,相比最强单一基础检索器提升+0.16。
英文摘要
Time-series diagnostic systems rarely rely on retrieving relevant historical cases, and when they do, retrieval is evaluated only indirectly through downstream prediction. We introduce READ-Bench, a benchmark for historical-case retrieval across 12 diagnostic datasets, centered on multivariate time series, that defines relevance by shared fault or event type rather than signal shape, so visually different traces of the same fault count as relevant while similar-looking traces of different faults do not. Treating retrieval as a base retriever followed by a reranker, we evaluate classical distances, symbolic retrievers, self-supervised and foundation-model embedders, and their fusion, plus label-aware and language-model rerankers, under one protocol that varies supervision, pollution, and corpus scale with significance testing. Under a common channel-independent interface, pretrained representations offer no statistically detectable advantage over strong classical and symbolic baselines for search alone. The decisive factor is a small amount of resolved-case supervision at reranking, namely a Gaussian-process reranker that propagates a few neighbor labels in embedding space, which helps far more than more sophisticated representations or language-model reasoning and holds under pollution and at full corpus scale. Guided by these findings, we fuse a normal-residual-scored embedder with a dynamic time warping leg via reciprocal-rank fusion, then rerank with the Gaussian-process reranker, improving NDCG@10 over its own search stage on all 12 datasets, by +0.11 from reranking and +0.16 over the strongest single base retriever.
发表机构
- C3 AI(C3 AI公司)
- The Ohio State University(俄亥俄州立大学)
机构由 AI 辅助整理,请以论文原文为准。