arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05238cs.LG

感知与描述解耦:多变量时间序列与语言间的计算基础表示对齐

Decoupling Perception from Description: Computation-Grounded Representation Alignment between Multivariate Time Series and Language

Xinran Feng, Yi Xie, Chao Zhang, Ruikun Li, Wanyun Ling, Ziyue Li, Chenxi Liu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对多模态时间序列语言对齐的三难困境,提出CGTime模型,通过计算处理感知、LLM处理描述,在多变量理解任务上优于更大的通用模型。

中文摘要 AI 辅助

训练将时间序列与语言对齐的多模态模型会陷入自监督陷阱。常规做法是让大型语言模型(LLM)读取序列并生成描述,因此标签质量受限于模型本应学习的感知能力,数据能传授的知识永远不会超过标注者已掌握的内容。第二个差距加剧了这一问题:大多数数据集使用单一变量,但关键模式(跨通道相关性、领先滞后结构、共现异常)仅在多变量场景中出现,而标注用LLM的局限性在此处最为凸显。这两个问题形成了三难困境:现有方法可靠、现实或可扩展,但无法同时满足三者。我们通过将感知与描述解耦来解决该问题:确定性代码从真实开源多变量序列中计算一组统计量,LLM对这些预计算事实进行文本表述。LLM不擅长的感知由计算处理,而LLM负责表达。由此得到我们的40亿参数计算基础时间序列语言模型CGTime。CGTime在多变量理解任务上表现优于大得多的通用模型:在我们的保留基准上取得了最佳多变量事实分数(0.283,对比GPT-4o-mini的0.173和GPT-5.4-nano的0.203),该差距在针对所有基线的霍尔姆校正配对显著性检验中仍存在。它还能在生成的描述中更准确地陈述可验证的数值事实,并覆盖更广泛的统计属性。

英文摘要

Training multimodal models to align time series with language runs into a self-supervision trap. The usual recipe asks an LLM to read a series and write a description, so label quality is capped by the perceptual skill the model is supposed to learn. The data can never teach more than the labeler already knows. A second gap makes this worse: most datasets use a single variable, but the patterns that matter (cross-channel correlation, lead-lag structure, co-occurring anomalies) appear only with several variables, right where the labeling LLM's limits are most exposed. These two problems create a trilemma: existing methods are reliable, realistic, or scalable, but none achieves all three. We resolve this by decoupling perception from description. Deterministic code computes a set of statistics from real, open-source multivariate series; the LLM verbalizes those precomputed facts. Perception, which LLMs do poorly, is handled by computation, while the LLM handles expression. This produces CGTime, our 4B-parameter computation-grounded time-series-language model. CGTime outperforms far larger general-purpose models on multivariate understanding tasks: it attains the best multivariate fact score on our held-out benchmark (0.283 vs. 0.173 for GPT-4o-mini and 0.203 for GPT-5.4-nano), a gap that survives Holm-corrected paired significance tests against every baseline. It also states verifiable numerical facts in generated captions more accurately and covers a broader range of statistical properties.

补充信息

↑