arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ShapeLex:解耦局部形状符号化与全局尺度建模的文本控制时间序列生成

ShapeLex: Decoupling Local Shape Symbolization and Global Scale Modeling for Text-Controlled Time Series Generation

Subo Wei, Jianqi Gao, Mingyan Fan, Shaorong Xie, Xinzhi Wang, Yongpeng Dong

arXiv 2609.24003首次发表:更新:

发表机构

Shanghai University; Shanghai Institute of Applied Physics, Chinese Academy of Sciences(上海大学; 中国科学院上海应用物理研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ShapeLex通过将局部形状离散符号化与全局尺度连续建模解耦,实现文本控制的时间序列生成,在十二个数据集上优于现有方法,并自动合成监督信号提升可扩展性。

AI 中文摘要

文本控制的时间序列生成旨在合成遵循自然语言描述且忠实于真实数据分布的序列。现有范式通常在单一连续潜空间中耦合语义理解与序列建模,缺乏显式的局部语义锚点,也未将全局连续属性与局部离散形状分离。因此,关键的局部结构可能被平滑、遗漏或错位。我们提出形状词典(ShapeLex),将文本到序列的生成解耦为局部形状的离散符号化与全局属性的连续建模。ShapeLex首先从训练数据中归纳出可复用的离散形状单元词汇表,如上升、尖峰和急剧下降,形成可解释的符号空间。随后,一个自回归生成器根据文本描述选择形状,调整位置和持续时间等属性,并按时间顺序将它们组合成形状骨架。最后,一个混合密度尺度头对整体水平和波动性进行建模与采样,以恢复真实的全局尺度。在十二个公共数据集、真实用户编写的文本以及下游预测任务上的实验表明,ShapeLex生成的序列比现有方法更匹配真实数据分布。此外,配对监督信号从学习到的词汇表中自动合成,避免了随数据集规模增长而增加的标注成本,提高了可扩展性。

英文摘要

Text-controlled time series generation aims to synthesize sequences that follow natural-language descriptions while remaining faithful to real data distributions. Existing paradigms often couple semantic understanding and sequence modeling in a single continuous latent space, lacking explicit local semantic anchors and separation between global continuous attributes and local discrete shapes. As a result, key local structures may be smoothed, missed, or misplaced. We propose Shape Lexicon (ShapeLex), which decouples text-to-sequence generation into discrete symbolization of local shapes and continuous modeling of global attributes. ShapeLex first induces a reusable vocabulary of discrete shape units, such as rises, spikes, and sharp drops, from training data, forming an interpretable symbolic space. An autoregressive generator then selects shapes according to the textual description, adjusts attributes such as position and duration, and composes them in temporal order into a shape skeleton. Finally, a mixture-density scale head models and samples the overall level and volatility to restore realistic global scale. Experiments on twelve public datasets, real user-written text, and downstream forecasting tasks show that ShapeLex generates series that better match real data distributions than existing methods. In addition, paired supervision is automatically synthesized from the learned vocabulary, avoiding annotation costs that grow with dataset size and improving scalability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑