GALA:面向文本到时间序列合成的生成感知跨模态对齐
GALA: Generation-Aware Cross-Modal Alignment for Text-to-Time-Series Synthesis
AI总结:
本研究针对文本到时间序列合成中条件表示与信号模态不匹配的问题,提出GALA两阶段跨模态对齐方法,在TSFragment-600K数据集上实现SOTA,打破了生成器内部文本编码器的保真度与贴合度权衡。
AI中文摘要:
从自然语言合成时间序列正成为可控时间序列生成最具表现力的形式。然而,现有的文本条件生成器要么采用现成文本编码器的冻结字幕嵌入,要么对编码器进行端到端适配,仅让去噪损失作为副作用塑造嵌入。无论哪种情况,条件表示都未与信号模态进行刻意匹配,导致其不适用于指导生成。我们通过引入GALA(Generation-Aware cross-modaL Alignment,面向文本条件时间序列生成的生成感知跨模态对齐)解决该问题。GALA是一种两阶段方法:首先将预训练文本编码器与时间序列基础模型通过辅助生成损失适配两个编码器以生成的方式,对比耦合到共享嵌入空间;随后冻结得到的字幕嵌入以驱动流匹配生成器。在涵盖四个领域和三种片段长度的TSFragment-600K数据集上,GALA达到新的SOTA,在36个指标列中排名第一,在长度24/48/96时的平均排名为1.08/1.08/1.42,而最强基线的平均排名为1.92/2.00/1.75。我们进一步发现,生成器内部文本编码器会在保真度与字幕贴合度之间产生权衡,而基于对齐嵌入的条件生成可打破该权衡:FID、CTTP和JFTSD均同步提升。对辅助损失进行 ablation 会同时降低FID、CTTP和JFTSD,这表明生成项是对齐的必要组成部分,而非附加项。
英文摘要:
Synthesizing time series from natural language is emerging as the most expressive form of controllable time series generation. However, existing text-conditioned generators either take caption embeddings frozen from off-the-shelf text encoders, or adapt the encoder end-to-end, letting the denoising loss shape the embeddings only as a by-product. In either case, the conditioning representation is never deliberately matched to the signal modality, leaving it ill-suited to guide generation. We address this by introducing GALA: Generation-Aware cross-modaL Alignment for text conditional time series generation. GALA is a two-stage approach that first contrastively couples a pretrained text encoder with a time-series foundation model into a shared embedding space with both encoders adapted to generation by an auxiliary generative loss, and then freezes the resulting caption embedding to drive a flow-matching generator. On TSFragment-600K, spanning four domains and three fragment lengths, GALA sets a new state of the art, ranking first in 30 of 36 metric columns and reaching an average rank of 1.08/1.08/1.42 at lengths 24/48/96 against 1.92/2.00/1.75 for the strongest baseline. We further find that generator-internal text encoders force a trade-off between fidelity and caption adherence, whereas conditioning on the aligned embedding breaks it: FID, CTTP, and JFTSD all improve at once. Ablating the auxiliary loss degrades FID, CTTP and JFTSD together, it indicates the generative term is a necessary component of the alignment rather than an add-on.