arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ChronoSSM:自回归状态空间模型中用于时间感知表示的训练方法

ChronoSSM: Training for Temporally Aware Representations in Autoregressive State Space Models

Adrien Schoen, Nachiketa Ratnakar Patil, Arjun Bhagoji, Francesco Bronzino

arXiv 2608.10120首次发表:更新:

发表机构

ENS de Lyon; CNRS; UCBL1; Indian Institute of Technology Bombay; Institut universitaire de France(里昂高等师范学院; 法国国家科学研究中心; 里昂第一大学; 印度理工学院孟买分校; 法国大学研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ChronoSSM是一种自回归状态空间模型,通过联合建模事件与时间戳的共享主干,在四个领域的实验中,能生成更具时间信息的表示且不降低内容生成质量。

AI 中文摘要

现代序列模型,从Transformer到状态空间模型(SSM),已在多个领域实现了强大的生成式建模,但这些模型通常被训练为预测事件内容,而将事件发生的时间视为次要问题。在事件与显式时间信息关联的数据挖掘场景中,这种分离会限制时间推理、异常检测以及事件时间顺序的准确重建。一种常见策略是将时间视为辅助信号,使用仅为事件预测学习到的表示训练独立的时间模型。然而,这种两阶段方法隐含假设:为事件预测优化的表示已包含足够的时间结构。我们提出ChronoSSM,一种自回归状态空间模型(SSM),其通过结合 token 和时间生成目标训练的共享主干,联合建模事件和时间戳。我们将时间监督更新主干的联合训练机制,与仅使用冻结事件表示学习时间的两阶段机制进行对比。在涵盖密集和部分时间戳监督的四个领域中,联合训练始终使冻结表示中可恢复的到达间隔信息更多,且整体内容生成质量无系统性下降。我们的结果表明,时间监督可生成更具时间信息的表示,且不会显著降低自回归事件建模的性能。

英文摘要

Modern sequence models, from Transformers to State Space Models, have enabled powerful generative modeling across diverse domains, yet they are typically trained to predict what happens while treating when it happens as a secondary concern. In data-mining settings where events are associated with explicit timing information, this separation can limit temporal reasoning, anomaly detection, and faithful reconstruction of event chronology. A common strategy is to treat timing as an auxiliary signal, training a separate timing model using representations learned solely for event prediction. However, this two-stage approach implicitly assumes that representations optimized for event prediction already contain sufficient temporal structure. We introduce ChronoSSM, an autoregressive State Space Model (SSM) that jointly models events and timestamps with a shared backbone trained using combined token and temporal generation objectives. We compare the joint regime, where temporal supervision updates the backbone, with the two-stage regime, where timing is learned only using the frozen event representations. Across four domains spanning dense and partial timestamp supervision, joint training consistently makes inter-arrival information more recoverable from frozen representations without any systematic degradation in content-generation quality overall. Our results show that temporal supervision can produce more temporally informative representations without materially degrading autoregressive event modeling.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑