发表机构
Johns Hopkins University(约翰斯·霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出 ChronoState 基准,证实冻结骨干语言模型可在直接监督下将隐式流逝时间与符号状态组合,但其泛化性有限,且性能不及提示注入时间戳基线。
AI 中文摘要
语言模型系统中的时间决策通常同时依赖于符号任务状态和流逝的挂钟时间,例如缓存过期、作业完成、配额重置、截止日期或过时会话。我们研究流逝时间是否可以作为非 token 的系统侧标量提供,并由冻结骨干语言模型与可见符号状态组合。我们引入 ChronoState,这是一个组合式时间-状态基准,其中符号状态出现在提示中,流逝秒数 τ 通过隐式计时注入通道提供,模型选择强制选择的时间动作。此处的“隐式”指对用户可见的 token 序列隐藏,而非对模型计算隐藏。使用 Qwen2.5-3B-Instruct 作为冻结的 bf16 骨干,搭配 31 维正弦加对数时间编码、门控 FiLM 残差调制以及秩为 8 的 LoRA 动作表面,隐式时间 CI 达到 0.9305 ± 0.0134 的准确率和 0.9410 ± 0.0103 的平衡准确率。无时间和打乱时间的对照组分别降至 0.5511 ± 0.0042 和 0.3323 ± 0.0097,高打乱时间错误状态一致性支持训练分布内对注入标量的因果依赖。在保留的模板、持续时间和多约束组合上的泛化性仍较强,但保留的配额家族迁移准确率较弱,为 0.5065 ± 0.0559,而公平的提示加 LoRA 时间戳基线达到 0.9893 ± 0.0052。因此,ChronoState 支持一个狭义结论:隐式流逝时间可在直接监督下与符号任务状态组合,但不能确立自主时间跟踪、广泛的未见家族抽象或优于提示注入时间戳的性能。
英文摘要
Temporal decisions in language-model systems often depend on both symbolic task state and elapsed wall-clock time, such as cache expiration, job completion, quota resets, deadlines, or stale sessions. We study whether elapsed time can be supplied as a non-token, system-side scalar and composed with visible symbolic state by a frozen-backbone language model. We introduce ChronoState, a compositional temporal-state benchmark in which symbolic state appears in the prompt, elapsed seconds tau are supplied through a hidden chronometric-injection channel, and the model selects a forced-choice temporal action. Here, "hidden" means hidden from the user-visible token sequence, not from model computation. Using Qwen2.5-3B-Instruct as a frozen bf16 backbone with a 31-dimensional sinusoidal-plus-log time encoding, gated FiLM residual modulation, and a rank-8 LoRA action surface, hidden-time CI reaches 0.9305 +/- 0.0134 accuracy and 0.9410 +/- 0.0103 balanced accuracy. No-time and shuffled-time controls fall to 0.5511 +/- 0.0042 and 0.3323 +/- 0.0097, respectively, with high shuffled-time wrong-state consistency supporting causal dependence on the injected scalar within the trained distribution. Generalization remains strong for held-out templates, durations, and multi-constraint compositions, but held-out quota-family transfer is weak at 0.5065 +/- 0.0559, while a fair prompt+LoRA timestamp baseline reaches 0.9893 +/- 0.0052. Thus, ChronoState supports a narrow conclusion: hidden elapsed time can be composed with symbolic task state under direct supervision, but does not establish autonomous time tracking, broad unseen-family abstraction, or superiority over prompt-injected timestamps.
CommentsSubmitted to SPIE Defense and Commercial Sensing 2027