WaveTLM:通过任务编译实现可靠的时间序列语言建模
WaveTLM: Reliable Time-Series Language Modeling through Task Compilation
浏览论文内容
中文总结 AI 辅助
WaveTLM通过任务编译将自然语言请求转换为类型化任务状态,并由任务原生执行器生成可靠输出,在ExecTS-QA上实现99.40%契约覆盖率,显著优于基线,确保时间序列任务输出的可靠性。
中文摘要 AI 辅助
时间序列语言模型为跨时间任务提供了共享的自然语言接口,但看似合理的文本并不能保证可靠的任务输出。响应可能看起来合理,同时却对所需对象产生幻觉:数值序列可能违反形状、尺度、通道顺序或时间对齐,文本决策可能超出合法标签空间。我们形式化了可靠的时间序列语言建模,将任务对象可靠性与预测质量区分开来。我们引入了ExecTS-QA,一个基于契约的基准,涵盖预测、插补、分类、异常检测和波形分析。我们进一步提出了WaveTLM,一个统一的编译器-执行器模型,其任务编译器将用户请求、可见参数和基于波形的证据转换为类型化任务状态,而任务原生执行器则构建数值张量、合法决策或结构化记录。在ExecTS-QA上,单个WaveTLM检查点实现了99.40%的契约有效覆盖率,而最强评估的字符串优先基线为37.83%,同时在所有五个任务族中保持了平衡的预测性能。在SciTS、TSQA、IRTS-ToolBench和ARFBench上的评估提供了额外的迁移证据。代码、构建脚本和ExecTS-QA数据集将在发表后公开。这些结果表明,任务编译可以将合理的语言生成转换为可靠的时间序列输出。
英文摘要
Time-series language models provide a shared natural-language interface across temporal tasks, but plausible text does not guarantee reliable task outputs. Responses may appear reasonable while hallucinating the required object: numerical sequences can violate shape, scale, channel order, or temporal alignment, and textual decisions can fall outside the legal label space. We formulate reliable time-series language modeling, separating task-object reliability from predictive quality. We introduce ExecTS-QA, a contract-grounded benchmark spanning forecasting, imputation, classification, anomaly detection, and waveform analysis. We further propose WaveTLM, a unified compiler-executor model whose task compiler transforms user requests, visible arguments, and wave-grounded evidence into typed task states, while task-native executors construct numerical tensors, legal decisions, or structured records. On ExecTS-QA, a single WaveTLM checkpoint achieves 99.40% contract-valid coverage, compared with 37.83% for the strongest evaluated string-first baseline, while retaining balanced predictive performance across all five task families. Evaluations on SciTS, TSQA, IRTS-ToolBench, and ARFBench provide additional evidence of transfer. The code, construction scripts, and ExecTS-QA dataset will be publicly released upon publication. These results show that task compilation can convert plausible language generation into reliable time-series outputs.
发表机构
- Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。