TW3Cast:轻量微调基础模型的冻结路由器,用于GIFT-Eval时间序列预测,完全在训练集上选择
TW3Cast: A Frozen Router of Lightly Fine-Tuned Foundation Models for Time-Series Forecasting on GIFT-Eval, Selected Entirely on the Training Split
- TW3 Partners
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
TW3Cast通过冻结路由表选择轻量微调的基础模型,在GIFT-Eval基准上以平均MASE排名第3,无需智能体或语言模型,完全基于训练集回测决策。
AI中文摘要:
TW3Cast是一个时间序列预测系统,截至2026年9月14日,在GIFT-Eval基准的130个参赛条目中,按平均MASE排名达到第3位。其上方两个条目属于排行榜的智能体类别,即使用智能体或语言模型来推理、生成或选择预测的多步骤系统。TW3Cast不运行任何智能体,也不使用语言模型。其选择是一个在训练集上一次性计算后冻结的表格,其专家模型是在这些训练集上轻量微调的公开基础模型。对于97种数据集、频率和预测长度配置,该表格提供四种模式之一:专家模型,即对Chronos-2、TiRex或Toto进行LoRA或全量微调,其训练数据通过显式规则进行清洗和丰富;包含专家模型的分位数混合;基础模型的混合;或在从训练集划分出的回测上进行的选拔锦标赛。表格中的每个决策都在该回测上做出。一旦专家模型在锦标赛中胜出即被接纳,因此一个候选模型仅需几兆字节和几分钟的GPU时间,失败的候选不会改变任何内容。三个保护机制防止选择过程自身偏差:双重准确性和校准标准、对训练期间见过序列的候选模型的不对称边际,以及保守的逐窗口门控。选择规则本身在时间元回测中选定。最佳基础模型单独服务的平均MASE排名为33.8,锦标赛在所有配置上服务的排名为38.0,而完整路由器的排名为19.4。路由表、专家索引、固定的基础模型版本、提交的分数文件以及公开分数的带日期快照均已发布,本文中的每个排行榜数字均可通过一个脚本从这些数据重新生成。
英文摘要:
TW3Cast is a time-series forecasting system that reaches position 3 of 130 entries on the GIFT-Eval benchmark by mean MASE rank, as of 2026-09-14. The two entries above it belong to the leaderboard's agentic category, multi-step systems that use agents or language models to reason about, generate or select forecasts. TW3Cast runs no agent and no language model. Its selection is a table computed once on the training split and then frozen, and its experts are public foundation models lightly fine-tuned on those training splits. For each of the 97 dataset, frequency and horizon configurations, the table serves one of four modes: a specialist, which is a LoRA or full fine-tune of Chronos-2, TiRex or Toto whose training data was cleaned and enriched by explicit rules; a quantile blend that contains a specialist; a blend of base models; or a selection tournament played on a backtest carved from the training split. Every decision in the table was taken on that backtest. A specialist is admitted the moment it beats the tournament there, so a candidate costs a few megabytes and minutes of GPU time, and a failed candidate changes nothing. Three guarded mechanisms protect the selection from its own biases: a dual accuracy and calibration criterion, an asymmetric margin against candidates that saw the series during training, and conservative per-window gates. The selection rules themselves were chosen inside a temporal meta-backtest. The best base model served alone reaches a mean MASE rank of 33.8, the tournament served on every configuration reaches 38.0, and the full router reaches 19.4. The routing table, the expert index, the pinned base-model revisions, the submitted score file and the dated snapshot of the public scores are released, and every leaderboard number in this paper regenerates from them by one script.