发表机构
University of Illinois Chicago; Texas A&M University; Massachusetts Institute of Technology; Tsinghua University(伊利诺伊大学芝加哥分校; 德克萨斯农工大学; 麻省理工学院; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出CRAFTER智能体,挖掘黑盒冻结预测器的残差以生成纠正特征,在6个数据集和6个主干上优于现有系统,使改进翻倍、误差降27%,可归因特征来源。
AI 中文摘要
冻结的预训练预测器常以结构化、可重复的方式失效,通过微调修复成本高昂。本文研究纠正特征发现:挖掘冻结预测器残差的可解释特征,以驱动轻量的事后纠正器。现有自动化特征工程对数据生成过程建模,而纠正特征对模型失效过程建模。本文提出CRAFTER(基于特征的时序探索与推理的纠正残差智能体),其保留主干冻结状态,通过两个互补生成器挖掘残差:对原始输入通道的组合搜索,以及提出命名特征组合、二元标志和短可执行代码的大语言模型(LLM)。无论候选来源如何,单个基于验证集的门控对其接受或否决,且由验证集选择的纠正器应用已接受特征或保持预测不变。该与来源无关的流程还可让现有特征工程系统在相同条件下被评估,使CRAFTER成为将预测改进归因于特征来源的工具。在6个公开数据集和6个冻结主干上,CRAFTER在所有特征预算下均优于每个专用特征工程系统,大致使仅纠正器实现的改进翻倍,且将最弱主干的误差降低多达27%。这些增益在不同LLM后端间具有鲁棒性,即使应用于微调后的主干之上仍能保持。
英文摘要
Frozen pretrained forecasters often fail in structured, recurring ways that are costly to repair through fine-tuning. We study corrective feature discovery: mining interpretable features of a frozen forecaster's residual to drive a lightweight post-hoc corrector. Prior automated feature engineering models the data-generating process; corrective features instead model the model-failure process. We present CRAFTER (Corrective Residual Agent with Feature-based Temporal Exploration and Reasoning), which keeps the backbone frozen and mines its residual with two complementary generators: a compositional search over the raw input channels, and a large language model (LLM) that proposes named feature combinations, binary flags, and short executable code. A single validation-grounded gate accepts or rejects every candidate regardless of its origin, and a validation-selected corrector applies the accepted features or leaves the forecast unchanged. This source-agnostic pipeline also allows prior feature-engineering systems to be evaluated under identical conditions, making CRAFTER an instrument for attributing forecast improvements to the feature source alone. Across six public datasets and six frozen backbones, CRAFTER surpasses every dedicated feature-engineering system at every feature budget, roughly doubling the improvement achieved by the corrector alone and reducing the error of the weakest backbones by up to 27%. These gains are robust across different LLM backends and persist even when applied on top of fine-tuned backbones.
Comments23 pages, 14 tables, 4 figures