发表机构
The University of Tokyo; Harbin Institute of Technology; RWTH Aachen University; HKUST(东京大学; 哈尔滨工业大学; 亚琛工业大学; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究上下文学习器何时应扩展假设空间,提出将预测失败转化为结构证据并基于价值边界决策的框架,通过结构修正环境验证,发现模型决策与预测信号脱节。
AI 中文摘要
学习系统在熟悉的模型族内能快速适应。更困难的步骤发生在更早阶段:从可能是噪声、异常、族内变化或族外结构的观测中,决定开启一个更丰富的模型族是否值得其代价。我们将此视为一个代价高昂的序列决策:预测失败必须转化为结构证据,证据转化为扩展的价值,价值转化为行动。结构修正环境(Structural Revision Environment)从每个来源产生匹配的失败,独立于证据改变扩展价格和剩余时间范围,并允许精确的贝叶斯计算和一次性决策的精确规范解。其解表明,修正是价值边界而非证据阈值:同一历史在不同价格、时间范围和查询下有不同的最优行动,局部修复与扩展之间的边界由规则预测的输入和修复无法覆盖的输入设定,而对更丰富模型族的信念在决策之前很久就已跨越。在该环境中训练的Transformer仅从效用中复现了这一边界。三个训练后血统的语言模型在其预测中携带失败敏感信号,但该信号未反映在其修正决策中,并且给定扩展的收益,它们读取该信号却不将其与价格和时间范围权衡。三个允许推理的模型以参考比例权衡所述收益,但仍未将历史转化为对扩展所能带来价值的估计。对元训练学习器进行受控后训练会移动先验和预测的锐度,但两者均不移动标准。
英文摘要
Learning systems adapt quickly inside a familiar family of models. The harder step comes earlier: deciding, from observations that could be noise, an exception, a change within the family or structure outside it, whether opening a richer family is worth its cost. We treat this as a costly sequential decision: prediction failure must be turned into structural evidence, evidence into a value of expansion, and value into action. The Structural Revision Environment produces matched failures from each source, varies the price of expansion and the remaining horizon independently of the evidence, and admits exact Bayesian calculations and an exact normative solution of the one-shot decision. Its solution shows that revision is a value boundary and not an evidence threshold: one history has different optimal actions under different prices, horizons and announced queries, the boundary between local repair and expansion is set by the inputs a rule predicts and a repair cannot cover, and belief in the richer family crosses long before the decision does. Transformers trained in the environment reproduce this boundary from utility alone. Language models of three post-training lineages carry a failure-sensitive signal in their predictions that is not reflected in their revision decisions, and given the gain of expanding they read it without weighing it against price and horizon. Three models allowed to reason weigh the stated gain in the reference's proportions and still do not turn the history into an estimate of what expansion would buy. Controlled post-training of the meta-trained learners moves the prior and the sharpness of predictions, and neither moves the criterion.
Comments43 pages, 11 figures