发表机构
Shandong University; National University of Singapore; Tsinghua University(山东大学; 新加坡国立大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出局部预测充分性原则,通过控制接口处预测缺陷来指导递归压缩训练,在多个任务中匹配或提升宿主模型性能,并更好地保留预测信息。
AI 中文摘要
递归计算反复压缩或重用中间状态,产生了一个简单的张力:在更长的递归路径中必须保持有用的信息,在到达最终预测之前也面临更多丢失的机会。现有的重建或局部预测目标提供了可操作的监督,但不能确保保留的信息对于后续递归计算仍然充分。我们将局部预测充分性与递归预测闭包联系起来:控制各个接口处的局部预测缺陷可以控制根部的最终差异。然后我们将这一原则转化为一个可操作的训练过程。从变分特征化出发,我们推导出有限预测测试和经验预测缺陷,用于衡量压缩过程中保留的预测价值。其预测敏感性定义了参数更新的边际松弛半空间约束,并且仅当预测保留否则将被违反时,我们将宿主优化器提出的更新投影到它们的交集上。在时间图、语言记忆、视觉-语言-动作控制和递归自我改进等任务中,该方法在匹配压缩预算下匹配或改进相应的宿主模型,在更重的递归或记忆需求下获得更大收益,同时更好地保留连续变换中的预测信息。关键的是,相同的任务无关预测保留原则通过宿主兼容干预在所有四个设置中实例化,同时保持端点任务、主干和评估协议固定。这些结果确立了递归接口处的预测保留作为递归压缩的一般训练原则。
英文摘要
Recursive computation repeatedly compresses or reuses intermediate states, creating a simple tension: information that must remain useful across longer recursive paths is also exposed to more opportunities for loss before reaching the final prediction. Existing reconstruction or local-prediction objectives provide tractable supervision, but do not ensure that the retained information remains sufficient for subsequent recursive computation. We identify local predictive sufficiency with recursive predictive closure: controlling local predictive deficiencies at individual interfaces controls the resulting discrepancy at the root. We then turn this principle into a tractable training procedure. Starting from a variational characterization, we derive finite predictive tests and an empirical predictive deficiency that measures predictive value retained across compression. Its predictive sensitivities define margin-relaxed half-space constraints on parameter updates, and we project the host optimizer's proposed update onto their intersection only when predictive preservation would otherwise be violated. Across temporal graphs, language memory, vision-language-action control, and recursive self-improvement, the method matches or improves the corresponding host models under matched compression budgets, with larger gains under heavier recursive or memory demands, while better preserving predictive information across successive transformations. Crucially, the same task-agnostic predictive-preservation principle is instantiated across all four settings through host-compatible interventions while keeping the endpoint task, backbone, and evaluation protocol fixed. These results establish predictive preservation at recursive interfaces as a general training principle for recursive compression.