arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自改进测试时智能:推理阶段的反馈驱动适配、学习与扩展综述

A Survey on Self-Improving Test-Time Intelligence: Feedback-Driven Adapting, Learning, and Scaling at Inference

Shuaicheng Niu, Guohao Chen, Yaofo Chen, Zhiquan Wen, Jinwu Hu, Zeshuai Deng, Deyu Chen, Shuhai Zhang, Renjie Chen, Zihao Lian, Shoukai Xu, Gang Dai, Yunbei Zhang, Wei Luo, Yifan Zhang, Mingkui Tan, Cheng Deng

arXiv 2609.01679首次发表:更新:

发表机构

South China University of Technology; Nanyang Technological University; Tulane University; Pazhou Laboratory; National University of Singapore; Hohai University(华南理工大学; 南洋理工大学; 杜兰大学; 琶洲实验室; 新加坡国立大学; 河海大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本综述提出反馈驱动的测试时智能(TTI)作为统一视角,关联测试时适配、学习与扩展,梳理相关方法、应用与挑战,为推理阶段自改进AI研究提供基础与路线图。

AI 中文摘要

AI系统在部署过程中改进自身行为的能力正变得愈发重要。随着推理不再局限于固定训练模型的静态执行,越来越多的研究致力于探索模型如何通过利用测试时信息和额外计算来动态优化自身行为。这些进展主要沿两个方向发展:一是利用测试时信号修改模型状态的方法,二是通过采样、工具使用等额外推理资源改进预测的方法。然而,这些方向常被不同领域以不同术语分开研究,导致其联系难以被察觉。本综述将反馈驱动的测试时智能(Test-Time Intelligence, TTI)作为统一视角,用于理解这类部署时的改进。我们用这一视角关联测试时适配、测试时学习与测试时扩展,既凸显它们的区别,也强调其在混合系统中日益增加的重叠。该统一框架有助于连接此前零散的思路,为研究推理时自改进提供更清晰的概念基础。我们回顾了视觉、语言、多模态学习、生成模型、机器人学和医疗保健领域的主要方法范式、代表性应用及开放挑战,旨在为测试时自改进AI系统的研究提供连贯基础与研究路线图。

英文摘要

The ability of AI systems to improve their behavior during deployment is becoming increasingly important. As inference moves beyond the static execution of a fixed trained model, a growing body of work studies how models can refine their behavior on the fly by exploiting test-time information and additional computation. These developments have largely evolved along two directions: methods that modify the model's state using test-time signals, and methods that improve predictions through extra inference-time resources such as more sampling and tool use. However, these directions are often studied in separate communities with different terminology, making their connections harder to see. In this survey, we present feedback-driven Test-Time Intelligence (TTI) as a unified perspective for understanding such deployment-time improvement. We use this view to relate test-time adaptation, test-time learning, and test-time scaling, highlighting both their distinctions and their growing overlap in hybrid systems. This unified framework helps connect previously fragmented ideas and provides a clearer conceptual foundation for studying inference-time self-improvement. We review major methodological paradigms, representative applications, and open challenges across vision, language, multimodal learning, generative models, robotics, and healthcare. Our goal is to provide a coherent foundation and research roadmap for the study of self-improving AI systems at test time.

Commentsaccepted by Machine Intelligence Research

DOI:10.1007/s11633-026-1694-1

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑