上下文自适应推理:统一的统计和基础模型视角
Context-Adaptive Inference: A Unified Statistical and Foundation-Model View
浏览论文内容
中文总结 AI 辅助
研究现代预测系统的上下文自适应推理,统一了统计、元学习和基础模型的相关方法,证明特定条件下显式与隐式自适应等价,提出设计原则、评估指标并识别开放问题。
中文摘要 AI 辅助
现代预测系统需根据具体情况调整行为。本文统一了三种通常分开处理的上下文自适应推理传统:统计中的显式自适应(如变系数模型等)、元学习和迁移中的快速特定任务自适应、大基础模型中通过提示等的隐式自适应。在共同目标下形式化这些方法,证明了显式参数自适应和隐式路由在特定条件下与核岭回归等价。提出实用设计原则和评估指标,识别了可扩展性等方面的开放问题。
英文摘要
Modern predictive systems are expected to adapt their behavior to the specific situation they are facing. A clinical model should not treat every patient the same; a retrieval-augmented model should change its answer when given different evidence; a mixture-of-experts model should route different inputs to different experts. We call this capability context-adaptive inference: before predicting, the system uses information about the current context to specialize its parameters or computation for that instance. This article provides a unified view of context-adaptive inference across three traditions that are usually treated separately: (i) explicit adaptation in statistics (e.g. varying-coefficient models, local regression, hierarchical sharing), (ii) rapid task-specific adaptation in meta-learning and transfer, and (iii) implicit adaptation in large foundation models via prompting, retrieval, and expert routing. We formalize these approaches under a common objective: to map context $c$ to adapted parameters $θ(c)$, then to predict via $f(x; θ(c))$. Under squared loss, linear prediction heads, and fixed features, we prove that explicit parameter adaptation and implicit routing are mathematically equivalent to kernel ridge regression on joint features of inputs and context. Building on this bridge, we propose practical design principles and evaluation metrics including adaptation-efficiency, routing stability, and context-specific robustness to guide when to specialize, how to constrain that specialization, and how to audit context-adaptive models in deployment. Finally, we identify open problems in identifiability, robustness under distribution shift, and efficient large-scale adaptation, outlining design principles for methods that are scalable, reliable, and transparent in real-world settings.
发表机构
- University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
- Carnegie Mellon University(卡内基梅隆大学)
- University of Washington(华盛顿大学)
- The Ohio State University(俄亥俄州立大学)
机构由 AI 辅助整理,请以论文原文为准。