推理策略的测试时自适应:基于贝叶斯非参数记忆
Test-Time Adaptation of Reasoning Strategies with Bayesian Nonparametric Memory
浏览论文内容
中文总结 AI 辅助
本文提出贝叶斯速查表(HDP-GMM),在测试时通过软更新记忆并动态重组行为聚类,实现低成本自适应,在多个推理基准上超越现有记忆模块。
中文摘要 AI 辅助
尽管现代大型语言模型(LLMs)已被训练为通过言语化的思维链进行推理,但由于达到最终答案的路径次优,生成成本大幅增长。此外,当在观察各种输入查询(例如通过自我反思)时发现新的见解,现有的机制有限,无法将这些发现延续应用于后续问题。人们可以将此类策略或行为的列表视为一个不断增长的速查表,并在推理时从该记忆模块中检索元素。在本工作中,我们考虑结构化的速查表,其中包含学习到的行为聚类。我们引入了一个层次狄利克雷过程高斯混合模型(HDP-GMM)作用于行为嵌入,该模型在不同领域间共享组件,同时允许领域特定的混合权重,并使用后验预测来检索查询的相关行为;我们将其称为“贝叶斯速查表”。该机制允许在在线测试时训练(TTT)设置中进行廉价的自适应,在每个样本后软更新混合的充分统计量,并在合成行为足够新颖时创建新组件。我们证明,与现有记忆模块相比,贝叶斯速查表在诸如AIME'25、Omni-MATH和PhysReason等推理基准上取得了明显的性能提升,即使在冷启动设置下也是如此。我们表明,贝叶斯速查表是一个自适应重组的记忆模块,因为行为可以通过一步折叠吉布斯采样重新分配到组件中。我们的研究结果强调了贝叶斯启发的记忆模块在有效测试时自适应中的价值,以及结构在元认知推理中的作用。
英文摘要
While modern large language models (LLMs) have been trained to reason through verbalized chains-of-thought, the generation cost grows substantially due to suboptimal paths to reach the final answer. Furthermore, as new insights are discovered while observing various input queries (e.g. through self-reflection), limited mechanisms exist for carrying forward these findings to be applied to subsequent problems. One can view the list of such strategies or behaviors as a growing cheatsheet, with elements retrieved from this memory module at inference-time. In this work, we consider structured cheatsheets, with learned clusters of behaviors. We introduce a Hierarchical Dirichlet Process Gaussian Mixture Model (HDP-GMM) over behavior embeddings, which shares components across domains while allowing domain-specific mixing weights, and uses the posterior predictive to retrieve relevant behaviors for a query; we call this a $\textit{Bayesian Cheatsheet}$. This mechanism allows for cheap adaptation in an online test-time training (TTT) setting, softly updating the mixture's sufficient statistics following each sample and enabling the creation of new components when the synthesized behaviors are sufficiently novel. We demonstrate that Bayesian Cheatsheet achieves clear performance gains relative to existing memory modules across reasoning benchmarks such as AIME'25, Omni-MATH, and PhysReason, even in the cold-start setting. We show that the Bayesian Cheatsheet is an adaptively reorganizing memory module, as behaviors can be re-assigned to components through a single step of collapsed Gibbs sampling. Our findings highlight the value of Bayesian-inspired memory modules for effective test-time adaptation and the role of structure in metacognitive reasoning.
发表机构
- IBM Research AI(IBM研究院AI)
机构由 AI 辅助整理,请以论文原文为准。