发表机构
Autopoiesis Sciences(自创生科学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出元认知引导方法,通过科学家交互轨迹中的对比干预识别Kimi 2.6的低维控制结构,实现推理时动态调控探索、收敛与批判性评估,并在自主研究系统中验证其有效性。
AI 中文摘要
长期科学发现需要智能体在证据变化时交替进行探索、纪律性执行和批判性重新评估。当前语言模型主要基于科学产物进行训练,并使用结果级信号进行优化,为这些过程级的科学判断转变提供的监督有限。我们研究这种判断是否可以从科学家交互轨迹中恢复,并用于控制冻结前沿模型的内部计算。利用在真实科学研究过程中收集的对比干预,我们在Kimi 2.6(一个万亿参数的混合专家模型)中识别出一个协调的、低维的控制结构。残差分析、注意力权重子空间对齐和跨层奇异值分解收敛于一个跨越关键层的中间深度控制面。我们引入了元认知引导(Metacognitive Steering),一种推理时控制器,它读取模型的认知状态,并动态组合针对探索、程序收敛或批判性重新评估的层特定干预,而无需修改模型参数。行为分析表明,这种控制产生了更持续的探索、显式剪枝和证据响应的综合。我们将该方法应用于Columbus-1(一个自主研究系统)中,该系统在BlueZ中识别了八个独立复现的、攻击者可利用的漏洞,并指导了设计、模拟和制造一枚十英尺长的火箭,该火箭旨在使用不可节流的固体发动机进行推进式着陆。这些结果共同表明,过程级科学判断可以为模型推理策略的可解释、动态控制提供监督。
英文摘要
Long-horizon scientific discovery requires agents to alternate between exploration, disciplined execution, and critical reassessment as evidence changes. Current language models are trained primarily on the products of science and optimized using outcome-level signals, providing limited supervision for these process-level shifts in scientific judgment. We investigate whether such judgment can be recovered from scientist interaction traces and used to control the internal computation of a frozen frontier model. Using contrastive interventions collected during real scientific research, we identify a coordinated, low-dimensional control structure within Kimi 2.6, a trillion-parameter mixture-of-experts model. Residual analysis, attention-weight subspace alignment, and cross-layer singular value decomposition converge on a mid-depth control surface spanning key layers. We introduce Metacognitive Steering, an inference-time controller that reads the model's cognitive regime and dynamically composes layer-specific interventions for exploration, procedural convergence, or critical reassessment without modifying model parameters. Behavioral analyses show that this control produces more sustained exploration, explicit pruning, and evidence-responsive synthesis. We operationalize the method in Columbus-1, an autonomous research system that identified eight independently reproduced, attacker-reachable vulnerabilities in BlueZ and directed the design, simulation, and fabrication of a ten-foot rocket intended to land propulsively using non-throttleable solid motors. Together, these results show that process-level scientific judgment can provide supervision for interpretable, dynamic control over a model's reasoning strategy.