arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.18553cs.LGcs.AI

循环语言模型中的操作原语内省:过程质量挖掘、可执行分支和读出控制边界

Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary

Jan Kirin

首次发表
浏览论文内容

中文总结 AI 辅助

研究语言模型能否读取计算质量及外部干预能否改善结果,通过在Ouro-RLTT中实验,利用隐藏状态等预测成功,构建相关机制,虽有可读属性但干预未提升能力,此为操作原语内省。

中文摘要 AI 辅助

我们在一个冻结的26亿参数循环变压器Ouro-RLTT中测试了两个问题:语言模型能否读取正在进行的计算质量,以及外部干预能否将该读出转化为更好的结果。在GSM8K上,一个严格的预答案探测排除了答案区域和黄金值,但仍能预测最终成功:隐藏状态加上长度和对数概率捷径的AUROC达到0.797,而仅捷径的AUROC为0.731。低容量挖掘还能读取角色专用属性。非循环控制复制了候选质量读出。我们构建了分支/携带/修剪机制,没有冻结干预能产生经过验证的能力提升。我们将这种可读但尚未可用的属性称为操作原语内省。

英文摘要

Can a language model read the quality of its ongoing computation, and can an external intervention turn that readout into better outcomes? We test both questions in a frozen 2.6B looped transformer, Ouro-RLTT. On GSM8K, a strict pre-answer probe excludes the answer region and gold value yet predicts success: hidden states plus length/log-probability features reach AUROC 0.797 versus 0.731 for those surface features alone (increment +0.066; task-clustered 95% CI [+0.021,+0.112]; 170 tasks). On Horizon Logic, a prospectively extended task-disjoint study gives an increment of +0.111 (CI [+0.056,+0.169]), independently replicated on the new cohort (+0.095) and robust to an adversarial malformed-sibling shortcut. Recurrence also moves candidate-quality readability to progressively earlier physical depth; the trend replicates across the Ouro family and qualitatively in out-of-family Huginn, although their transfer geometry differs. The readout converts into validated decision-level gains. Hidden-state-based scores improve risk-coverage over shortcut-only scores in four sealed selective-prediction arms, and terminal selection beats matched random even when every candidate is well formed (27/32 correct selections versus 64.8% expected; p = 0.0086). Generative control does not convert: directional steering is negative, a branch screen is bounded, and exact-compute loop allocation and minimal LoRA direction-binding detect no gain. These tests run through bit-exact branch/carry/prune machinery over Ouro's 192-slot recurrent cache, including a suffix-recompute splice saving up to 88% of per-branch layer passes. We call this decision-usable but not generatively controllable property operational proto-introspection. All load-bearing values use source-item-disjoint splits and antisymmetrized pairwise evaluation.

补充信息

↑