解读点之间的信息:解码跨填充令牌的隐藏计算
Reading Between the Dots: Decoding Hidden Computation across Filler Tokens
浏览论文内容
中文总结 AI 辅助
研究前沿语言模型对无内容填充令牌的隐藏计算,通过分析两个前沿模型在四个任务家族中的表现,介绍无监督解码管道,能从隐藏状态恢复中间值,证明隐藏计算可从残差流读取。
中文摘要 AI 辅助
前沿语言模型可对诸如点或计数序列等无内容填充令牌进行多步推理,在无可见思维链情况下产生正确答案。在四个任务家族中,两个前沿模型以结构化、清晰方式对填充令牌进行计算。我们引入无监督解码管道,仅以隐藏状态为输入,在两个模型和所有四个任务上以80 - 95%的准确率恢复中间值。
英文摘要
Frontier LLMs can perform multi-step reasoning over content-free filler tokens like dots or counting sequences, producing correct answers with no visible chain-of-thought (CoT). This is a limit case for behavioral oversight, where surface tokens carry no information about the underlying reasoning. But hidden from the output is not the same as hidden from us. On four task families (fact retrieval, parallel numeric composition, string manipulation, and in-context computation), two open-weights frontier models (DeepSeek V3, Kimi K2) compute over filler tokens in a legible way: attention routes the question through the filler region to the answer, logit-lens readouts show retrieved facts emerging early and their composition crystallizing in late layers, and KV-cache transplants at filler positions causally swap outputs between examples. We introduce an unsupervised decoding pipeline that takes only hidden states as input and recovers intermediate values with 82-94% accuracy (best LLM judge) across both models and all four tasks, without ground-truth labels or training. Even without a judge, the hidden values are already directly in the pipeline's top-2 tokens 35-85% of the time. The uplift persists whether the filler is prefilled or the model generates the filler itself. On these cleanly decomposable tasks, hidden computation that defeats behavioral CoT monitoring is readable from the residual stream, which suggests that monitorability is a property of the model's full computational trace rather than only its surface tokens.
发表机构
- Harvard University(哈佛大学)
- Cambridge Boston Alignment Initiative(剑桥波士顿对齐计划)
- Massachusetts Institute of Technology(麻省理工学院)
- Anthropic
机构由 AI 辅助整理,请以论文原文为准。