思维链熵度量的是什么?对脚手架、路由和内容的通道审计
What Does Chain-of-Thought Entropy Measure? A Channel Audit of Scaffolding, Routing, and Content
- Department of Statistics and Data Science, The Wharton School, University of Pennsylvania, Philadelphia, PA, USA(美国宾夕法尼亚大学沃顿商学院统计与数据科学系)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本研究通过通道审计分离思维链熵中的脚手架、路由与内容成分,证明内容约定在分叉与压缩上更优,并揭示重述答案对准确性的显著贡献。
中文摘要 AI 辅助
思维链令牌上的熵决定哪些令牌获得策略梯度、哪些被剪枝以及运行是否崩溃,然而每一个这样的统计量读取的是混合了三种选择的下一令牌分布:是否发出连接性脚手架、发出哪个连接词,以及实质性延续应该是什么。指定一个脚手架词汇子集可以精确地将这三者分开,对于熵、Kullback-Leibler散度和softmax策略的一阶熵速度都是如此。我们证明,在一个显式开放区域上,原始约定和内容约定在哪个位置是更大的分叉点上存在分歧,并通过内容通道加上一个由测量见证者认证的泄漏项来界定答案多样性。在二十三种配置中,脚手架一侧承载了原始高熵集合的多达41%;在匹配分词器的阶梯上,耦合仅在数学语料库步骤发生变化,而脚手架的熵份额在蒸馏过程中持续增长;来自一个通道相关性的闭式预测在54点范围内追踪选择保留,误差在五点以内,且无需拟合。在压缩方面,内容约定在每个单元上都优于原始惊异度;随后的一次答案泄漏审计纠正了我们自己的头条控制:重新馈送的链从重述答案中获得了其准确性的四分之一到一半,一旦剥离,没有令牌评分器能胜过随机连续块。
英文摘要
Entropy over chain-of-thought tokens decides which tokens receive the policy gradient, which get pruned, and whether a run has collapsed, yet each such statistic reads a next-token distribution mixing three choices: whether to emit connective scaffolding, which connective, and what the substantive continuation should be. Designating a scaffold vocabulary subset separates the three, exactly, for entropy, Kullback--Leibler divergence, and the first-order entropy velocity of a softmax policy. We prove the raw and content conventions disagree about which position is the larger fork on an explicit open region, and bound answer diversity by the content channel plus a leakage term a measured witness certifies. Across twenty-three configurations the scaffold side carries up to 41% of the raw high-entropy set; on a matched-tokenizer ladder, coupling changes only at the math-corpus step while the scaffold's entropy share keeps growing through distillation; a closed-form forecast from one channel correlation tracks selection retention over a 54-point range to five points, unfitted. On compression, the content convention beats raw surprisal in every cell; an answer-leakage audit then corrects our own headline control: re-fed chains earn a quarter to a half of their accuracy from restated answers, and once stripped, no token scorer beats a random contiguous block.