arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

思维的方差:策略方差、关键分叉与局部信用分配

The Variance of Thought: Policy Variance, Critical Forks, and Local Credit Assignment

Yingru Li

arXiv 2608.22467首次发表:更新:

AI 中文总结

该研究针对长 horizon 语言模型任务的信用分配问题,分析策略方差,得出其作为发现预算、受基尼离散度约束及剩余 horizon 决定估计成本等结论,提出自举法消除成本的思路,支持对数值参数化。

AI 中文摘要

长 horizon 语言模型任务——多步推理与工具使用智能体——均受信用分配限制。我们通过策略方差σ_π²(s)=Var_{a~π}[Q_π(s,a)]分析该问题,在确定性马尔可夫决策过程(MDP)中,该方差是回报方差的唯一来源,且在被称为关键分叉的状态处以离散脉冲形式注入。得出三个结果:(i)策略方差是发现预算:观察优势为c的动作需要Ω(c²/σ_π²(s))次采样,该边界在经典两点分叉上是精确的;(ii)策略方差受策略的基尼离散度约束,σ_π²(s)≤1−||π(·|s)||₂²,这是一种无需展开的临界性必要条件,可仅从 logits 计算;(iii)剩余 horizon 决定估计成本:在下游成功概率为P的分叉处,蒙特卡洛优势估计的信噪比为√P量级,因此其样本成本按1/P缩放,分支采样也共享该成本。自举法通过将生存概率的乘积转换为和来消除该成本,前提是值表示具有乘法准确性,这支持对数值参数化。

英文摘要

Long-horizon language-model tasks --- multi-step reasoning and tool-using agents alike --- are limited by credit assignment. We analyze it through the policy variance $σ_π^2(s)=\operatorname{Var}_{a\simπ}[Q_π(s,a)]$, which in a deterministic MDP is the sole source of return variance and is injected in discrete pulses at states we call critical forks. Three results follow. (i) Policy variance is a discovery budget: observing an action of advantage $c$ requires $Ω(c^2/σ_π^2(s))$ draws, a bound that is exact on the canonical two-point fork. (ii) Policy variance is bounded by the policy's Gini dispersion, $σ_π^2(s)\le 1-\|π(\cdot|s)\|_2^2$, a rollout-free necessary condition for criticality computable from logits alone. (iii) The remaining horizon sets the estimation cost: at a fork whose downstream success probability is $P$, the Monte Carlo advantage estimate has signal-to-noise ratio of order $\sqrt{P}$, so its sample cost scales as $1/P$ --- a cost that branched sampling shares. Bootstrapping removes it by converting a product of survival probabilities into a sum, provided the value representation is multiplicatively accurate, which argues for log-value parameterization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑