arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26866cs.LG

边际正确的工具缓存可逆转组归一化策略更新

Marginally Correct Tool Caches Can Reverse Group-Normalized Policy Updates

Shivam Gupta

首次发表
浏览论文内容

中文总结 AI 辅助

本文证明在双动作模型中,即使工具缓存保持边际奖励分布,共享随机结果也可能逆转组归一化策略更新,表明边际有效性不足以保证训练等价性。

中文摘要 AI 辅助

工具结果缓存减少了智能体训练中的重复执行,但也耦合了 rollout 的随机性。我们研究了一个双动作模型,其中独立执行和共享执行保留了每个 rollout 的条件奖励分布。尽管存在这种边际一致性,每组共享一个随机结果可以逆转预期的组归一化策略更新。我们推导出一个精确的有限组表达式:针对恒定备择假设,共享更新遵循获胜概率减去失败概率,而非期望奖励之差。伯努利特例产生了一个错误方向区域,并且随着组规模增大,更新方差下限非零。在不进行组标准差缩放的情况下进行中心化,在该模型中保留了期望回报方向,使用了现有的估计器控制。穷举有限求和验证了 540 种配置和 3,240 次估计器评估,并配有独立的顺序序列检查器。一项实现审计在固定的、未修改的 TVCache 栈中使用 256 次脚本化 rollout 复现了共享路径。这些结果不衡量语言模型训练性能,也不反驳 TVCache 的确定性输出契约。它们确立了边际输出有效性本身不能证明随机缓存与训练等价。

英文摘要

Tool-result caching reduces repeated execution in agent training, but also couples rollout randomness. We study a two-action model in which independent and shared execution preserve every rollout's conditional reward distribution. Despite this marginal agreement, sharing one stochastic result per group can reverse the expected group-normalized policy update. We derive an exact finite-group expression: against a constant alternative, the shared update follows the probability of winning minus the probability of losing, rather than the difference in expected reward. A Bernoulli specialization yields a wrong-direction region and a non-vanishing update-variance floor as group size grows. Centering without group standard-deviation scaling preserves the expected-return direction in this model, using an existing estimator control. Exhaustive finite sums verify 540 configurations and 3,240 estimator evaluations, with a separate ordered-sequence checker. An implementation audit reproduces the sharing path in a pinned, unmodified TVCache stack using 256 scripted rollouts. These results do not measure language-model training performance or refute TVCache's deterministic-output contract. They establish that marginal output validity alone cannot certify a stochastic cache as training-equivalent.

补充信息

↑