arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CREDO:基于方差引导的评分规则演化用于重放修正的信用分配

CREDO: Variance-Guided Rubric Evolution for Replay-Corrected Credit Assignment

Xuchun Hu

arXiv 2609.24174首次发表:更新:

AI 中文总结

针对长视界语言智能体的稀疏奖励问题,提出CREDO框架,通过演化语义评分规则与选择性重放修正信用分配,实现方差引导的预算分配,并给出理论保证与评估协议。

AI 中文摘要

长视界语言智能体仅获得稀疏的终端反馈,而中间评分规则提供了结构化但可能错误指定的进展评估。在可重置的训练环境中,反事实延续推演可以衡量局部信用,但穷举重放成本高昂。我们提出Credo框架,该框架将演化的语义评分规则与选择性的、基于执行的信用修正相结合。一个冻结的评判器将可见转换映射到评分规则特征,一个信用头预测与已实现转换相关的期望终端奖励变化。独立采样的双侧重放利用其记录的包含概率来修正预测残差。我们推导了条件无偏性以及一个方差分解,该分解将两个设计选择联系起来:保留哪些评分规则特征,以及在何处分配固定的期望重放预算。由此产生的准则根据策略分数敏感性和缺失的重放覆盖度对预测误差进行加权;其分配规则额外考虑了延续成本。我们还描述了一种与终端留一法优势相结合的实用混合方案,并将其裁剪的、令牌归一化的PPO实现与理想的策略梯度估计器区分开来。本初步报告提供了方法、证明、精确的有限模型审计以及受控评估协议。它并未声称在语言智能体基准上具有经验优越性。

英文摘要

Long-horizon language agents receive sparse terminal feedback, while intermediate rubrics provide structured but potentially misspecified assessments of progress. In resettable training environments, counterfactual continuation rollouts can measure local credit, but exhaustive replay is costly. We propose Credo, a framework that couples evolving semantic rubrics with selective, execution-based credit correction. A frozen judge maps visible transitions to rubric features, and a credit head predicts the change in expected terminal reward associated with the realized transition. Independently sampled two-sided replays correct prediction residuals using their recorded inclusion probabilities. We derive conditional unbiasedness and a variance decomposition that connects two design choices: which rubric features to retain, and where to allocate a fixed expected replay budget. The resulting criterion weights prediction errors by policy-score sensitivity and missing replay coverage; its allocation rule additionally accounts for continuation cost. We also describe a practical mixture with terminal leave-one-out advantages and distinguish its clipped, token-normalized PPO implementation from the ideal policy-gradient estimator. This preliminary report provides the method, proofs, an exact finite-model audit, and a controlled evaluation protocol. It makes no claim of empirical superiority on language-agent benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑