arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00874cs.AI

超越时钟:衡量自适应修订的价值

A Score Is Not a Policy: Measuring the Value of Adaptive Revision

Ayushi Chadha

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对分层潜在推理器的元级控制问题,发现学习的自适应修订策略未优于最佳强制时机策略,强固定调度可捕获大部分时机价值,提出需从三方面评估元级控制。

中文摘要 AI 辅助

随着智能体系统成为复合系统,越来越重要的决策已超出任务执行本身:高级控制器应何时保留指导另一进程的策略,又应何时对其进行修订?我们在分层潜在推理器中研究这一元级控制问题,该推理器的管理器可保留或替换对下层计算的承诺。在三个预先提交的训练种子中,学习到的修订时机产生了性质不同的策略,从几乎确定性的早期时钟到更依赖状态的调度分布,然而 none 在相同冻结检查点上评估时,均未优于最佳强制时机策略。这将状态依赖性与决策价值区分开来:控制器可根据内部状态改变其行动,而无需将这种改变转化为可复现的任务性能收益。对原始检查点的深入干预研究表明,时机本身具有重要性且对顺序敏感,而穷举枚举显示,在该决策预算下,强固定调度捕获了时机可测价值的大部分。反事实 PERSIST/REPLAN 诊断进一步表明,当可预测性由决策位置而非位置内判别主导时,分数级证据为何可能具有误导性。综上,这些结果主张应沿三个独立轴评估学习到的元级控制:其分数是否依赖状态,该依赖是否改变已实现的行为,以及这些改变是否捕获了强非自适应策略之外的结果价值。

英文摘要

As agentic systems become compound systems, increasingly important decisions move above task execution itself: when should a higher-level controller preserve the strategy guiding another process, and when should it revise it? We study this meta-level control problem in a hierarchical latent reasoner whose manager can retain or replace a commitment governing lower-level computation. Across three precommitted training seeds, learned revision timing produces qualitatively different policies, ranging from an almost deterministic early clock to substantially more state conditioned schedule distributions, yet none outperforms the best forced timing policy evaluated on the same frozen checkpoint. This separates state dependence from decision value: a controller can vary its actions with internal state without turning that variation into a reproducible task-performance benefit. A deeper intervention study on the original checkpoint shows that timing itself is consequential and order-sensitive, while exhaustive enumeration reveals that a strong fixed schedule captures most of the measurable value available from timing at this decision budget. Counterfactual PERSIST/REPLAN diagnostics further show why score-level evidence can be misleading when predictability is dominated by decision position rather than within-position discrimination. Together, these results argue that learned meta-level control should be evaluated along three separate axes: whether its score depends on state, whether that dependence changes realized behavior, and whether those changes capture outcome value beyond a strong non-adaptive policy.

补充信息

↑