发表机构
Cabal AI; Para AI(Cabal AI; Para AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究智能体计算机使用强化学习中单次运行结果的误导性,通过验证器引导修复35B智能体,发现成功率受上游方差主导,修复能力分两层,框架级修复有条件,还发现自身过度断言,发布库用于k种子报告,首次进行多模态分段聚合策略内自蒸馏更新。
AI 中文摘要
报告了单次运行中智能体计算机使用强化学习的情况,但这些数字具有误导性。通过在五个神谕分级环境中使用验证器引导修复一个35B计算机使用智能体(CUA),我们发现修复策略的成功率主要受上游方差影响。评估方差可忽略不计,训练种子效应较小。在最难的单元中,数据抽取的份额占主导。修复能力有两层,框架级修复仅在纠正行动是任务唯一剩余阻碍时才转化为任务成功。通过跨种子复制发现了自身的过度断言,压力测试表明单次运行改进可能有错误的结果。我们发布了一个库用于常规的k种子报告。该装置是首个对真实35B CUA策略进行的多模态分段聚合策略内自蒸馏(SA - OPSD)更新。
英文摘要
Agentic computer-use RL is reported in single runs, and those numbers mislead. Using verifier-guided repair of a 35B computer-use agent (CUA) across five oracle-graded environments, we show a repaired policy's success rate is dominated by upstream variance: a variance-components decomposition across three cells (crossed data-draw $\times$ seed grid, bootstrap CIs) finds evaluation variance negligible ($σ_{\mathrm{eval}} \approx 0$) and the training-seed effect small everywhere ($\leq 10\%$); instead it splits between the data draw and run-to-run nondeterminism, the data draw's share rising to dominant ($48\%$) on the hardest cell. There the run-to-run distribution is bimodal (Hartigan dip $p=0.07$, $k=10$), so a single run has roughly a 30% chance of the failure mode and mean$\pm$std is the wrong summary. On that footing, two findings hold. First, repairability is two-tier in how constrained the corrective action is: a single fixed token installs reliably (done-detection $0.97\pm0.06$), while open-ended corrections are only partial -- spatial-coordinate clicks (grounding $0.53\pm0.35$) and a generative field-fill ($0.14\pm0.04$). Second, the frame-level repair transfers to task success only when the corrective action is the task's sole remaining blocker (LinkedIn 8/20 vs. base 0/15, Fisher $p=0.006$). We caught two of our own over-claims -- a sample-efficiency curve and a 'grounding cannot be bought' boundary -- only by replicating across seeds; a stress test makes the stakes external: a single-run improvement of the size this field publishes would have the wrong sign roughly one-third of the time in a comparable regime. We release a library (cua_reliability) for routine k-seed reporting. The apparatus is, to our knowledge, the first multimodal segment-aggregated on-policy self-distillation (SA-OPSD) update on a real 35B CUA policy.
Comments15 pages, 3 figures. Reliability protocol and library (cua_reliability), completion verifier, and leakage-free held-out plus bootstrap evaluation harness are open-source