发表机构
SAP Lab(SAP实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对终端智能体训练中数据生成与验证的挑战,提出元智能体流程并诊断三类失败,证明可解性区间是模型特定的,需将校准与审计作为一等评估标准。
AI 中文摘要
使用像Claude Opus这样的前沿模型作为元智能体来生成终端任务和验证器以进行强化学习训练越来越常见。然而,一个可运行的Docker镜像和可执行的测试套件并不能保证终端智能体训练的端到端流程是可靠的。基于这一差距,我们提出了一个元智能体流程,诊断出三类失败:基准无效性、测试框架脆弱性和奖励错位。提示词重新设计和上下文扩展将基线可解性提高了5.6倍,但一个9B模型在Claude Opus生成的任务上,在20步内平均pass@2饱和于81.3%。在不改变训练配置的情况下,添加困难任务将平均pass@2降至20.6%,这强有力地证明可解性区间是模型特定的。这些发现表明,元智能体的可靠性需要将可解性区间校准、验证器审计和基础设施错误核算作为一等评估标准,而非事后诊断。
英文摘要
Using a frontier model like Claude Opus as a meta-agent to generate terminal tasks and verifiers for RL training is increasingly common. Yet a runnable Docker image and executable test suite do not guarantee a faithful end-to-end pipeline for terminal agent training. We present a meta-agent pipeline motivated by this gap, diagnosing three classes of failure: benchmark invalidity, harness brittleness, and reward misalignment. Prompt redesign and context extension raise baseline solvability 5.6 times, but a 9B model saturates at 81.3% mean pass@2 within 20 steps on Claude Opus-generated tasks. Adding hard tasks reduces mean pass@2 to 20.6% without changing the training configuration, a strong evidence that the solvability band is model-specific. These findings demonstrate that meta-agent reliability requires solvability-band calibration, verifier audits, and infrastructure error accounting as first-class evaluation criteria, not post-hoc diagnost.