AI 中文总结
研究针对训练后智能体优化问题,提出协同控制框架,交替优化控制与模型参数。通过基于大语言模型的控制批评家分析失败轨迹并提出更新,再用改进控制生成的轨迹微调模型,案例研究显示该方法能提升智能体性能。
AI 中文摘要
训练后的自动化人工智能研究智能体,不仅需要优化模型参数,还需优化运行时控制,后者决定了研究轨迹的生成、评估和学习方式。现有流程通常在固定控制下训练模型,这导致模型更新与决定轨迹质量的静态框架不匹配。我们引入了协同控制框架,在训练后联合优化智能体控制和模型参数。该框架在控制优化和模型优化之间交替进行。基于大语言模型的控制批评家分析失败轨迹,识别控制层面的失败模式,并提出有效的局部更新。然后,模型在改进控制生成的高质量轨迹上进行微调,将有效的框架提炼到模型参数中。一个200多小时的自主案例研究表明,协同控制可以从系统崩溃中恢复,提高推理效率,并在无需人工干预的情况下发现集成策略。这些结果表明,联合控制和模型优化是超越固定控制训练后提升智能体的有效方法。
英文摘要
Post-training agents for automated AI research requires optimizing not only model parameters, but also the runtime harness that shapes how research trajectories are generated, evaluated, and learned from. Existing pipelines typically train models under a fixed harness, including prompts, tools, skills, middleware, and memory, while leaving the data-generating process outside the optimization objective. This creates a mismatch between model updates and the static scaffolding that determines trajectory quality. We introduce Co-Harness, a framework that jointly optimizes the agent harness and model parameters during post-training. Co-Harness alternates between harness optimization and model optimization. An LLM-based HarnessCritic analyzes failed trajectories, identifies harness-level failure modes, and proposes validated local updates. The model is then fine-tuned on high-quality trajectories generated by the improved harness, distilling effective scaffolding into model parameters. A 200+ hour autonomous case study further shows that Co-Harness can recover from system crashes, improve inference efficiency, and discover ensemble strategies without human intervention. These results suggest that joint harness and model optimization is an effective way to improve agents beyond fixed-harness post-training.