科学领域终端环境的自监督扩展
Self-Supervised Scaling of Terminal Environments for Scientific Domains
- Tencent HY LLM Frontier(腾讯HY大模型前沿团队)
- University of Georgia(佐治亚大学)
- University of Maryland, College Park(马里兰大学帕克分校)
- National University of Singapore(新加坡国立大学)
- Indiana University(印第安纳大学)
- University of Illinois at Chicago(伊利诺伊大学芝加哥分校)
- Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出软件在环重构框架,利用现有科学软件自动生成参考输出与验证目标,构建终端智能体训练环境,实验表明微调后模型在Terminal-Bench 2上性能显著提升。
AI中文摘要:
终端智能体正越来越多地部署在软件工程之外的科学及其他专业领域。构建训练环境需要可执行的参考行为以及一个领域特定的验证器,用以区分语义正确性与表面看似合理的产物。为每项任务编写这些组件需要重复的工程工作,并限制了复用性。我们提出了软件在环重构(software-in-the-loop reconstruction),这是一种自监督框架,可从现有软件工作流(即将结构化输入映射到输出的可执行程序)中获取参考输出和验证目标。对于每个工作流,我们执行多种输入配置,并将案例划分为公开观测和隐藏评估。给定指令、输入模式以及公开的输入-输出观测,智能体在无法访问源工作流的情况下构建一个可编辑程序。候选程序在隐藏配置上依据工作流输出进行评估。一个分层验证器结合了领域特定的语义比较、结构有效性以及反捷径检查,而公开反馈则支持迭代修订。该构建过程无需为每项任务编写参考解决方案即可纳入额外的工作流和配置。我们用500个工作流和涵盖六个领域的46个软件家族实例化了SWR。在每项任务三次尝试中,Qwen3.8-Max解决了838个任务并产生了1,422条已验证轨迹,我们将其过采样至3,000个仅重构的训练示例。对Qwen3.8-27B进行监督微调,使其在三个随机种子上的平均Terminal-Bench 2性能从47.94%提升至53.56%,并在所有四项报告的评估中,于四个匹配词元语料对照组中取得了最高平均值。这些结果表明,现有科学软件能够为终端智能体提供可扩展的、行为验证的监督。
英文摘要:
Terminal agents are increasingly deployed beyond software engineering in science and other specialized domains. Constructing training environments requires executable reference behavior and a domain-specific verifier that distinguishes semantic correctness from superficially plausible artifacts. Authoring these components for each task requires repeated engineering and limits reuse. We introduce software-in-the-loop reconstruction, a self-supervised framework that obtains reference outputs and verification targets from existing software workflows, executable programs mapping structured inputs to outputs. For each workflow, we execute multiple input configurations and partition cases into public observations and hidden evaluations. Given the instruction, input schema, and public input--output observations, an agent constructs an editable program without access to the source workflow. The candidate is evaluated on hidden configurations against workflow outputs. A hierarchical verifier combines domain-specific semantic comparison, structural validity, and anti-shortcut checks, while public feedback supports iterative revision. The construction admits additional workflows and configurations without authoring a reference solution for each task. We instantiate SWR with 500 workflows and 46 software families across six domains. Across three attempts per task, Qwen3.8-Max solves 838 tasks and produces 1,422 verified trajectories, which we oversample to 3,000 reconstruction-only training examples. Supervised fine-tuning of Qwen3.8-27B improves mean Terminal-Bench 2 performance from 47.94% to 53.56% across three seeds and achieves the highest mean among four matched-token corpus controls on all four reported evaluations. These results indicate that existing scientific software can provide scalable, behaviorally verified supervision for terminal agents.