发表机构
ServiceNow; Mila; Université de Montréal(ServiceNow公司; 米拉研究所; 蒙特利尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
StarHarness是一种固定模型权重的harness进化框架,经分层搜索优化后,在ITBench等3个企业任务基准上性能提升20-35个百分点,且可跨GPT、Qwen模型迁移,缓解模型-环境不匹配问题。
AI 中文摘要
我们提出StarHarness,这是一种在保持模型权重固定的情况下进化特定环境智能体harness(工具适配层)的框架。进化后的harness可包含提示与任务框架、工具接口、技能、MCP支持的提供商、子智能体结构及智能体循环配置。StarHarness通过根据基线失败行为对任务分层来构建紧凑的进化池,将提议者可见的搜索任务与提议者隐藏的选择任务分离,并保留预留任务用于评估泛化能力。在ITBench SRE、EnterpriseOps-Gym ITSM和AutomationBench Finance基准测试中,每个环境经过4至12次接受的变更后,harness进化使完整基准性能相比默认harness提升了20至35个百分点。这些增益在进化中排除的任务上依然存在,且无需重新进化即可在GPT和Qwen模型系列间迁移。轨迹分析将改进归因于接口修复、环境规范及压缩搜索的操作知识,在多种场景下减少了假阳性诊断并缩短了轨迹。因此,StarHarness提供了一种实用方法,可减少工具丰富的企业任务中持续存在的模型-环境不匹配问题。
英文摘要
We present StarHarness, a framework for evolving environment-specific agent harnesses while keeping model weights fixed. The evolved harness can include prompt and task framing, tool interfaces, skills, MCP-backed providers, subagent structure, and agent-loop configuration. StarHarness constructs a compact evolution pool by stratifying tasks according to baseline failure behavior, separates proposer-visible search tasks from proposer-hidden selection tasks, and reserves held-out tasks for evaluating generalization. Across ITBench SRE, EnterpriseOps-Gym ITSM, and AutomationBench Finance, harness evolution improves full-benchmark performance by 20-35 percentage points over the default harness after 4-12 accepted changes per environment. These gains persist on tasks excluded from evolution and transfer without re-evolution across GPT and Qwen model families. Trace analysis links the improvements to interface repairs, environment conventions, and operational knowledge that compresses search, with fewer false-positive diagnoses and shorter trajectories in several settings. StarHarness therefore offers a practical way to reduce persistent model-environment mismatch in tool-rich enterprise tasks.