arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

JIT-Agent:通过即时测试框架演进扩展测试框架智能

JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution

Guibin Zhang, Leo Lu, Fangzhou Xie, Kang Zhu, Junhao Wang, Zhifei Xie, Zhaochen Yu, Zihang Liu, Zhongxiang Sun, Qiankun Li, Yue Liao, Heng Chang, Xiaobin Hu, Qibing Ren, Wangchunshu Zhou, Chuanrui Hu, Yafeng Deng, Shuicheng Yan

arXiv 2608.25593首次发表:更新:

发表机构

LV-NUS Lab(LV-NUS实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

JIT-Agent是首个专为即时生成智能体测试框架设计的模型,可定制、修复、自我演进测试框架,提升DeepSeek、GLM等模型在多基准任务的性能,确立测试框架智能为智能体能力的独立可扩展维度。

AI 中文摘要

智能体的能力并非仅由模型决定,智能体测试框架(涵盖内存管理、规划策略、动作协议及工具/技能编排)对底层基础模型的贡献可占主导地位。然而测试框架设计仍为人工、任务特定且本质上不可扩展的。我们提出JIT-Agent,一种测试框架智能模型,可即时为任意现成的智能体LLM合成任务自适应的智能体测试框架。我们将智能体测试框架形式化为一种可组合、可机器生成的产物,受固定的四模块协议约束,并训练JIT-Agent为给定任务定制测试框架、修复测试框架以实现稳定可靠执行,还可通过从不断扩展的先前测试框架配置档案中提炼性能信号实现自我演进。配备JIT-Agent作为测试框架助手后,DeepSeek-V4-Flash在DeepSearchQA上超越GPT-5.6(提升9.1),在OdysseyBench上超越GPT-5.6(提升4.3),而已表现强劲的GLM-5.2提升幅度最高达20.2个百分点。在受控评估中,JIT-Agent生成的测试框架性能可与OpenCode、Claude Code等成熟智能体运行时相媲美,且持续改进DeepSeek V4、Mimo-V2.5和Qwen3.6等多规模模型家族。据我们所知,JIT-Agent是首个专为即时测试框架生成设计的模型,确立测试框架智能为智能体能力中可训练、可迁移且可复合的维度,与模型缩放正交。

英文摘要

Agent capability is not determined by the model alone. The agent harness, encompassing memory management, planning strategy, action protocol, and tool/skill orchestration, can dominate the contribution of the underlying foundation model. Yet harness design remains manual, task-specific, and fundamentally unscalable. We present JIT-Agent, a harness intelligence model trained to synthesize task-adaptive agent harnesses on the fly for arbitrary off-the-shelf agentic LLMs. We formalize the agent harness as a composable, machine-generatable artifact governed by a fixed four-module protocol, and train JIT-Agent to customize harnesses for a given task at hand, repair harnesses for stable and reliable execution, and self-evolve by distilling performance signals from an expanding archive of prior harness configurations. Equipped with JIT-Agent as a harness helper, DeepSeek-V4-Flash surpasses GPT-5.6 on DeepSearchQA (+9.1) and OdysseyBench (+4.3), while the already strong GLM-5.2 gains up to +20.2 points. Across controlled evaluations, JIT-Agent-generated harnesses are performance-competitive with mature agent runtimes such as OpenCode and Claude Code and consistently improve multi-scale model families of DeepSeek V4, Mimo-V2.5, and Qwen3.6. To our knowledge, JIT-Agent is the first model purpose-built for just-in-time harness generation, establishing harness intelligence as a trainable, transferable, and compounding dimension of agent capability orthogonal to model scaling.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑