发表机构
Tencent; Indiana University; University of Maryland, College Park; University of Georgia; National University of Singapore(腾讯; 印第安纳大学; 马里兰大学帕克分校; 佐治亚大学; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究智能体框架演进中行为定位难的问题,提出通过Harness手册和行为引导的渐进式披露,以行为为中心自动合成框架表示并辅助规划,提高行为定位和编辑计划质量,助力复杂智能体系统发展。
AI 中文摘要
现代人工智能智能体的能力不仅取决于其基础模型,还取决于其框架,框架用于构建提示、管理状态、调用工具和协调执行。随着模型、API、环境和需求的发展,框架必须不断修改。在进行此类更改之前,开发人员或编码智能体必须识别实现目标行为的所有代码位置。这很困难,因为生产框架庞大、紧密耦合且行为分散,而修改请求描述系统应做什么,存储库按文件和模块组织。代码搜索、存储库索引和长上下文处理便于检查,但仍需手动恢复行为到代码的映射。行为定位因此是框架演进的核心瓶颈。我们引入了Harness手册,这是一种以行为为中心的表示,通过静态分析和LLM辅助结构化从框架代码库自动合成,将每个行为与其相应源链接起来。我们还引入了行为引导的渐进式披露(BGPD),它引导智能体从高级行为到相关实现细节,并根据当前源验证候选位置。在来自两个开源框架的各种修改请求上,手册辅助规划提高了行为定位和编辑计划质量,同时使用更少的规划器令牌,在分散站点、很少执行的路径和跨模块交互方面收益最大。因此,不断发展复杂的智能体系统不仅取决于生成编辑,还取决于确定这些编辑应在何处进行。
英文摘要
The capability of a modern AI agent depends not only on its foundation model but also on its harness, which constructs prompts, manages state, invokes tools, and coordinates execution. As models, APIs, environments, and requirements evolve, the harness must be continually modified. Before such a change can be made, a developer or coding agent must identify all code locations that implement the target behavior. This is difficult because production harnesses are large, tightly coupled, and behaviorally distributed, while modification requests describe what the system should do and repositories are organized by files and modules. Code search, repository indexing, and long-context processing ease inspection, but still leave this behavior-to-code mapping to be recovered by hand. Behavior localization is therefore a central bottleneck in harness evolution. We introduce the Harness Handbook, a behavior-centric representation synthesized automatically from a harness codebase via static analysis and LLM-assisted structuring, linking each behavior to its corresponding source. We also introduce Behavior-Guided Progressive Disclosure (BGPD), which guides agents from high-level behaviors to relevant implementation details and verifies candidate locations against the current source. On diverse modification requests from two open-source harnesses, Handbook-Assisted planning improves behavior localization and edit-plan quality while using fewer planner tokens, with the largest gains on scattered sites, rarely executed paths, and cross-module interactions. Evolving complex agentic systems thus depends not only on generating edits, but also on determining where those edits should be made.
Comments29 pages, 6 figures. Project page: https://ruhan-wang.github.io/Harness-Handbook/