arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从提示到管控:从零开始构建Coderlet

From Prompt to Harness: Coderlet from Scratch

Mengfan Li

arXiv 2608.09480首次发表:更新:

AI 中文总结

本文研究了一种紧凑的编程智能体管控设计,通过模型、执行、状态边界连接各组件,实现模型生成到环境动作的转换,可逐步优化并在可执行工件中实现。

AI 中文摘要

单独的模型并不能决定编程智能体的行为方式。模型所感知的内容、动作如何进入环境、反馈如何返回、一次运行如何影响下一次运行,都取决于管控(harness)的组织方式。最小示例通常仅展示模型与工具之间的基本交互,而生产系统则将这些关系分散到复杂的组件和依赖项中。本文通过跟踪单个请求在上下文形成、模型决策、环境动作、观察返回和状态延续中的流程,研究了一种紧凑的管控设计。三个边界——模型、执行和状态边界——连接模型服务、工具环境和持久状态,而请求生命周期则决定了这些转换发生的顺序。它们共同展现了管控的核心作用:将模型生成内容转化为环境动作,将运行时反馈传递到后续决策,并允许状态在多个请求间延续。在此运行时结构之上,管控还可通过持续的自举在多次运行中逐步优化。该设计在可执行工件中实现,详见此https URL。

英文摘要

A model alone does not determine how a programming agent acts. What the model sees, how actions enter the environment, how feedback returns, and how one run affects the next all depend on how the harness is organized. Minimal examples usually show only the basic interaction between a model and tools, while production systems spread these relationships across complex components and dependencies. This paper studies a compact harness design by following a single request through context formation, model decision, environmental action, observation return, and state continuation. Three boundaries---model, execution, and state---connect the model service, tool environment, and persistent state, while the request lifecycle determines the order in which these transitions occur. Together, they show the harness's core role: turning model generations into environmental actions, carrying runtime feedback into later decisions, and allowing state to continue across requests. On top of this runtime structure, a harness can also be gradually refined across runs through continued bootstrapping. The design is realized in the executable artifact https://github.com/lilinxi/Coderlet.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑