长期分裂的帝国必须统一:三种大语言模型智能体框架的架构趋同
The Empire, Long Divided, Must Unite: Architectural Convergence in Three LLM Agent Harnesses
浏览论文内容
中文总结 AI 辅助
通过对LangChain deepagents、Earendil pi、DeepSeek dsh三种编码智能体框架的多案例研究,发现它们虽初始理念迥异却趋同于含五个元素的架构,同时指出外部可验证性是未趋同的关键维度。
中文摘要 AI 辅助
智能体框架是将语言模型转化为自主智能体的组件:它是构建模型上下文、协调工具、运行循环并在长周期运行中持久化状态的周边代码。这一层而非其封装的模型,正日益成为智能体行为的约束因素。我们对三种基于截然不同理念构建的开源编码智能体框架开展了源代码层面的多案例研究:LangChain的deepagents(配备全套工具)、Earendil的pi(极端极简主义)以及DeepSeek的dsh(一切皆为插件)。通过读取每个框架的固定提交版本并跟踪其提交历史,我们发现两个成熟框架朝着相反方向发展(deepagents减少了定制脚手架,pi则增加了持久基础设施),却都趋同于包含五个重复元素的中间架构形式:通用循环、仅追加的可重放会话记录、作为数据保留的模型特性、上下文的渐进式展示以及显式扩展接缝。第三个框架作为保留样本后续读取,展现了全部五个元素,且在一个接缝处直接复用了另一个框架的实现。因此,我们不主张独立发明,并将这种趋同分解为平行发现、扩散和字面复用。最后,一个关键维度未出现趋同,甚至根本不存在:外部可验证性,即外部方无需信任运行时即可检查的防篡改记录。我们认为这种缺失并非疏忽,而是一种预测性缺口,是针对来源敏感领域的智能体框架将产生差异的下一个维度。
英文摘要
An agent harness is what turns a language model into an autonomous agent: the surrounding code that builds the model's context, mediates its tools, runs the loop, and persists state across a long-horizon run. This layer, not the model it wraps, is increasingly the binding constraint on agent behaviour. We present a source-level, multi-case study of three open coding-agent harnesses built from deliberately opposing philosophies: LangChain's deepagents (batteries-included), Earendil's pi (radical minimalism), and DeepSeek's dsh (everything-is-a-plugin). Reading each at a pinned commit and following its commit history, we find that the two mature harnesses have travelled in opposite directions (deepagents subtracting authored scaffolding, pi accreting durable infrastructure), yet converged toward one architectural middle form of five recurring elements: a commoditised loop, an append-only replayable session record, model quirks kept as data, progressive disclosure of context, and explicit extension seams. A third harness, read afterward as a held-out check, exhibits all five, and in one seam reuses another's implementation outright. We therefore do not claim independent invention, and decompose the convergence into parallel discovery, diffusion, and literal reuse. Finally, one load-bearing dimension shows no convergence, and indeed no presence: external verifiability, a tamper-evident record an outside party can check without trusting the runtime. We read this absence not as an oversight but as a predictive gap, the next axis on which harnesses for provenance-sensitive domains will differ.