arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.16851cs.AI

AgentBrew:从强大教师到弱小语言模型智能体的终身知识酿造

AgentBrew: Lifelong Knowledge Brewing from Strong Teachers to Weak LLM Agents

Yangqin Jiang, Chao Huang

首次发表
浏览论文内容

中文总结 AI 辅助

研究语言模型智能体知识酿造问题,提出AgentBrew方法,含失败触发教师和学生感知合成两个组件,无需训练,能将教师交互经验提炼到学生外部记忆,经实验验证可产生强大且可部署的智能体。

中文摘要 AI 辅助

部署语言模型智能体通常在测试时需要一个精简的学生模型,即便训练时有更强的教师模型可用。我们研究知识酿造,即将教师的交互经验提炼到学生的持久外部记忆中。关键在于,这无需权重更新、专家示范、真实标签或测试时访问教师模型。此设置带来两个挑战:环境仅提供稀疏的二元反馈,且教师编写的笔记必须专门适配能力弱得多的学生以可具体执行。为克服这些障碍,我们提出AgentBrew,它由两个耦合组件组成。一是失败触发的教师——拉尔夫循环,通过将学生失败转化为经环境验证的笔记来减轻稀疏反馈。二是学生感知合成,将教师知识校准到弱执行器的操作粒度,产生特定于模型的可操作指导。在编码、数学和工具使用任务上的广泛评估和全面消融实验表明,这种不对称的、无需训练的酿造范式能产生能力强大且可部署的语言模型智能体。

英文摘要

Deploying LLM agents typically requires a compact test-time student, even if a stronger teacher is available during training. We study knowledge brewing: distilling a teacher's interactive experience into a persistent external memory for the student. Crucially, this requires no weight updates, expert demonstrations, ground-truth labels, or test-time teacher access. This setting poses two challenges: environments provide only sparse, binary feedback, and teacher-authored notes must be inherently tailored to be concretely executable by a substantially weaker student. To address these hurdles, we propose AgentBrew, comprising two coupled components. First, a failure-triggered teacher--Ralph Loop mitigates sparse feedback by transforming student failures into environment-validated notes. Second, student-aware synthesis calibrates teacher knowledge to the weak executor's operational granularity, yielding model-specific, actionable guidance. Extensive evaluations and comprehensive ablations across coding, math, and tool-use tasks demonstrate that this asymmetric, training-free brewing paradigm produces highly capable yet deployable LLM agents.

发表机构

  • The University of Hong Kong(香港大学)

机构由 AI 辅助整理,请以论文原文为准。

↑