arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向小型语言模型智能体的感知工具环境蒸馏

Harness-Aware Distillation for Small Language Model Agents

Moonseok Choi, Taehong Moon, Giung Nam, Juho Lee

arXiv 2610.02858首次发表:更新:

发表机构

KAIST AI(韩国科学技术院人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出感知工具环境蒸馏(HAD),通过对比教师有无工具环境信息的动作偏好和有效性检查,提升小型语言模型智能体在长时程任务中的性能,无需额外奖励或标签。

AI 中文摘要

语言模型智能体在部署时配备一个工具环境,即模型周围的软件,用于管理其上下文、工具和反馈。当这样的智能体被蒸馏为较小的模型时,工具环境保持不变,因此学生模型主要需要工具环境无法提供的教师特定能力,例如根据工具环境信息正确行动。然而,标准蒸馏模仿教师的全部输出,并将工具环境视为输入的一部分。我们提出了感知工具环境蒸馏(HAD),将蒸馏重点放在教师超越工具环境所提供的内容上。HAD补充了在线策略蒸馏,包含两个组件:一个动作偏好,对比同一教师在有和没有工具环境信息时的动作,并在学生自身推理后进行评分;以及一个有效性检查,丢弃其偏好动作与工具环境记录相矛盾的偏好对。我们表明,这种对比为学生提供了仅模仿教师无法提供的信息,且HAD不需要任务奖励、成功标签或未来信息。在多个长时程智能体基准和模型上,HAD优于使用相同固定工具环境的在线策略蒸馏基线。我们的分析显示,HAD比基线进入更少的无效循环,并且更频繁地从错误中恢复,这表明它自适应地将可学习的反馈保留在其权重中,同时从工具环境读取状态信息。

英文摘要

Language model agents are deployed with a harness, the software around the model that manages its context, tools, and feedback. When such an agent is distilled into a smaller one, the harness stays in place, so the student mainly needs the teacher-specific abilities that the harness cannot provide, such as acting correctly on harness information. Standard distillation, however, imitates the teacher's full outputs and treats the harness as part of the input. We propose Harness-Aware Distillation (HAD), which focuses distillation on what the teacher adds beyond the harness. HAD complements on-policy distillation with two components: an action preference that contrasts the same teacher's actions with and without the harness information, scored after the student's own reasoning, and a validity check that drops preference pairs whose preferred action contradicts the harness records. We show that the contrast gives the student information that imitating the teacher alone cannot provide, and HAD needs no task rewards, success labels, or future information. Across multiple long-horizon agent benchmarks and models, HAD outperforms on-policy distillation baselines with the same fixed harness. Our analysis shows that HAD enters fewer unproductive loops and recovers from errors more often than the baselines, and suggests that it adaptively keeps learnable feedback in its weights while reading state information from the harness.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑