arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于熵的智能体选择性引导:从不完美的VLM教师学习自主策略

Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers

Giovanni Bonetta, Matteo Merler, Davide Zago, Rossella Cancelliere, Bernardo Magnini

arXiv 2609.01567首次发表:更新:

发表机构

Fondazione Bruno Kessler; University of Torino(布鲁诺·凯塞勒基金会; 都灵大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出SAGE框架,仅在智能体不确定时查询VLM教师,将引导提炼为轻量级RL策略,在稀疏奖励任务中学习的策略优于无引导RL,还减少了VLM使用量。

AI 中文摘要

视觉语言模型(VLM)为交互式决策提供了有用的先验,但直接将其用作策略成本高昂且脆弱:必须每一步都查询,无法从环境交互中改进,还会重复系统性错误。我们研究如何从在线、昂贵、不完美但信息丰富的VLM教师中学习低成本自主策略。我们提出SAGE(Selective Agent Guidance via Entropy,基于熵的智能体选择性引导)框架,仅在学习者不确定时查询VLM,训练时执行建议的动作,并将引导提炼为轻量级强化学习(RL)策略。由于VLM建议并非总是可靠,SAGE可使用环境衍生的优势对教师动作提炼进行加权,而非将所有建议视为同等有用。在稀疏奖励视觉推理和导航任务中,SAGE学习的策略在评估时无需VLM引导即可行动,在多个环境中优于无引导RL,包括所学策略超过其VLM教师的场景。结果表明,当VLM能帮助智能体发现高奖励轨迹时,选择性引导最有益;当无引导探索已成功或教师动作无法产生有益经验时,选择性引导的作用较小。SAGE还通过仅在部分训练步骤中提示教师,且部署时无需VLM调用,减少了VLM使用量。总体而言,我们的结果表明,VLM无需用作固定策略即可发挥作用;它们可作为临时、不完美的引导源,其价值通过交互得到检验并被内化。

英文摘要

Vision-Language Models (VLMs) provide useful priors for interactive decision-making, but using them directly as policies is expensive and brittle: they must be queried at every step, do not improve from environment interaction, and can repeat systematic errors. We study how to learn a cheap autonomous policy from an online, expensive, and imperfect but informative VLM teacher. We propose SAGE (Selective Agent Guidance via Entropy), a framework that queries a VLM only when the learner is uncertain, executes the suggested action during training, and distills guidance into a lightweight Reinforcement Learning (RL) policy. Because VLM advice is not always reliable, SAGE can weight teacher-action distillation using environment-derived advantages rather than treating all suggestions as equally useful. Across sparse-reward visual reasoning and navigation tasks, SAGE learns policies that act without VLM guidance at evaluation time and improves over unguided RL in several environments, including settings where the learned policy exceeds its VLM teacher. The results show that selective guidance is most beneficial when the VLM can help the agent discover high-reward trajectories, and less useful when unguided exploration already succeeds or teacher actions do not lead to informative experience. SAGE also reduces VLM usage by prompting the teacher only on a fraction of training steps and requiring no VLM calls at deployment. Overall, our results suggest that VLMs don't need to be used as fixed policies to be useful; they can instead act as temporary, imperfect sources of guidance whose value is tested and internalized through interaction.

Comments9 pages, 3 figures, 4 tables in the main text, 27 pages, 4 figures, 9 tables including Appendix

Journal refEMNLP 2026 Findings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑