发表机构
Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出ViSkill框架,将成功交互编码为视觉技能卡片形成闭环反馈,在Sokoban等数据集上总体成功率达0.89,冷启动后升至0.91,优于基线且收敛更快。
AI 中文摘要
技能增强型智能体通过将成功轨迹提炼为可复用策略来提升样本效率,但现有多数方法以文本为中心,将空间布局与动作-状态对应关系线性化为语言,丢失了关键几何结构。近期研究已开始纳入视觉证据,但多将技能构建与更新和策略优化分开进行,未充分探索二者的相互改进。本文提出ViSkill,一种视觉原生技能学习框架,它将成功交互编码为复合视觉技能卡片,供VLM智能体直接使用。检索到的技能同时指导推理和奖励塑造,而成功轨迹会被提炼回技能库,形成技能积累与策略改进相互增强的闭环反馈。可选的冷启动机制可进一步加速早期学习。在Sokoban、FrozenLake和PrimitiveSkill上评估显示,ViSkill的总体成功率为0.89,采用冷启动初始化后升至0.91,优于所有评估的专有及开源基线,且收敛速度快于标准PPO。
英文摘要
Skill-augmented agents improve sample efficiency by distilling successful trajectories into reusable strategies. Yet most existing approaches remain text-centric, linearizing spatial layouts and action-state correspondences into language that loses critical geometric structure. Recent efforts have begun incorporating visual evidence, but construct and update skills separately from policy optimization, leaving their mutual improvement underexplored. We propose ViSkill, a visual-native skill learning framework that encodes successful interactions as composite visual skill cards directly accessible to VLM agents. Retrieved skills guide both inference and reward shaping, while successful trajectories are distilled back into the library, forming a closed feedback loop in which skill accumulation and policy improvement reinforce each other. An optional cold-start mechanism further accelerates early-stage learning. Evaluated on Sokoban, FrozenLake, and PrimitiveSkill, ViSkill achieves an overall success rate of 0.89, rising to 0.91 with cold-start initialization, outperforming all evaluated proprietary and open-source baselines while converging faster than standard PPO. Our code is available at https://github.com/ZJU-REAL/ViSkill.
CommentsCode: https://github.com/ZJU-REAL/ViSkill