AI 中文总结
本研究提出结构感知微调(SAFT)方法,利用LoRA适配器和内在结构先验优化视觉语言模型奖励信号,实现更快策略收敛与更好对齐,为稳定文本条件强化学习提供可扩展路径。
AI 中文摘要
设计有效的奖励函数仍然是强化学习(RL)的主要瓶颈。近期研究采用大型基础视觉语言模型(VLM)作为奖励模型,通过计算文本-观测相似度来规避手动奖励工程。尽管该方法颇具前景,但这类奖励信号通常存在噪声且不可靠,限制了其在部署阶段的直接实用性。本文提出结构感知微调(SAFT),这是一种简单的自监督方法,无需真实标注即可在线优化这些不完善的奖励信号。SAFT利用内在结构先验,通过LoRA适配器对VLM的潜在空间进行正则化。我们在一系列基础模型能力范围内对SAFT进行严格评估,以证明其通用性。结果显示,SAFT可持续去除奖励空间的噪声,相较于基础模型能实现更快的策略收敛和显著提升的对齐度(EPIC距离),这表明失败往往可归因于结构脆弱性而非语义误解。通过用任务固有的结构归纳偏差替代大量人工偏好标注,SAFT为稳定文本条件强化学习提供了可扩展的路径,并凸显了将任务结构作为通用归纳偏差纳入的更广泛价值。
英文摘要
Designing effective reward functions remains a major bottleneck in Reinforcement Learning (RL). Recent work uses large foundation Vision-Language Models (VLMs) as reward models, computing text-observation similarity to bypass manual reward engineering. Although promising, these rewards are often noisy and unreliable, limiting their direct utility during deployment. We present Structure-Aware Fine-Tuning (SAFT), a simple, self-supervised method that refines these imperfect reward signals online without access to ground-truth supervision. SAFT leverages intrinsic structural priors to regularize the VLM's latent space via LoRA adapters. We rigorously evaluate SAFT across a spectrum of base model capabilities to demonstrate its versatility. Our results show that SAFT consistently denoises the reward landscape, yielding faster policy convergence and substantially improved alignment (EPIC distance) relative to the underlying base model, suggesting that failures can often be attributed to structural brittleness rather than semantic misunderstanding. By replacing extensive human preference annotation with structural inductive biases inherent to the task, SAFT offers a scalable path for stabilizing text-conditioned RL and underscores the broader value of incorporating task structure as a general inductive bias.