arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03204cs.CLcs.AI

测试时对齐大型视觉语言模型:一种轨迹引导的结构化采样方法

Aligning Large Vision-Language Models at Test Time: A Trajectory-Guided Structured Sampling Approach

Tianbao Jiang, Weicong Ni, Gerard de Melo, Linlin Wang

AI总结:

针对现有视觉语言模型测试时对齐方法资源密集、训练与推理分布不匹配的问题,提出轨迹引导结构化采样的测试时对齐方法,在多模态推理数据集上提升准确率且推理开销可控。

AI中文摘要:

后训练强化学习(RL)算法通常用于将大型视觉语言模型(LVLMs)与人类意图及视觉推理任务的要求对齐。然而,现有的基于RL的对齐方法往往资源密集,且存在训练目标与推理时分布不匹配的问题。为弥合这一差距,我们提出一种新颖的测试时对齐方法,利用轨迹引导的结构化采样进行动态推理时优化,以实现与视觉接地的更好对齐并确保逻辑一致性。我们的方法首先通过轨迹学习算法构建推理记忆库,该算法将复杂问题求解分解为预定义推理模式的有序序列;随后,通过从推理记忆库收集轨迹以建立全局结构化推理先验,再使用迭代马尔可夫链蒙特卡洛(MCMC)算法对推理轨迹进行局部多目标优化,从而完成推理时对齐。在多个多模态推理数据集上的实验表明,我们的方法可显著提升准确率,且不会产生过高的推理开销。这些结果表明,轨迹引导的测试时采样是传统后训练对齐的可扩展且有效的替代方案,尤其适用于复杂视觉推理任务。

英文摘要:

Post-training reinforcement learning (RL) algorithms are commonly used to align large vision-language models (LVLMs) with human intent and the requirements of visual reasoning tasks. However, existing RL-based alignment methods are often resource-intensive and encounter mismatches between training objectives and inference-time distributions. To bridge this gap, we propose a novel test-time alignment approach that leverages trajectory-guided structured sampling for dynamic inference-time refinement, achieving better alignment with visual grounding and ensuring logical consistency. Our approach begins with curating a reasoning memory bank via a trajectory learning algorithm, which decomposes complex question solving into ordered sequences of predefined reasoning patterns. It subsequently accomplishes inference-time alignment by first collecting trajectories from reasoning memory bank to establish a global structural reasoning prior, and then using an iterative Markov Chain Monte Carlo (MCMC) algorithm for localized multi-objective refinement of the reasoning trace. Experiments across multiple multimodal reasoning datasets demonstrate that our approach significantly improves accuracy without incurring prohibitive inference overhead. These results establish trajectory-guided test-time sampling as a scalable and effective alternative to traditional post-training alignment, particularly for complex visual reasoning tasks.

↑