arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34905cs.CV

ReSight-SMC:基于岛屿SMC与视觉侦察的两阶段功率采样

ReSight-SMC: Two-Stage Power Sampling via Island SMC with Visual Scouts

发表机构清华大学
查看机构详情
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Yaowen Zhang, Xiangyu Qiu, Junyi Hu, Zhi Lu, Wenwen Tian, Aoqin Wang, Junhai Luo, Zhenming Peng

首次发表
浏览论文内容

中文总结 AI 辅助

针对LVLM推理中功率采样未充分探索的问题,提出两阶段ReSight-SMC采样器,通过岛屿SMC与视觉侦察增强探索,在四个骨干和五个基准上超越Power-SMC,无需训练即媲美强化学习模型。

中文摘要 AI 辅助

功率采样已成为一种无需训练的LLM推理方法,通过锐化模型在完整回答上的分布,能够激发与强化学习相当的能力。尽管取得了成功,功率采样在大规模视觉语言模型(LVLMs)中仍未得到充分探索。我们将Power-SMC迁移到LVLM解码中,通过定义一个同时以图像和提示为条件的序列功率目标。这种直接迁移提供了一个强大的无需训练的基线,但留下了有限粒子多模态推理的两个方面未解决。在粒子层面,全局重采样可能导致谱系崩溃,而基于粒子的功率采样在多模态解码中无法通过不同的视觉线索使轨迹多样化,限制了在有限粒子预算下的探索。在回答层面,序列级锐化使得不同的推理轨迹即使支持相同的答案也会相互竞争。我们引入了ReSight-SMC,一种用于LVLM推理的无需验证器的两阶段功率采样器。其第一阶段使用祖先隔离的SMC岛屿来保留独立的轨迹族,并将有界数量的前缀条件视觉侦察路由到前缀相关的图像区域,同时抑制冗余重叠。每个侦察临时增加对图像标记的注意力并强调其路由区域。精确的重要性校正保留了基础LVLM序列功率目标。第二阶段按规范答案聚合终端重要性质量,对答案边缘进行幂运算,并采样一个答案及其支持轨迹。在四个LVLM骨干和五个基准上,ReSight-SMC在推理和感知基准组上均取得了比Power-SMC更强的聚合性能。无需后训练,它在聚合性能上与使用强化学习训练的骨干匹配模型保持竞争力。

英文摘要

Power sampling has emerged as a training-free approach to LLM reasoning, eliciting capabilities comparable to reinforcement learning by sharpening the model distribution over complete responses. Despite this success, power sampling remains underexplored in large vision-language models (LVLMs). We transfer Power-SMC to LVLM decoding by defining a sequence-power target conditioned on both the image and the prompt. This direct transfer provides a strong training-free baseline, but leaves two aspects of finite-particle multimodal inference unaddressed. At the particle level, global resampling can collapse genealogies, while particle-based power sampling does not diversify trajectories through distinct visual cues in multimodal decoding, limiting exploration under a finite particle budget. At the answer level, sequence-level sharpening makes distinct reasoning trajectories compete even when they support the same answer. We introduce ReSight-SMC, a verifier-free two-stage power sampler for LVLM inference. Its first stage uses ancestry-isolated SMC islands to preserve independent trajectory families and routes a bounded set of prefix-conditioned visual scouts to prefix-relevant image regions while discouraging redundant overlap. Each scout temporarily increases attention to the image tokens and emphasizes its routed region. Exact importance correction preserves the base LVLM sequence-power target. The second stage aggregates terminal importance mass by canonical answer, powers the answer marginal, and samples an answer together with a supporting trajectory. Across four LVLM backbones and five benchmarks, ReSight-SMC achieves stronger aggregate performance than Power-SMC over both the reasoning and perception benchmark groups. Without post-training, it remains competitive in aggregate with backbone-matched models trained using reinforcement learning.

↑