arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BUS:用于高级多模态推理的受脑启发的无监督自我反思

BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning

Jiacheng Yang, Tongying Xiao, Yunkai Dang, Cong Wang, Yuekun Yang, Qi Fan, Tianyu Ding, Wenbin Li, Feng Miao, Yang Gao

arXiv 2607.07361首次发表:更新:

发表机构

AAAI Press(美国人工智能协会出版社)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对VLM处理复杂视觉任务的不足,受脑启发提出BUS无监督训练框架,通过反向预测提升其反思推理能力,无需标注数据,与多种微调方法兼容,经实验验证在多任务中有效提升了VLM推理性能。

AI 中文摘要

当前视觉语言模型(VLM)在处理需要一致且细粒度推理的复杂视觉任务时常常遇到困难。近期方法旨在训练模型促进自我反思推理,但需要大量标注数据且在测试时缺乏明确反思行为。本文从神经科学获得灵感来弥合这一差距。人类大脑具有高效的反向预测能力。研究先验证主流VLM能进行类似人类大脑的反向预测,接着提出无监督自我反思框架BUS,它能在无真实标签数据上进行反向预测并提供学习信号,消除对标注数据的依赖,提升推理性能,还与多种微调方法兼容。最后在8个基准测试上的实验证明了BUS在复杂视觉任务中的有效性。

英文摘要

Current Vision-Language Models (VLMs) often struggle to handle complex visual tasks that require consistent and fine-grained reasoning. Recent methods aim to train models to facilitate self-reflective reasoning, i.e., reviewing and improving the generated reasoning. However, they require large volumes of annotated data and lack explicit reflective behavior during test time. By contrast, humans perform explicit and efficient self-reflection through mechanisms such as backward prediction, i.e., predicting which current states are likely to precede a given future state. Inspired by neuroscience, this work proposes a novel solution to address these challenges. We first observe and investigate the phenomenon that mainstream VLMs can perform backward prediction, similar to the human brain. A label-free training framework named Brain-inspired Unsupervised Self-reflection (BUS) is proposed to leverage and exploit backward prediction capability to enhance reflective reasoning in complex visual tasks. BUS enables self-verification of reflective reasoning based on backward prediction, providing explicit learning signals under unsupervised conditions. In this way, BUS eliminates reliance on annotated data while improving reasoning performance. Designed as a model-agnostic plug-in, our framework is compatible with popular fine-tuning methods, such as Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL). Initialized from Qwen3-VL-8B, it improves HR-Bench-8K (+8.0%), HR-Bench-4K (+7.7%), V* Bench (+6.3%), and MME-RealWorld-Lite (+5.8%), proving backward prediction is key to advancing reflective reasoning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑