arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Flow-of-Thought:一种用于视觉推理的框架

Flow-of-Thought: A Framework for Visual Reasoning

Mariia Baidachna, Nicolas Pugeault

arXiv 2610.09746首次发表:更新:

发表机构

University of Glasgow(格拉斯哥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出Flow-of-Thought框架,通过生成视觉草图作为中间推理步骤模拟心理意象,在空间推理任务上取得高准确率,并验证了连续视觉轨迹作为可解释表示的有效性。

AI 中文摘要

心理意象,即“用心灵之眼观看”,是人类认知的一个重要方面。尽管大型语言模型(LLMs)和视觉变换器(ViTs)取得了快速进展,但在需要空间理解的任务上仍然表现不佳。为了解决这一问题,我们引入了Flow-of-Thought(FoT),一个将视觉草图生成作为中间推理步骤进行整合的框架,模拟人类的心理意象。我们在$SO(2)$群轨道和累积最短路径上训练坐标感知的轨迹流场,然后冻结学习到的动态;相同与不同决策通过使用前景加权重建能量来比较竞争性的生成假设。在锁定测试中,FoT在俄罗斯方块上达到100.0%的准确率,在彩色形状上达到99.0%。在冻结迁移下,轨道训练的2D流在BLINK多视图任务上优于其仅端点控制的对照(在133个公共验证对上为72.2%对比63.9%),支持连续视觉轨迹作为空间推理在某些分布外设置中的有效且可解释的表示。

英文摘要

Mental imagery, ``seeing with the mind's eye'' is an essential aspect of human cognition. Despite rapid progress Large Language Models (LLMs) and Vision Transformers (ViTs) still underperform on tasks requiring spatial understanding. To address this, we introduce Flow-of-Thought (FoT), a framework that integrates the generation of visual sketches as intermediate reasoning steps, mimicking mental imagery in humans. We train coordinate-aware trajectory flow fields on $SO(2)$ group orbits and cumulative shortest paths, then freeze the learned dynamics; same vs. different decisions compare competing generative hypotheses using foreground-weighted reconstruction energy. On locked tests FoT reaches 100.0% accuracy on Tetris and 99.0% on colored shapes. Under frozen transfer, the orbit-trained 2D flow improves over its endpoint-only control on BLINK Multi-view (72.2% vs. 63.9% on 133 public validation pairs), supporting continuous visual traces as an effective and interpretable representation for spatial reasoning in some out-of-distribution settings.

CommentsNeurIPS WiML 2026 version: OpenReview version: https://openreview.net/forum?id=QBcqVOacYO

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑