arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VISTA:交互世界中的视觉推理驾驭框架

VISTA: A Visual Harness for Reasoning in an Interactive World

Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu, Kaiming He

arXiv 2610.02200首次发表:更新:

发表机构

Massachusetts Institute of Technology(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

VISTA是一种视觉驾驭框架,通过无损视觉记忆和主动检索,释放多模态模型在交互环境中的推理潜力,在ARC-AGI-3上达到满分并显著超越基线。

AI 中文摘要

我们证明多模态模型具备强大的推理能力,且适当的驾驭框架能够释放其潜力,使其解决各种交互环境中的任务。我们引入VISTA,一种视觉驾驭框架,为通用多模态模型提供长时程视觉能力。VISTA使模型能够通过视觉观察直接感知环境,并维护一种无损视觉记忆,以原始形式保留过往观察。模型能够主动检索这些观察,并在推理过程中重新组织其视觉输入。在ARC-AGI-3基准上,VISTA将Claude Opus 5.0的相对人类行动效率得分从40.68提升至完美的100.00,且模型使用比首次参与的人类参与者少57.4%的动作完成了全部25个公开游戏。VISTA的简洁设计还使其能够以最小化适配自然扩展到多样化的视觉环境。在覆盖多种视觉游戏和谜题的三个额外基准上,它显著优于使用相同底层模型但仅配备最小化驾驭框架的基线。我们的结果凸显了VISTA作为通用视觉驾驭框架在复杂视觉环境中推进多模态智能体的潜力。

英文摘要

We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows the model to directly perceive the environment through visual observations and maintains a lossless visual memory that preserves past observations in their original form. The model can actively retrieve these observations and reorganize its visual input as it reasons. On ARC-AGI-3, VISTA improves Claude Opus 5.0's Relative Human Action Efficiency score from 40.68 to a perfect 100.00, with the model completing all 25 public games using 57.4% fewer actions than first-time human participants. VISTA's simple design also allows it to extend naturally to diverse visual environments with minimal adaptation. Across three additional benchmarks covering a diverse range of visual games and puzzles, it substantially outperforms baselines using the same underlying model with minimal harnesses. Our results highlight VISTA's potential as a general-purpose visual harness for advancing multimodal agents in complex visual environments.

CommentsTech report. An early version of this manuscript was in a blogpost published in Aug 5, 2026: https://vista-research.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑