See, Symbolize, Act: Grounding VLMs with Spatial Representations for Better Gameplay
看见、符号化、行动:通过空间表示接地VLMs以实现更好的游戏表现
AI总结 本文研究了通过提供视觉帧和场景符号表示来提升VLMs在交互环境中的表现,发现准确的符号信息能提升性能,但自身提取符号的模型性能受模型能力和场景复杂度影响。
Comments 11 pages, 13 figures. Accepted to LMReasoning Workshop at AAAI 2026