Graph-of-Mark: Promote Spatial Reasoning in Multimodal Language Models with Graph-Based Visual Prompting
图标记:通过基于图的视觉提示提升多模态语言模型的空间推理能力
机构 * University of Bologna(博洛尼亚大学)
AI总结 本文提出Graph-of-Mark,一种基于图的视觉提示方法,通过在输入图像上叠加场景图来增强多模态语言模型的空间推理能力,实验表明其在视觉问答和定位任务中提升了11个百分点的准确率。
Comments Please cite the definitive, copyrighted, and peer-reviewed version of this article published in AAAI 2026, edited by Sven Koenig et al., AAAI Press, Vol. 40, No. 36, Technical Track, pp. 30726-30734, 2026. DOI: https://doi.org/10.1609/aaai.v40i36.40329