arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

期刊&会议

AAAI Conference on Artificial Intelligence · 会议 · Artificial Intelligence

2026-03-27 至 2026-03-27 共收录 2
2603.06663 2026-03-27 cs.CV cs.AI

Graph-of-Mark: Promote Spatial Reasoning in Multimodal Language Models with Graph-Based Visual Prompting

图标记:通过基于图的视觉提示提升多模态语言模型的空间推理能力

Giacomo Frisoni, Lorenzo Molfetta, Mattia Buzzoni, Gianluca Moro

机构 * University of Bologna(博洛尼亚大学)

AI总结 本文提出Graph-of-Mark,一种基于图的视觉提示方法,通过在输入图像上叠加场景图来增强多模态语言模型的空间推理能力,实验表明其在视觉问答和定位任务中提升了11个百分点的准确率。

Comments Please cite the definitive, copyrighted, and peer-reviewed version of this article published in AAAI 2026, edited by Sven Koenig et al., AAAI Press, Vol. 40, No. 36, Technical Track, pp. 30726-30734, 2026. DOI: https://doi.org/10.1609/aaai.v40i36.40329

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12760 2026-03-27 cs.MM cs.AI cs.CV

A User-Friendly Framework for Generating Model-Preferred Prompts in Text-to-Image Synthesis

一种用户友好的生成模型偏好提示框架

Nailei Hei, Qianyu Guo, Zihao Wang, Yan Wang, Haofen Wang, Wenqiang Zhang

AI总结 本文提出UF-FGTG框架,通过粗细粒度提示数据集和自适应特征提取模块,自动优化用户输入提示以生成更高质量的图像。

Comments Accepted by The 38th Annual AAAI Conference on Artificial Intelligence (AAAI 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏