arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14473cs.CV

PuzzleMate:面向以自我为中心视角的拼图辅助的多模态大语言模型基准测试

PuzzleMate: Benchmarking MLLMs for Egocentric Puzzle Assistance

  • IIIT Hyderabad(海得拉巴国际信息技术学院)
  • IIT BHU(印度理工学院(巴纳拉斯印度教大学校区))
  • IISER Thiruvananthapuram(蒂鲁文南特布勒姆印度科学教育与研究所)
  • Univ. Grenoble Alpes, Inria, CNRS, Grenoble INP, LJK(格勒诺布尔阿尔卑斯大学,法国国家信息与自动化研究所,法国国家科学研究中心,格勒诺布尔国立理工学院,让·库尔实验室)

机构由 AI 辅助整理,请以论文原文为准。

Avijit Dasgupta, Shayon Dasgupta, Zakaria Laskar, C. V. Jawahar, Karteek Alahari

AI总结:

本文提出PuzzleMate框架和基准,通过用户研究评估多模态大语言模型在自我中心拼图辅助中的顺序推理能力,发现关键瓶颈及性能差距。

AI中文摘要:

个人AI助手有潜力从数字界面发展为具身伴侣,能够引导用户完成复杂的物理活动。为了使这些助手成为日常生活中不可或缺的一部分,它们必须不仅仅识别物体;它们必须提供与用户实时进度相一致的精确、逐步的指令。尽管多模态大语言模型(MLLMs)在通用视觉理解方面显示出潜力,但它们在细粒度操作任务中提供基于上下文的、顺序性指导的能力在很大程度上仍未得到验证。在本文中,我们选择拼图作为这一能力的战略性测试平台。与通用物体识别不同,拼图解决要求高精度的空间推理、区分微小几何变化的能力以及对顺序逻辑的严格遵循。我们通过PuzzleMate(一个专注于通过自我中心视角捕捉的拼图解决的新框架)来研究这一能力。我们在用户参与的研究中部署了PuzzleMate,以评估最先进的MLLMs如何感知当前拼图状态并生成可操作的下一步指令。我们的分析揭示了限制其有效性的七个关键瓶颈。基于这些见解,我们提出了一个基准测试,能够系统地评估MLLMs在拼图解决中的推理能力。我们的研究结果揭示了当前模型如GPT-5.2和Gemini-2.5-Pro存在的显著性能差距;尽管这些MLLMs能力很强,但它们在应对拼图辅助所必需的复杂推理和顺序逻辑方面存在困难。

英文摘要:

Personal AI assistants hold the potential to evolve from digital interfaces into embodied companions capable of guiding users through complex physical activities. For these assistants to become integral to daily life, they must do more than identify objects; they must provide precise, step-by-step instructions that align with a user's real-time progress. While Multimodal Large Language Models (MLLMs) show promise in general visual understanding, their ability to deliver grounded, sequential guidance for fine-grained manipulation tasks remains largely unverified. In this paper, we choose the jigsaw puzzle as a strategic testbed for this capability. Unlike general object recognition, puzzle solving demands high-precision spatial reasoning, the ability to distinguish between minute geometric variations, and a rigorous adherence to sequential logic. We investigate this capability through PuzzleMate, a novel framework focused on jigsaw puzzle solving captured through an egocentric viewpoint. We deploy PuzzleMate in a user-in-the-loop study to evaluate how well state-of-the-art MLLMs perceive the current puzzle state and generate actionable next-step instructions. Our analysis reveals seven key bottlenecks that limit their effectiveness. Building on these insights, we propose a benchmark that enables systematic evaluation of MLLMs' reasoning capabilities for puzzle solving. Our findings reveal a substantial performance gap in current models like GPT-5.2 and Gemini-2.5-Pro; while these MLLMs are highly capable, they struggle to navigate the intricate reasoning and sequential logic essential for jigsaw puzzle assistance.

↑