arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21946cs.LG

用于SeePhys Pro的多智能体辩论与视觉信息提取:ICML 2026 AI4Math赛道3挑战赛的第一名技术报告

Multi-Agent Debate and Visual Information Extraction for SeePhys Pro: A 1st-Place Technical Report from ICML 2026 AI4Math Track 3 Challenge

Jiseok Kwak, Suhyeon Jo, Taewoo Kim, Yeongmin Kim, Byeonghu Na, Il-chul Moon

首次发表
浏览论文内容

中文总结 AI 辅助

针对回答含图像的大学物理问题的SeePhys Pro任务,提出两阶段框架,包括视觉信息提取和多智能体辩论推理阶段,提高了准确率,在挑战赛中获第一名,分析得出编排收益及图形辅助价值与问题图像占比相关的结论。

中文摘要 AI 辅助

本技术报告介绍了我们在第三届数学人工智能研讨会上对挑战赛赛道3:SeePhys Pro的方法,该任务是回答大学水平的物理问题,其陈述和图形可能部分或全部以图像形式给出。当决定性信息存在于图形而非文本中时,视觉物理问题对大语言模型来说变得更加困难,并且随着更多问题转移到图像中,这种模态差距会扩大。我们用两阶段框架解决该任务:一个视觉信息提取阶段,将图形内容重新表达为求解器可读文本以弥合模态差距;一个推理阶段,通过多智能体辩论协调三个异构求解器。我们的分析得出两个发现:编排的收益来自可靠的答案选择而非额外的辩论,图形辅助的价值与问题锁定在图像中的程度成比例。由此产生的管道在公共分割上将单智能体基线的整体准确率从0.643提高到0.802,并在公共和私有排行榜上均获得第一名(私有整体0.743)。

英文摘要

This technical report presents our approach to Challenge Track~3: SeePhys Pro at the 3rd AI for Math Workshop, where the task is to answer college-level physics questions whose statement and figure may be given partly or entirely as an image. Visual physics problems become substantially harder for large language models when the decisive information resides in a figure rather than in the text, and this modality gap widens as more of the problem migrates into the image. We address the task with a two-stage framework: a visual information extraction stage that re-expresses figure content as solver-readable text to close the modality gap, and a reasoning stage that orchestrates three heterogeneous solvers through multi-agent debate. Our analysis yields two findings: the gain from orchestration comes from reliable answer selection rather than from additional debate, and the value of a figure aid scales with how much of the problem is locked inside the image. The resulting pipeline improves overall accuracy over a single-agent baseline from 0.643 to 0.802 on the public split, and won 1st place on both the public and the private leaderboard (private overall 0.743).

↑