arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Exemplar2VQA:一种通过多智能体编码的可扩展示例驱动视觉问答生成框架

Exemplar2VQA: A Scalable Exemplar-Driven Visual Question Answering Generation Framework via Multi-Agent Coding

Jiayu Ying, Qijian Tian, Ruijie Xu, Xinnan Zhu, Daoguo Dong, Jiachen Xu, Xin Tan

arXiv 2609.37655首次发表:更新:

发表机构

East China Normal University; Shanghai Jiao Tong University; Fudan University; Tencent(华东师范大学; 上海交通大学; 复旦大学; 腾讯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Exemplar2VQA通过多智能体编码和几何工具库,以示例驱动方式可扩展生成大规模空间QA数据,微调Qwen2.5-VL后显著提升多种基准性能,弥合具身AI模拟到现实差距。

AI 中文摘要

多模态大语言模型(MLLMs)在空间智能方面的进展受到复杂、可扩展的3D问答(QA)数据稀缺的瓶颈制约。虽然手动标注劳动密集,但直接利用LLMs合成这些QA对往往因其在空间和几何计算方面的固有缺陷而失败。我们引入了Exemplar2VQA,一种可扩展的示例驱动视觉问答生成框架,通过多智能体编码在模拟环境中快速合成大规模空间QA对。通过为协作智能体配备精心设计的几何工具库,Exemplar2VQA通过确定性代码执行绕过了LLMs的空间推理缺陷。关键的是,该框架表现出显著的通用性:以多样化的静态对象中心空间查询模板作为示例,它无缝且自主地将它们扩展为大规模、高保真的合成数据集。仅使用Exemplar2VQA生成的合成室内数据对Qwen2.5-VL(3B/7B)进行微调,在各种不同基准上产生了显著的性能提升。此外,其有效性不仅限于域内室内数据集,还稳健地扩展到室外和混合场景基准。这些结果确立了Exemplar2VQA作为弥合具身AI中模拟到现实差距的可扩展且强大的范式。我们的代码位于此https URL。

英文摘要

Advancing spatial intelligence in Multimodal Large Language Models (MLLMs) is bottlenecked by the scarcity of complex, scalable 3D question-answer (QA) data. While manual annotation is labor-intensive, directly utilizing LLMs to synthesize these QA pairs often fails due to their inherent deficiencies in spatial and geometric computation. We introduce Exemplar2VQA, a scalable exemplar-driven visual question answering generation framework that rapidly synthesizes large-scale spatial QA pairs in simulated environments via multi-agent coding. By equipping collaborative agents with a meticulously designed library of geometric utilities, Exemplar2VQA bypasses LLMs' spatial reasoning flaws through deterministic code execution. Crucially, the framework exhibits remarkable versatility: taking diverse static object-centric spatial query templates as exemplars, it seamlessly and autonomously scales them into massive, high-fidelity synthetic datasets. Fine-tuning Qwen2.5-VL (3B/7B) exclusively on Exemplar2VQA-generated synthetic indoor data yields significant performance improvements across various diverse benchmarks. Furthermore, its effectiveness is not limited to in-domain indoor datasets but also robustly extends to outdoor and mixed-scene benchmarks. These results establish Exemplar2VQA as a scalable and powerful paradigm for bridging the sim-to-real gap in Embodied AI. Our code is at https://github.com/yingjiayu12/Exemplar2VQA

CommentsAccepted to NeurIPS 2026. 33 pages, 11 figures, 11 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑