arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过置信度校准和增量推理实现特定任务的多模态问答代理,用于QANTA 2026

Task-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026

Nirjhar Das, Md. Al-Mamun Provath

arXiv 2607.09623首次发表:更新:

发表机构

Department of Computer Science; Engineering, Chittagong University of Engineering \& Technology, Chattogram, Bangladesh(; )

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对QANTA 2026共享挑战,开发特定任务双代理架构,抢答代理用带置信度校准的模型及数值推理策略,加分代理用多种推理整合信息,强调仅托管环境下的高效推理与校准,系统取得高分,证明轻量级策略在资源受限问答中性能强。

AI 中文摘要

我们提交了在ICML 2026高效多模态问答研讨会上的QANTA 2026共享挑战的成果。QANTA评估多模态知识竞赛系统,该系统在现实效率约束下根据增量显示的文本和图像回答金字塔式问题。挑战包括两个不同任务:抢答问题,需在不确定时决定何时回答;加分问题,强调准确答案选择和人工采纳。为实现不同目标,我们开发特定任务的双代理架构。抢答代理使用带有置信度校准回答的GPT-4o-mini-class模型及特定领域数值推理策略减少过度自信预测;加分代理使用带有导入感知推理、结构化关系推理和多模态证据整合的GPT-4o-class模型。我们的方法强调在仅托管环境中的高效推理策略和置信度校准。我们的系统在排行榜上获得最高分0.402,包括抢答分数0.238和加分效应分数0.164。结果表明轻量级、特定任务推理策略能在资源受限的多模态问答基准测试中表现出色。

英文摘要

We present our submission to the QANTA 2026 shared challenge at the ICML 2026 Workshop on Efficient Multimodal Question Answering (EMM-QA). Quanta evaluates multimodal quizbowl systems that answer pyramid-style questions from incrementally revealed text and accompanying images while operating under realistic efficiency constraints. The challenge consists of two distinct tasks: Tossup questions, which require deciding when to answer under uncertainty, and Bonus questions, which emphasize accurate answer selection and human adoption. To address these differing objectives, we develop a task-specific two-agent architecture. Our Tossup agent utilizes a GPT-4o-mini-class model (referred to as GPT-4.1-mini in the competition logs) with confidence-calibrated answering and a domain-specific numeric reasoning policy that reduces overconfident predictions from isolated quantitative clues. Our Bonus agent uses GPT-4o-class model (referred to as GPT-4.1) with leadin-aware reasoning, structured relational reasoning, and multimodal evidence integration to improve exact answer selection. Rather than relying on a retrieval pipeline or model ensembles, our approach emphasizes efficient reasoning policies and confidence calibration within a hosted-only environment. Our system achieved the highest overall leaderboard score of 0.402, including a Tossup score of 0.238 and a Bonus Effect score of 0.164. The results demonstrate that lightweight, task-specific reasoning strategies can provide strong performance on resource-constrained multimodal question answering benchmarks.

Comments10 pages, 1 figure. Accepted at the EMM-QA 2026 Workshop, ICML 2026 (Non-Archival). Rank #1 overall system in the QANTA 2026 Challenge

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑