arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19355cs.MMcs.CV

GRACE:基于适配器组合与证据感知校准的教育视觉问答的接地推理

GRACE: Grounded Reasoning via Adapter Composition and Evidence-Aware Calibration for Educational Visual Question Answering

Xinjin Li, Yudi Xia, Xi Zhao, Yiliu Xu, Yining Liu, Cheng Lu, Yujian Long, Yu Ma, Jinghan Cao, Liang Fan, Yeyun Xu

首次发表
浏览论文内容

中文总结 AI 辅助

针对教育视觉问答的问题-选项捷径问题,提出GRACE框架,利用结构化教育状态实现参数高效多模态适配,在ScienceQA上提升了多项准确率。

中文摘要 AI 辅助

教育视觉问答(VQA)要求模型结合语言与视觉证据解决面向课程的多项选择题。与传统开放式VQA相比,教育类示例常包含结构化评估元数据、图表或图像上下文,以及语义相近的选项,这为问题-选项捷径创造了极大机会。我们在该场景下开发并评估了一种针对冻结多模态大语言模型的参数高效适配框架。我们提出GRACE(Ground Reasoning via Adapter Composition and Evidence-Aware Calibration,基于适配器组合与证据感知校准的接地推理),该框架利用每个问题的教学状态来专门化轻量级语言与视觉适配。该状态包含推理可见的主题、分组技能、年级、视觉上下文、问题意图及选项结构线索。GRACE使用因子特定提示与轻量级视觉适配器,随后应用证据感知选项校准,在共享多模态上下文下对所有候选选项打分。在ScienceQA数据集上,GRACE将共享适配器基线的整体准确率从90.5%提升至93.1%,图像上下文问题的准确率从88.7%提升至91.2%。移除教学组合、选项校准或视觉适配器分别会使整体准确率降低1.4、1.0和1.5个百分点。这些受控结果表明,结构化教育状态是参数高效多模态适配的有效路由信号。

英文摘要

Educational visual question answering, or VQA, requires models to solve curriculum-oriented multiple-choice questions using both language and visual evidence. Compared with conventional open-ended VQA, educational examples often include structured assessment metadata, diagrams or image contexts, and semantically close answer options, creating strong opportunities for question-option shortcuts. We develop and evaluate a parameter-efficient adaptation framework for a frozen multimodal large language model in this setting. We introduce GRACE, Grounded Reasoning via Adapter Composition and Evidence-Aware Calibration, a framework that uses the pedagogical state of each question to specialize lightweight language and vision adaptation. The state combines inference-visible subject, grouped skill, grade, visual-context, question-intent, and option-structure cues. GRACE uses factor-specific prompts and lightweight visual adapters, then applies evidence-aware option calibration to score all candidates under a shared multimodal context. On ScienceQA, GRACE improves a shared-adapter baseline from 90.5 percent to 93.1 percent overall accuracy and from 88.7 percent to 91.2 percent on image-context questions. Removing pedagogical composition, option calibration, or the visual adapter reduces overall accuracy by 1.4, 1.0, and 1.5 points, respectively. These controlled results show that structured educational state is an effective routing signal for parameter-efficient multimodal adaptation.

发表机构

  • Columbia University(哥伦比亚大学)
  • Carnegie Mellon University(卡内基梅隆大学)
  • University of California, Berkeley(加州大学伯克利分校)
  • Stevens Institute of Technology(史蒂文斯理工学院)
  • Georgetown University(乔治城大学)
  • San Francisco State University(旧金山州立大学)
  • Loughborough University(拉夫堡大学)
  • Texas A&M University(德克萨斯农工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑