arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MedBenchAgent:迈向医学视觉语言模型(VLM)基准构建的系统化自动化

MedBenchAgent: Towards Systematic Automation of Medical VLM Benchmark Construction

Yulin Fu, Junren Wang, Guangjing Yang, Zhangyuan Yu, Wanran Sun, Jiabao Zhou, Jin Yin, Qicheng Lao

arXiv 2610.11312首次发表:更新:

发表机构

Beijing University of Posts and Telecommunications; West China Hospital(北京邮电大学; 华西医院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出多智能体框架MedBenchAgent,以约束编译方式自动推导医学VLM基准规范,其任务空间F1达90.9%,优于现有方法,可实现医学VLM基准的系统化自动化构建。

AI 中文摘要

借助丰富标注的影像数据集与大型语言模型(LLM),大规模构建医学视觉语言模型(VLM)基准的可行性日益提升,但现有自动化工作主要聚焦于在预定义的基准规范内生成评估项。本文研究更广泛的问题:如何自动推导基准规范本身——即评估什么、哪些标注支持每项任务,以及如何将这些证据转化为可靠的评估项。我们将基准构建表述为约束编译过程,其中基准规范从评估需求、异构标注及医学知识中逐步推导得出。基于此表述,我们提出MedBenchAgent,这是一个多智能体框架,带有基准中间表示(BIR),可编码各构建阶段的任务定义、证据映射、评估协议及项规范。MedBenchAgent将规划(推导并验证规范)与实例化(在锁定的规范下构建并审核项)相分离。MedBenchAgent的任务空间F1值达90.9%,优于直接任务归纳(79.2-80.0%)和先验引导的归纳(85.1%);在正确识别的任务中,抽样的1000个项里有994个通过人工审核。我们进一步展示其对特定医学领域的可移植性,并评估12个VLM,揭示了被聚合分数掩盖的任务与设置特定的差异。这些结果确立了约束编译作为一种可扩展且可审核的框架,适用于超越问题生成的医学VLM基准构建。

英文摘要

Large-scale construction of medical vision-language model (VLM) benchmarks is increasingly feasible with richly annotated imaging datasets and large language models (LLMs), yet existing automation largely focuses on generating evaluation items within predefined benchmark specifications. We study the broader problem of automatically deriving the specification itself: what to evaluate, which annotations support each task, and how to translate this evidence into reliable evaluation items. We formulate benchmark construction as constrained compilation, in which the benchmark specification is progressively derived from evaluation requirements, heterogeneous annotations, and medical knowledge. Based on this formulation, we introduce MedBenchAgent, a multi-agent framework with a Benchmark Intermediate Representation (BIR) that encodes task definitions, evidence mappings, evaluation protocols, and item specifications across construction stages. MedBenchAgent separates planning, which derives and verifies the specification, from instantiation, which constructs and audits items under the locked specification. MedBenchAgent achieves a Task-Space F1 of 90.9%, outperforming direct task induction (79.2-80.0%) and prior-guided induction (85.1%); 994 of 1,000 sampled items from correctly identified tasks pass human audit. We further demonstrate portability to a specialized medical domain and evaluate twelve VLMs, revealing task- and setting-specific variation obscured by aggregate scores. These results establish constrained compilation as a scalable and auditable framework for medical VLM benchmark construction beyond question generation.

Comments25 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑