arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.08106cs.AI

ChartBmkAgent:基于稀疏错误分类规范的管控多智能体图表问答基准构建

ChartBmkAgent: Harness-Governed Multi-Agent Construction of Chart QA Benchmarks from Sparse Error-Taxonomy Specifications

Langxi Huang, Pingping Zhang, Lanyun Zhu, Chunyang Jiang, Jiawei Shao, Haocheng Yuan, Peilin Chen

首次发表
浏览论文内容

中文总结 AI 辅助

ChartBmkAgent通过管控多智能体从稀疏错误分类规范构建图表问答样本,将能力差距转化为针对性诊断证据,实验表明样本能揭示模型能力差异并保持目标对齐。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)发展迅速,而传统基准开发滞后,延迟了对新观察到的能力差距的调研。此类调研需要表达丰富的任务格式和按需构建流程:信息丰富的图表使图表问答(Chart QA)适用于探测耦合的感知与推理。自动化图表问答构建旨在通过将已识别的差距按需转化为针对性样本来缩短基准开发周期。然而,当前方法通常将目标引导与从零生成分离:目标引导系统往往需要准备好的数据、图表或模板,而从零生成系统主要确保工件有效性,而未明确控制新合成的需求与内容是否与外部指定的诊断目标保持一致。我们引入了ChartBmkAgent,它通过从稀疏错误分类规范构建完整的图表问答样本,将已识别的能力差距转化为有针对性的诊断证据。在整个构建过程中,中央管控机制管理专业智能体,要求提供与原始错误类别对齐的各阶段证据,并记录每次接受决策的依据。在300个覆盖全分类的样本上,多模态大语言模型的准确率从32.7%到84.3%不等,且具有不同的类别特征,表明生成的样本能揭示能力差异。在三个源模型比较中,针对性后续测试得分为50.0%,而匹配对照组为82.2%(p=8.96×10⁻⁶);所有六项跨模型比较方向一致,证明了针对性验证和诊断数据生成的有效性。多个评估模型评估了每个样本是否测试了其指定的错误类别;86.4%的样本符合该标准,为目标保持提供了经验证据。

英文摘要

Multimodal large language models (MLLMs) advance rapidly, while conventional benchmark development lags behind, delaying investigation of newly observed capability gaps. Such investigation requires an expressive task format and an on-demand construction process: information-rich charts make chart question answering (Chart QA) suitable for probing coupled perception and reasoning. Automated Chart QA construction is intended to shorten the benchmark-development cycle by turning identified gaps into targeted samples on demand. Current methods, however, commonly separate target guidance from scratch generation: target-guided systems often require prepared data, charts, or templates, while scratch-generation systems primarily ensure artifact validity, without explicitly controlling whether newly synthesized requirements and content remain aligned with an externally specified diagnostic target. We introduce ChartBmkAgent, which turns an identified capability gap into targeted diagnostic evidence by constructing complete Chart QA samples from sparse error-taxonomy specifications. Throughout construction, a central harness governs specialized agents, requires stage-specific evidence of alignment with the original error category, and records the basis for each acceptance decision. On 300 taxonomy-wide samples, MLLM accuracies ranged from 32.7% to 84.3% with distinct category profiles, showing that generated samples reveal capability differences. Across three source-model comparisons, targeted follow-ups scored 50.0% versus 82.2% on matched controls ($p=8.96\times10^{-6}$); all six cross-model comparisons had the same direction, demonstrating targeted validation and diagnostic-data generation. Multiple evaluator models assessed whether each sample tested its specified error category; 86.4% met this criterion, providing empirical evidence of target preservation.

发表机构

  • City University of Hong Kong(香港城市大学)
  • Tongji University(同济大学)
  • Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑