arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16690cs.CV

AnchorScore:一种基于CLIP的多模态大语言模型(MLLM)标注难度诊断方法

AnchorScore: A CLIP-Based Diagnostic of MLLM Annotation Difficulty

Yan Ma, Lizhuo Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出基于CLIP的AnchorScore,可低成本预先对类别按MLLM标注难度排名,在课堂行为、动作识别等数据上与MLLM准确率相关性高,可用于混合路由、提示消歧等场景。

中文摘要 AI 辅助

多模态大语言模型(MLLM)被广泛用于自动标注,但不同类别的准确率差异极大(例如,三个课堂子数据集的13个类别中,准确率在12%至98%之间),且测量成本高昂:对一个27B参数的MLLM在5416张验证图像上进行评估大约需要14小时,而用冻结的CLIP模型对相同图像进行处理仅需约3分钟。目前,针对如何预先获取一种低成本信号以对类别按预期的MLLM标注难度进行排名的研究仍不充分。基于配套研究中提出的AnchorProxy结构(即每类的零样本CLIP准确率),本文系统评估了其全框架形式(在此处称为AnchorScore),将其作为一种预先诊断工具,用于标记MLLM最不可能可靠标注的类别。在课堂行为数据(SCB5,13个类别,6个MLLM)上,AnchorScore与每类MLLM准确率的Spearman相关系数为0.769,p值为0.002,样本量n=13。在n=13的情况下,其他替代难度预测因子(DINOv2、ResNet-50、SigLIP或MLLM自表述不确定性)均未显示出显著的类别级相关性。跨模型一致性控制实验表明,AnchorScore主要捕获的是共享的类别难度因子,而非特定于CLIP的信号。在Stanford40 Actions上的独立复现实验产生了几乎相同的效果(相关系数rho=0.817,p<0.001);该关联在动作识别数据上最强,在医学和卫星图像上则有所减弱。本文提出了三个实际应用方向:可部署的CLIP/MLLM混合路由策略(预测类别路由:相比仅使用CLIP,准确率提升高达23个百分点,同时MLLM计算成本约节省44%)、针对困难类别的提示消歧(探索性)以及用于人工验证的审查优先级预测。AnchorScore不估计MLLM的精确准确率,它提供了一种低成本的排名信号,将昂贵的MLLM评估引导至最具信息价值的类别。

英文摘要

Multimodal large language models (MLLMs) are widely used for automated annotation, yet their per-class accuracy varies widely (e.g., 12%-98% across the 13 classes of three classroom sub-datasets) and is expensive to measure: evaluating one 27B MLLM on 5,416 validation images takes roughly 14 hours, whereas a frozen-CLIP pass over the same images completes in about 3 minutes. A low-cost signal for ranking classes by expected MLLM annotation difficulty a priori remains underexplored. Building on the AnchorProxy construct (per-class zero-shot CLIP accuracy) introduced in the companion study, this paper systematically evaluates its full-frame formulation, termed AnchorScore here, as an a priori diagnostic that flags the classes MLLMs are least likely to annotate reliably. On classroom behavior data (SCB5, 13 classes, 6 MLLMs), AnchorScore correlates with per-class MLLM accuracy (Spearman rho = 0.769, p = 0.002, n = 13). None of the alternative difficulty predictors (DINOv2, ResNet-50, SigLIP, or MLLM self-verbalized uncertainty) showed a significant class-level correlation at n = 13. A cross-model consensus control suggests AnchorScore primarily captures a shared class-difficulty factor rather than a CLIP-specific signal. An independent replication on Stanford40 Actions yields a nearly identical effect (rho = 0.817, p < 0.001); the association is strongest on activity-recognition data and attenuates on medical and satellite imagery. Three practical applications follow: a deployable hybrid CLIP/MLLM routing strategy (predicted-class routing: up to +23 pp over CLIP-only at roughly 44% fewer MLLM calls), prompt disambiguation on hard classes (exploratory), and review-priority prediction for human verification. AnchorScore does not estimate exact MLLM accuracy; it provides a low-cost ranking signal that directs expensive MLLM evaluation to the classes where it is most informative.

发表机构

  • School of Foreign Studies, Changsha University of Science and Technology(长沙理工大学外国语学院)
  • School of Education, Hunan Agricultural University(湖南农业大学教育学院)
  • School of Information and Intelligence, Hunan Agricultural University(湖南农业大学信息与智能科学学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑