全能评判者还是全能偏差?通过平衡、解耦的视角诊断多模态评判者
OmniJudge or OmniBias? Diagnosing Multimodal Judges through Balanced, Decoupled Lenses
浏览论文内容
中文总结 AI 辅助
本研究推出平衡解耦的多模态评判者诊断基准D3-Omni,发现全能评判者存在模态相关维度表现差、漏检失败等系统性盲点,为评估多模态模型提供新视角。
中文摘要 AI 辅助
能够联合评判文本生成图像(T2I)、文本生成视频(T2V)和文本生成语音(TTS)的多模态理解模型正日益被用作“全能评判者”,用于评估和自动标注。然而,现有基准和训练数据往往过度强调正例,并混淆不同的失败模式,因此评判者可能在未识别失败的情况下仍能获得高分,其能力缺口也保持隐藏状态,导致这些模型的评分可靠性尚不明确。受此启发,我们推出D3-Omni,这是一个用于诊断细粒度多模态理解的平衡且解耦的基准,覆盖三个任务的53个正交二元维度(17/22/14)和10671个样本(3526/1998/5147)。我们不重新生成可能在维度间泄露信息的输出,而是固定已验证的完全正例种子,并通过受控的提示重写和原子化、维度隔离的扰动推导负例。所得的D3设计具有三重特性:双平衡,有助于缓解负例稀缺和每个维度的标签不平衡;解耦,使每个错误可归因于单一能力;动态,引导构建朝向标签分布中代表性不足的区域,生成模型套件实现了接近1:1的每个维度奇偶性和所有总分级别的均匀分布。在这种平衡视角下,即使是强大的全能评判者也往往在模态相关维度上表现挣扎,确认满足需求的可靠性远高于检测违反需求的可靠性,且将名义上不同的属性视为大致单一决策,表明聚合准确率可能隐藏系统性盲点,而平衡和解耦的视角有助于揭示这些盲点并进而解决它们。
英文摘要
Multimodal understanding models that can jointly judge text-to-image (T2I), text-to-video (T2V) and text-to-speech (TTS) generation are increasingly used as "OmniJudges" for evaluation and automatic annotation. How reliably they understand what they score remains unclear, since existing benchmarks and training data tend to overemphasize positive examples and to conflate distinct failure modes, so a judge may score well without recognizing failures while its capability gaps stay hidden. Motivated by this, we introduce D3-Omni, a balanced and decoupled benchmark for diagnosing fine-grained multimodal understanding, covering 53 orthogonal binary dimensions (17/22/14) and 10,671 samples (3,526/1,998/5,147) across the three tasks. Rather than re-generating outputs, which may leak information across dimensions, we fix verified fully positive seeds and derive negatives through controlled prompt rewriting and atomic, dimension-isolating perturbations. The resulting D3 design is Dual-balanced, which helps alleviate negative-sample scarcity and per-dimension label imbalance; Decoupled, so that each error is attributable to a single capability; and Dynamic, steering construction toward under-represented regions of the label distribution as generative models improve.The suite reaches near 1:1 per-dimension parity and a uniform distribution over all total-score levels. Under this balanced view, even strong OmniJudges tend to struggle on modality-related dimensions, to confirm satisfied requirements far more reliably than they detect violated ones, and to treat nominally distinct attributes as largely a single decision, suggesting that aggregate accuracy may hide systematic blind spots that a balanced and decoupled lens can help expose and, in turn, address.