发表机构
Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对多参考图像生成基准的缺陷,构建TRACE-Bench基准,采用算子分解方法评估9个主流模型,发现解耦与属性绑定是核心瓶颈,最优模型属性保真度仅0.74
AI 中文摘要
尽管近期面向多参考图像生成的统一多模态模型取得了进展,但现有基准仍围绕预定义任务类型(如“主体合成”)组织,这类设置不适用于该组合场景,导致覆盖范围碎片化、复杂度不受控,且几乎无诊断价值。鉴于不同多参考任务共享一组通用原子操作,我们采用面向能力的视角,形式化定义了四个算子:锚定($f$)、解耦($g$)、应用($\bigoplus$)和合成($C$)。任意多参考提示均可表示为这些算子的组合公式,其结构复杂度由算子槽位数量量化。基于该形式化方法,我们构建了TRACE-Bench,包含约1600个评估案例,覆盖1至8个槽位,由631个公式模板和约4000张参考图像构成,涵盖多样艺术风格与现实主体。公式结构直接驱动了面向单能力评分的算子对齐评估协议,以及用于递归故障定位的诊断树分析。对9个主流模型的评估揭示了整体评分无法体现的洞见:主要瓶颈在于解耦($g$)与属性绑定($\bigoplus$),而非场景级合成($C$),即便是最优模型在属性保真度上仅得0.74。项目页面:this https URL
英文摘要
Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., "subject composition"), which are ill-suited to this combinatorial setting and lead to fragmented coverage, uncontrolled complexity, and little diagnostic value. Recognizing that diverse multi-reference tasks share a common set of atomic operations, we adopt a capability-oriented perspective and formalize four operators: Anchor ($f$), Disentangle ($g$), Apply ($\oplus$), and Compose ($C$). Any multi-reference prompt can then be represented as a compositional formula over these operators, whose structural complexity is quantified by the number of operator slots. Building on this formulation, we construct TRACE-Bench, comprising approximately 1,600 evaluation cases across slot counts 1--8, built from 631 formula templates and around 4,000 reference images spanning diverse artistic styles and real-world subjects. The formula structure directly drives an operator-aligned evaluation protocol for per-capability scoring and a diagnostic tree analysis for recursive failure localization. Evaluating 9 leading models reveals insights invisible to holistic scoring: the primary bottleneck lies in disentanglement ($g$) and attribute binding ($\oplus$) rather than scene-level composition ($C$), with even the best model scoring only 0.74 on attribute fidelity. Project page: https://amuseum-whr.github.io/TraceBench
CommentsAccepted to ACM Multimedia 2026 (ACM MM 2026)