检测图形设计中的一切:用于自回归检测的元素级奖励
Detect Anything in Graphic Design: Element-Level Rewards for Autoregressive Detection
查看机构详情
- BNRist, Tsinghua University(清华大学BNRist)
- Canva Research(Canva研究院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
提出DAD模型,将图形设计检测视为按组合顺序解构,结合失模检测与元素级优化方法EleRPO,在千万级数据集上达到人类水平,超越基线。
中文摘要 AI 辅助
图形设计,如海报、广告和信息图表,是传达信息和塑造理解的重要媒介。与自然图像不同,它们由具有明确组合顺序的分层元素组成。然而,现有的目标检测模型将这些元素视为无序集合,未利用组合顺序。为解决这一局限,我们提出了“检测图形设计中的一切”(DAD),一种将图形设计检测表述为组合解构的模型。它按组合顺序解码元素,利用较低层元素更好地检测较高层元素。DAD的关键特性是失模检测(amodal detection),即预测每个元素的完整边界框,包括被其上方元素遮挡的区域。基于这一表述,我们提出了元素相对策略优化(EleRPO),将GRPO从序列级监督扩展到元素级优化。EleRPO提供细粒度的训练信号,捕捉每个检测元素对整体检测质量的贡献,并与组合顺序协同作用以提高检测性能。为支持训练和评估,我们构建了一个包含1000万个图形设计的数据集。实验表明,DAD在所有基线上表现优异,并在失模检测中达到人类水平,支持有效的图像到图层分解。EleRPO在九个检测基准上持续优于GRPO。
英文摘要
Graphic designs, such as posters, advertisements, and infographics, are an important medium for communicating information and shaping understanding. Unlike natural images, they consist of layered elements with explicit compositional order. However, existing object detection models treat these elements as an unordered set, leaving compositional order unexploited. To address this limitation, we present Detect Anything in Graphic Design (DAD), a model that formulates graphic design detection as compositional deconstruction. It decodes elements in compositional order, using lower-layer elements to better detect higher-layer ones. The key feature of DAD is amodal detection, which predicts the full bounding box of each element, including regions occluded by elements placed above it. Building on this formulation, we propose Element Relative Policy Optimization (EleRPO), which extends GRPO from sequence-level supervision to element-level optimization. EleRPO provides fine-grained training signals that capture how each detected element contributes to overall detection quality, and works synergistically with compositional order to improve detection performance. To support training and evaluation, we build a dataset of 10 million graphic designs. Experiments show that DAD outperforms all baselines and achieves human-level performance in amodal detection, supporting effective image-to-layer decomposition. EleRPO consistently improves over GRPO across nine detection benchmarks.