AI 中文总结
RelateAnything是一个53M参数的开放词汇关系预测模型,接受任意区域输入和推理时提供的谓词词汇表,无需对象标签,在跨数据集基准上平均召回率提升2.3-3.5倍,并构建了RA-4M语料库和OV-SGG-Bench基准。
AI 中文摘要
开放词汇检测在推理时接受任意类别列表,可提示分割返回不带类别名称的区域:分类体系已离开模型而成为输入。关系预测却并非如此。场景图模型仍在一个标注风格的50或56个谓词上进行训练和评估,其关系头以对象标签为条件,因此与一个检测器绑定。三个障碍解释了这一点,且主要都不是建模问题:没有关系语料库既是自由文本又经过验证;标签条件架构无法接受其未训练过的词汇表;标准指标奖励与训练语料库的一致性,因此更大的词汇表被视为回归。我们提出RelateAnything,一个53M参数模型,接受图像和来自任何来源的区域,返回在推理时以字符串形式提供的谓词词汇表上的评分关系。对象标签从不作为输入,因此区域来源可以在不重新训练的情况下更改,词汇表是文本嵌入的库,而非学习的分类器。它以20毫秒/帧的速度运行。在19,103个谓词上训练需要正-未标记监督和能区分反义词的文本编码器,对比编码器将反义词嵌入在余弦0.95处。为提供监督,我们构建了RA-4M,包含474k图像和4.3M个关系,覆盖10,102个自由文本谓词,针对编号框标记生成并经几何验证。为衡量它,我们构建了OV-SGG-Bench,六个轴在标准召回先验无法满足的数据集上评分。在三个跨数据集基准和第四个零样本基准上,RelateAnything的平均召回率是同等规模最强开放词汇方法的2.3-3.5倍,这些优势在真实检测器下依然存在,并以不到其参数2%的规模在两项指标上领先3B-VLM场景图模型。域内测量高估了迁移增益约5倍。模型、语料库和基准均已公开。
英文摘要
Open-vocabulary detection accepts any class list at inference, and promptable segmentation returns regions without class names: the taxonomy has left the model and become an input. Relation prediction has not. Scene-graph models are still trained and evaluated on the 50 or 56 predicates of one annotation style, their relation head conditioned on object labels and so tied to one detector. Three obstacles explain this, none primarily modelling: no relation corpus is both free-text and verified, a label-conditioned architecture cannot accept a vocabulary it was not trained on, and the standard metric rewards agreement with the training corpus, so a larger vocabulary scores as a regression. We present RelateAnything, a 53M-parameter model taking an image and regions from any source and returning scored relations over a predicate vocabulary supplied at inference as strings. Object labels are never an input, so the region source can change without retraining, and the vocabulary is a bank of text embeddings, not a learned classifier. It runs at 20 ms/frame. Training over 19,103 predicates requires positive-unlabeled supervision and a text encoder that separates antonyms, which contrastive encoders embed at cosine 0.95. To supply the supervision we build RA-4M, 474k images and 4.3M relations over 10,102 free-text predicates, generated against numbered box markers and geometrically verified. To measure it we build OV-SGG-Bench, six axes scored across datasets that the priors standard recall rewards cannot satisfy. On three cross-dataset benchmarks and a fourth zero-shot, RelateAnything has 2.3-3.5x the mean recall of the strongest open-vocabulary method of comparable scale, margins that survive a real detector, and leads a 3B-VLM scene-graph model on both metrics at under 2% of its parameters. In-domain measurement overstates transfer gains ~5x. Model, corpus and benchmark are public.