Falcon Perception-HD:基于强化学习的高密度感知
Falcon Perception-HD: High Density Perception via Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
该研究针对开放词汇自回归感知模型的目标错位问题,基于Falcon Perception设计RL框架,采用GRPO优化,在指代表达分割任务上提升了极密集场景性能,减少冗余需求并保留物体存在性知识。
中文摘要 AI 辅助
针对开放词汇场景下定位视觉实体的自回归感知模型,大多采用带最大似然的监督微调(SFT)进行训练,但该方法优化的代理目标(逐标记交叉熵)与精度、召回率等感知指标存在根本错位。本文探索后训练强化学习(RL),具体为GRPO,以直接对齐模型与评估指标。基于最新推出的Falcon Perception,我们设计了一个解决感知特定挑战的RL框架:针对集合结构输出的奖励设计与多头采样控制。我们发现RL对感知的多项优势:首先,RL在极密集场景(每场景多达500个物体)中实现了SOTA性能,而现有多数系统在此场景会急剧退化或崩溃;此外,它解决了自回归感知模型的常见问题,如掩码重复,并几乎完全消除了对NMS和坐标去重的需求,这既提升了性能与效率,又无需调整超参数;总体而言,我们在指代表达分割的所有难度级别(在PBench和SACO-Gold上评估)均观察到性能提升,且找到一种优雅的方法来保留物体是否存在的知识(通过MCC评估),无需在负样本上训练。我们证明,仅用惩罚假阴性和假阳性的简单奖励就足够。我们开发了两种混合自标注流水线,分别针对困难指代表达和极密集场景定制,并展示了它们对RL训练的益处。模型权重作为Falcon Perception的修订版发布,数据集将公开。
英文摘要
Autoregressive perception models trained to localize visual entities under the open-vocabulary setting are mostly trained using Supervised fine-tuning (SFT) with maximum likelihood, yet it optimizes a proxy objective (per-token cross-entropy) that is fundamentally misaligned with perception metrics such as precision and recall. In this paper, we explore post-training reinforcement learning (RL), specifically GRPO, to directly align these models with their evaluation metrics. Building up on the recently introduced Falcon Perception, we design an RL framework that addresses perception-specific challenges: reward design for set-structured outputs and multi-head sampling control. We discover multiple benefits from RL for perception: first, RL unlocks state-of-the-art performance in very dense scenes (up to 500 objects per scene), a regime where most existing systems degrade sharply or collapse; furthermore it fixes common issues in autoregressive perception models like mask repetitions and removes almost entirely the need for NMS and coordinate deduplication, which improve both performance and efficiency and remove the need for hyperparameters tuning; overall, we notice improvements on all levels of difficulties in referring expression segmentation (on PBench and SACO-Gold), and we find an elegant way to preserve the knowledge of whether an object exists or not (as evaluated by MCC) without training on negative samples. We show that a simple reward that penalizes false negatives and positives is sufficient. We develop two hybrid self-annotation pipelines, respectively tailored for difficult referring expressions and very dense scenes, and show their benefits on RL-training. Model weights are released as a Falcon Perception revision~\footnote{https://huggingface.co/tiiuae/Falcon-Perception}. Datasets will be published.
发表机构
- Tübingen AI Center(蒂宾根人工智能中心)
- University of Tübingen(蒂宾根大学)
- MIT-IBM Watson AI Lab(麻省理工学院-IBM沃森人工智能实验室)
- Technology Innovation Institute(技术创新研究院)
机构由 AI 辅助整理,请以论文原文为准。