AI 中文总结
针对DETR中查询知识碎片化问题,提出BS-O2G即插即用模块,通过构建预测感知图进行查询协作,在保持一对一匹配下提升检测性能并加速收敛。
AI 中文摘要
一对一(O2O)匹配通过将每个目标分配给单个正查询,使检测变换器(DETR)能够执行端到端的集合预测。然而,目标的最强分类、中心、尺度和重叠证据通常分布在多个查询中。这种不匹配导致只有被匹配的所有者对该目标获得正监督,而其他携带证据的查询则无法获得该目标的框目标。我们将此称为查询知识碎片化。为了在不采用一对多监督的情况下利用这种互补证据,我们提出了BS-O2G,一种即插即用模块,它从解码特征、框和类别分布构建稀疏的预测感知图,以在特征和优化空间中组织查询协作,同时保留原始的O2O匹配器、正标签和目标。一对图(O2G)校准通过该图传播相对消息,在前向传播中整合查询证据,而反向共享(BS)则重用其转置的分离邻接矩阵,将梯度路由到持久查询基向量上,而不改变前向传播中的解码器输入。在多种DETR方法、骨干网络、COCO和CrowdHuman上的实验表明,该方法在参数和FLOP增长可忽略、运行时开销适中的情况下,持续带来性能提升和更快的收敛速度,支持基于图的查询协作作为扩展正分配的替代方案。
英文摘要
One-to-one (O2O) matching enables Detection Transformers (DETRs) to perform end-to-end set prediction by assigning each object to a single positive query. However, the strongest classification, center, scale, and overlap evidence for an object is often distributed across multiple queries. This mismatch leaves only the matched owner positively supervised for the object, while other evidence-bearing queries receive no box target for it. We term this query knowledge fragmentation. To exploit such complementary evidence without one-to-many supervision, we propose BS-O2G, a plug-in that builds a sparse prediction-aware graph from decoded features, boxes, and class distributions to organize query collaboration in feature and optimization spaces while preserving the original O2O matcher, positive labels, and objective. One-to-Graph (O2G) calibration propagates relative messages over this graph to consolidate query evidence in the forward pass, whereas Backward Sharing (BS) reuses its transposed detached adjacency to route gradients across persistent query basis vectors without changing the decoder input in the forward pass. Experiments across diverse DETR methods, backbones, COCO, and CrowdHuman show consistent gains and faster convergence with negligible parameter/FLOP growth and modest runtime overhead, supporting graph-based query collaboration as an alternative to expanding positive assignments.
Commentspreprint