发表机构
BDNRC(BDNRC)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出DCA-MoE人群计数框架,含空间自适应跨层融合与密度路由专家机制,在NWPU-Crowd数据集上取得MAE 31.7等结果,验证了空间自适应融合与路由的可行性。
AI 中文摘要
人群计数需要在视角、头部尺度、遮挡和背景杂波等严重变化下恢复可靠的局部密度。尽管现代计数目标提供了强大的空间监督,但许多多级解码器仍采用空间不变的特征融合,并对每个位置应用单一感受野模式。我们提出DCA-MoE框架,在保留冻结的DINOv3编码器的同时,使两种决策都依赖于内容。空间自适应层融合(SALF)会对四个对齐的骨干特征预测逐位置权重,密度路由多感受野专家(DR-MoE)则为每个位置分配局部、中程和大上下文残差专家的软混合。EBC风格的头部重构块密度,同时DMCount监督和辅助路由平衡项训练解码器而不更新骨干。在NWPU-Crowd验证集上,基于DINOv3 ViT-L/16的最强配对配置获得31.7的MAE和72.2的RMSE;匹配的ViT-B/16完整模型则获得32.2/75.9的结果。跨数据集结果仍参差不齐,且若干组件基线目前报告的是来自单个种子的独立选择最小值。因此,证据支持空间自适应融合与路由的可行性,但仍需更广泛的配对和多种子评估以进行因果归因。
英文摘要
Crowd counting must recover reliable local density under severe variations in perspective, head scale, occlusion, and background clutter. Although modern counting objectives provide strong spatial supervision, many multi-level decoders still use spatially invariant feature fusion and apply one receptive-field pattern to every location. We propose DCA-MoE, a framework that makes both decisions content dependent while retaining a frozen DINOv3 encoder. Spatially Adaptive Layer Fusion (SALF) predicts position-wise weights over four aligned backbone features, and Density-Routed Multi-Receptive-Field Experts (DR-MoE) assigns each location a soft mixture of local, mid-range, and large-context residual experts. An EBC-style head reconstructs block density, while DMCount supervision and an auxiliary routing-balance term train the decoder without updating the backbone. On the NWPU-Crowd validation split, the strongest paired configuration, based on DINOv3 ViT-L/16, obtains 31.7 MAE and 72.2 RMSE; the matched ViT-B/16 full model obtains a paired 32.2/75.9. Cross-dataset results remain mixed, and several component baselines currently report independently selected minima from a single seed. The evidence therefore supports the feasibility of spatially adaptive fusion and routing, while broader paired and multi-seed evaluation remains necessary for causal attribution.