GramLoop:用于鲁棒密集预测的无训练Gram门控重放方法
GramLoop: Training-Free Gram-Gated Replay for Robust Dense Prediction
浏览论文内容
中文总结 AI 辅助
该研究提出无训练框架GramLoop,通过在冻结DINOv3的视觉骨干内添加推理计算,在不改变模型结构的情况下,提升了分布偏移下目标检测和语义分割等密集预测任务的鲁棒性,在COCO-O等基准上取得性能提升。
中文摘要 AI 辅助
我们旨在通过在视觉骨干网络内部添加推理计算,改进分布偏移下的冻结DINOv3密集预测模型,且不改变模型权重、任务适配器或预测头。挑战在于,重复的Transformer块计算必须细化密集特征,同时不破坏DINOv3用于保留空间结构的成对补丁关系。我们提出GramLoop,这是一种无训练框架,它重放短Transformer窗口,并通过最终层的余弦-Gram一致性控制每次重放。每个提案通过冻结的后缀传播,与标准DINOv3轨迹进行比较,并在重放窗口端点处通过逐补丁门接受。在存在损坏、扰动和自然偏移的目标检测和语义分割任务中,GramLoop在五个偏移基准上均优于配对的DINOv3基线。在COCO-O上,它将mAP提升0.252,有效鲁棒性提升0.250,同时保留干净ADE20K的性能。代码将被发布。
英文摘要
We aim to improve frozen DINOv3 dense-prediction models under distribution shift by adding inference computation inside the visual backbone, without changing model weights, task adapters, or prediction heads. The challenge is that repeated transformer-block computation must refine dense features without disrupting the pairwise patch relations that DINOv3 uses to preserve spatial structure. We introduce GramLoop, a training-free framework that replays a short transformer window and controls each replay through final-layer cosine-Gram consistency. Each proposal is propagated through the frozen suffix, measured against the standard DINOv3 trajectory, and accepted through a patchwise gate at the replay-window endpoint. Across object detection and semantic segmentation under corruptions, perturbations, and natural shifts, GramLoop improves all five shifted benchmarks over the paired DINOv3 baseline. On COCO-O, it improves mAP by +0.252 and Effective Robustness by +0.250, while preserving clean ADE20K performance. Code will be released at https://github.com/cheyan9/GramLoop.
发表机构
- Huazhong University of Science and Technology(华中科技大学)
- Tongji University(同济大学)
- King Abdullah University of Science and Technology(阿卜杜拉国王科技大学)
- University of California, Merced(加州大学默塞德分校)
- University of Amsterdam(阿姆斯特丹大学)
机构由 AI 辅助整理,请以论文原文为准。