arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向跨域小样本目标检测的、结合特征扰动的提示驱动模拟方法

Prompt-Driven Simulation with Feature Perturbation for Cross-Domain Few-Shot Object Detection

Linhai Zhuo, Junxi Cai, Tianwen Qian, Qingping Zheng, Yang Liu

arXiv 2608.01348首次发表:更新:

发表机构

Fuzhou University; Xiamen University(福州大学; 厦门大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出PSP-FSOD框架,结合提示驱动域模拟与特征扰动正则化,缓解跨域小样本目标检测的域偏移与标注数据不足问题,在多个基准上实现性能提升。

AI 中文摘要

数据增强通过模拟多样的视觉变化以扩展源分布并引入合成域偏移,是缓解跨域小样本目标检测(CD-FSOD)中严重域偏移与目标标注数据有限问题的简单却有效的策略。现有方法依赖Color-Jitter、Mosaic、Domain-RAG这类常规数据增强,在建模复杂域偏移方面存在局限,且常导致性能欠佳。本文提出PSP-FSOD这一框架,它将提示驱动的域模拟与特征扰动正则化相结合,以提升CD-FSOD的泛化能力。为实现可控的域合成,我们设计了提示驱动策略,利用大型视觉语言模型(VLM)的视觉定位能力联合建模前景与背景变化,生成语义一致但域多样的训练样本;同时采用感知定位的生成方案,引导目标放置并缓解语义-空间错位,进而改善前景适配。为确保训练的稳定性与鲁棒性,我们还引入噪声诱导的特征扰动机制,向多尺度中间特征注入经分布修正的高斯噪声,鼓励模型在扰动下做出一致预测,减少对域特定线索的依赖。大量实验表明,PSP-FSOD能生成高质量的域多样监督信号,学习域不变表示,在多个CD-FSOD基准上均实现了性能提升。

英文摘要

Data augmentation, which simulates diverse visual variations to expand the source distribution and induce synthetic domain shifts, is a simple yet effective strategy for mitigating severe domain shifts and limited labeled target data in cross-domain few-shot object detection (CD-FSOD). Existing approaches rely on conventional data augmentation, such as Color-Jitter, Mosaic, and background-centric adaptation (e.g., Domain-RAG), which are limited in modeling complex domain shifts and often lead to suboptimal performance. In this paper, we propose PSP-FSOD, a principled framework that integrates prompt-driven domain simulation with feature perturbation regularization to improve generalization in CD-FSOD. To enable controllable domain synthesis, we design a prompt-driven strategy that leverages the visual grounding capability of large VLMs to jointly model foreground and background variations, generating semantically consistent yet domain-diverse training samples. Moreover, we adopt a grounding-aware generation scheme that guides object placement and alleviates semantic-spatial misalignment, thereby improving foreground adaptation. To ensure training stability and robustness, we further introduce a noise-induced feature perturbation mechanism that injects Gaussian noise into multi-scale intermediate features with distribution correction, encouraging consistent predictions under perturbations and reducing reliance on domain-specific cues. Extensive experiments demonstrate that PSP-FSOD produces high-quality domain-diverse supervision and learns domain-invariant representations, consistently improving performance across CD-FSOD benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑