Phantom-Insight: Adaptive Multi-cue Fusion for Video Camouflaged Object Detection with Multimodal LLM
Phantom-Insight:基于多模态大语言模型的视频伪装目标自适应多线索融合检测方法
机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) ; School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
专题命中 视频多模态 :multimodal(title);MLLM(abstract,abstract_cn);分类 cs.CV
AI总结 该研究针对视频伪装目标检测的两大问题,提出Phantom-Insight方法,通过多线索融合与解耦学习优化SAM,在MoCA-Mask数据集上达最优性能,且泛化能力强。