PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation
PixDLM:一种用于无人机推理分割的双路径多模态语言模型
机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(中国教育部多媒体可信感知与高效计算重点实验室,厦门大学)
AI总结 本文提出PixDLM,一种用于无人机推理分割的多模态语言模型,通过构建DRSeg基准数据集,验证了该模型在处理高分辨率无人机图像中的有效性。
Comments Accepted to CVPR 2026 (highlight)