SeGDeP:用于推理分割的语义与几何感知解耦提示
SeGDeP: Semantic- and Geometric-Aware Decoupled Prompts for Reasoning Segmentation
浏览论文内容
中文总结 AI 辅助
SeGDeP提出语义与几何解耦提示的显式接口,通过GDPO优化平衡反馈,在推理分割任务中以极少参数实现高性能。
中文摘要 AI 辅助
推理分割将隐含的语言结论转化为精确的掩码,需要同时进行语义识别和空间定位。现有的多模态大语言模型(MLLM)分割器接口要么使用特殊触发器,要么将两种信号压缩到一个上下文中,尽管它们接收不同的监督并以不同的方式失败。这种耦合掩盖了失败是源于目标解释还是定位。我们提出了SeGDeP,一种显式的“什么-哪里”接口。一个语义提示分支和一个独立的几何投影路径将解析后的MLLM状态转化为语义特征和DETR预测的边界框,这些共同条件化SAM 3掩码解码器。训练首先对齐这一可执行接口,然后使用组奖励解耦策略优化(GDPO)来平衡格式、边界框IoU和掩码IoU反馈。SeGDeP-4B在八个RefCOCO系列分割上达到82.7的平均cIoU,在ReasonSeg验证/测试上达到66.0/59.6的gIoU,同时仅通过LoRA调整Qwen3-VL参数的0.38%。受控的逐阶段消融、梯度诊断和提示干预进一步表明,两条路径发展出互补的语义和几何专门化,而非重复相同的证据。
英文摘要
Reasoning segmentation converts an implicit linguistic conclusion into a precise mask, requiring both semantic identification and spatial grounding. Existing MLLM-segmenter interfaces either use a special trigger or compress both signals into one context, although they receive different supervision and fail differently. This coupling obscures whether a failure arises from target interpretation or from localization. We present SeGDeP, an explicit what-where interface. A semantic prompt branch and an independent geometric projection path transform resolved MLLM states into semantic features and a DETR-predicted box, which jointly condition a SAM 3 mask decoder. Training first aligns this executable interface, then uses group reward-decoupled policy optimization (GDPO) to balance format, box-IoU, and mask-IoU feedback. SeGDeP-4B reaches 82.7 average cIoU over eight RefCOCO-family splits and 66.0/59.6 gIoU on ReasonSeg val/test while adapting only 0.38% of Qwen3-VL parameters through LoRA. Controlled stage-wise ablations, gradient diagnostics, and prompt interventions further show that the two paths develop complementary semantic and geometric specialization rather than duplicating the same evidence.
发表机构
- Xidian University(西安电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。