发表机构
Massachusetts Institute of Technology; The University of Manchester; The Hong Kong Polytechnic University(麻省理工学院; 曼彻斯特大学; 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对现有指称遥感图像分割基准仅支持书面表达的问题,构建了首个语音查询指称分割基准RISBench-S,并提出高效双边网络模型AeroReformer2,在相关任务上取得优于基线的性能。
AI 中文摘要
口语为在密集遥感图像中指定任意目标提供了自然的免手持界面,但现有的指称遥感图像分割基准仅接受书面表达。为弥合这一差距,我们引入了数据集\textit{RISBench-S},这是一个从RISBench衍生而来的语音查询基准,在保留原始图像、掩码和数据划分的同时,添加了带有不同口音和语音的语音数据。其困难评估集结合了旋翼干扰、风干扰和混合干扰,并设置了三个信噪比水平。我们还提出了模型AeroReformer2,这是一种高效的双边网络,结合了保留边界的视觉路径、保留token的语音编码、核线性跨模态注意力以及分辨率细化头。该设计在两个尺度上对视觉特征进行条件约束,而无需实现密集的语音-视觉亲和矩阵,随后利用高分辨率视觉特征恢复精细边界。在干净测试划分上,采用Swin-Base的AeroReformer2达到了62.09%的平均交并比(mIoU)和68.22%的整体交并比(oIoU),分别比最强的音频适配遥感基线高出5.38和2.08个百分点;在困难集上也保持了最佳的mIoU,为54.09%。据我们所知,这是首个针对遥感图像的全句语音查询指称分割的基准与模型研究,代码将公开提供。
英文摘要
Spoken language offers a natural, hands-free interface for specifying an arbitrary target in dense remote-sensing imagery, yet existing referring remote-sensing image segmentation benchmarks accept only written expressions. To bridge this gap, we introduce \dataset, a spoken-query benchmark derived from RISBench that adds accent- and voice-diverse speech while preserving the original image, mask, and data splits. Its hard evaluation sets combine rotor, wind, and mixed interference with three signal-to-noise levels. We also propose \model, an efficient bilateral network that combines a boundary-preserving visual path with token-preserving speech encoding, kernel linear cross-modal attention, and a resolution refinement head. The design conditions visual features at two scales without materializing a dense speech--visual affinity matrix, then restores fine boundaries using high-resolution visual features. On the clean test split, \model with Swin-Base achieves 62.09\% mean intersection over union (mIoU) and 68.22\% overall intersection over union (oIoU), outperforming the strongest audio-adapted remote-sensing baseline by 5.38 and 2.08 percentage points, respectively. It retains the best hard-set mIoU at 54.09\%. To the best of our knowledge, this is the first benchmark and model study of full-sentence spoken-query referring segmentation for remote-sensing imagery. The code will be made publicly available.