HeatTok:基于热扩散分词的遥感图像理解增强方法
HeatTok: Enhancing Remote Sensing Image Understanding via Thermodiffusion-based Tokenization
浏览论文内容
中文总结 AI 辅助
本文提出HeatTok分词器,结合热扩散聚合与G-MRoPE编码,在VRSBench、EarthVQA数据集上实现遥感图像理解的最优性能,保留对象级语义完整性。
中文摘要 AI 辅助
多模态大语言模型(MLLMs)中当前的视觉分词器主要依赖基于图像块的划分方式,由于地理对象的不规则轮廓,这会导致遥感图像中出现严重的语义混合和对象碎片化问题。此外,现有的自适应方法难以提取精确的对象级分词,且缺乏针对不规则区域的专用几何位置编码。本文提出了HeatTok,一种由热扩散聚合驱动的语义感知分词器。受热传导物理原理启发,HeatTok自适应合并相邻同质区域,生成语义独立、与对象对齐的不规则分词。为使MLLMs能够感知这些不规则形状,我们设计了高斯多模态旋转位置编码(G-MRoPE),该编码通过二维高斯函数对分词空间分布进行建模,并显式注入中心、尺度和方向信息。在VRSBench和EarthVQA数据集上的大量评估表明,HeatTok能有效保留对象级语义完整性,并在合理的分词预算下达到了最优性能。代码可访问:this https URL。
英文摘要
Current visual tokenizers in Multimodal Large Language Models (MLLMs) predominantly rely on patch-based partitioning, which causes severe semantic mixture and object fragmentation in remote sensing imagery due to the irregular contours of geo-objects. Moreover, existing adaptive methods struggle to extract precise object-level tokens and lack dedicated geometric positional encodings for irregular regions. In this paper, we propose HeatTok, a semantic-aware tokenizer driven by thermodiffusion aggregation. Inspired by the physical principles of heat conduction, HeatTok adaptively merges adjacent homogeneous regions to generate semantically independent, object-aligned irregular tokens. To enable MLLMs to perceive these irregular shapes, we design the Gaussian Multimodal Rotary Positional Embedding (G-MRoPE), which models token spatial distributions via 2D Gaussians and explicitly injects center, scale, and orientation cues. Extensive evaluations on the VRSBench and EarthVQA datasets demonstrate that HeatTok effectively preserves object-level semantic integrity and achieves state-of-the-art performance under a reasonable token budget. The code is available: https://github.com/YingyingYan1/HeatTok.
发表机构
- Northwestern Polytechnical University(西北工业大学)
- Hong Kong University of Science and Technology(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。