发表机构
Harvard Medical School; Brigham and Women’s Hospital; Technical University of Munich(哈佛医学院; 布莱根妇女医院; 慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有跨模态描述方法需重新训练或模型运行时间长的问题,提出CrossFeat框架,通过特征描述符空间的跨模态函数实现单模态描述符跨模态运行,在多模态匹配中性能提升。
AI 中文摘要
目前,大部分关键点描述的进展都集中在单模态场景,其中图像变化源于视角、光照或对比度改变。多模态场景涉及由完全不同传感过程生成的图像,如多光谱成像、RGB-深度、卫星图像或医学成像,导致相同结构呈现出不同外观。跨模态描述的常见解决方案是为每对模态训练描述符,这需要在模态变化时重新训练;或训练大型模型,这会显著增加运行时间。相反,我们提出CrossFeat,这一框架可让现有单模态描述符跨模态运行。我们的方法在描述符空间中学习一个跨模态函数,将来自一种模态的特征映射到与另一种模态兼容的表示。为保留原始描述符捕获的结构信息,CrossFeat引入几何-外观解耦,仅改变外观,同时保留几何属性。在多个领域和数据集上的实验表明,该方法在多模态匹配中性能有所提升。
英文摘要
Most advances in keypoint descriptions address monomodal settings, where image variations arise from viewpoint, illumination, or contrast changes. Multimodal scenarios involve images produced by fundamentally different sensing processes, such as multispectral imaging, RGB-depth, satellite imagery, or medical imaging, causing the same structures to appear differently. A common solution to cross-modal description is to train descriptors for each modality pair, which requires retraining whenever the modalities change, or to train large models, which incur a significant increase in runtime. Instead, we propose CrossFeat, a framework that enables an existing monomodal descriptor to operate across modalities. Our method learns a crossing function in descriptor space that maps features from one modality to a representation compatible with another. To preserve the structural information captured by the original descriptor, CrossFeat introduces a geometry-appearance disentanglement such that only appearance is altered while the geometric properties are preserved. Experiments across multiple domains and datasets demonstrate improved performance in multimodal matching.
CommentsECCV 2026