发表机构
National Yang Ming Chiao Tung University(国立阳明交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对视觉SLAM中描述符在光照和视角变化下性能不佳的问题,提出Desc++轻量级增强模块,通过混合架构联合编码描述符和关键点几何信息,实验表明其能提高匹配精度、稳定轨迹估计,在准确性和效率间取得良好平衡。
AI 中文摘要
可靠的视觉数据关联是视觉同步定位与地图构建(V-SLAM)的基础,它直接决定相机位姿估计质量和地图一致性。大多数成熟实时系统使用的手工描述符在光照和视角变化下性能下降,基于学习的前端虽能解决此问题但需替换提取和匹配流程并引入大量计算开销。描述符增强可在原始格式内优化现有描述符,但当前方法依赖简化注意力机制,其有限的上下文建模限制了匹配质量。为解决上下文表达能力和效率之间的权衡,我们提出Desc++,这是一个轻量级增强模块,通过混合架构联合编码描述符表示和关键点几何信息,并在线性时间内通过结合无序全局注意力和几何感知顺序建模来聚合空间上下文。增强后的描述符保留其原始维度和匹配接口,可集成到已部署的V-SLAM系统中而无需修改流程。在描述符匹配、对应分析和四个不同V-SLAM系统的系统级基准测试中进行的实验表明,Desc++比现有最先进的增强方法提高了匹配精度,将这些提升转化为更准确和稳定的轨迹估计,并在准确性和效率之间取得了良好平衡,便于实际集成到现有的实时V-SLAM管道中。
英文摘要
Reliable visual data association is fundamental to visual SLAM (V-SLAM), as it directly determines the quality of the camera pose estimation and map consistency. However, the handcrafted descriptors used by most mature real-time systems degrade under illumination and viewpoint changes, while learning-based front-ends that address this weakness typically require replacing the extraction-and-matching pipeline and introduce substantial computational overhead. Descriptor enhancement offers a compromise by refining existing descriptors within their original format, yet current methods rely on simplified attention mechanisms whose limited contextual modeling constrains the achievable matching quality. To resolve this trade-off between contextual expressiveness and efficiency, we propose Desc++, a lightweight enhancement module that jointly encodes descriptor representations and keypoint geometry and aggregates spatial context through a hybrid architecture that combines order-agnostic global attention with geometry-aware sequential modeling in linear time. The enhanced descriptors retain their original dimensionality and matching interface, enabling integration into deployed V-SLAM systems without modifying the pipeline. Experiments across descriptor matching, correspondence analysis, and system-level benchmarks with four different V-SLAM systems demonstrate that Desc++ improves matching accuracy over the state-of-the-art enhancement method, translates these gains into more accurate and stable trajectory estimation, and achieves a favorable balance between accuracy and efficiency for practical integration into existing real-time V-SLAM pipelines.
Comments12 pages, 6 figures, and 9 tables