Cyclops:以彩色进行“梦境”的相机式激光雷达
Cyclops: LiDAR as a Camera That Dreams in Color
浏览论文内容
中文总结 AI 辅助
本文提出Cyclops框架,将稀疏NRS-LiDAR强度转为RGB视频,通过LBM等技术缓解帧间闪烁,合成RGB可使感知模型在多任务上优于LiDAR基线与传统相机。
中文摘要 AI 辅助
传统上,机器人感知高度依赖相机,因为相机能提供丰富的语义纹理,但在低光照或高动态范围环境中性能会显著下降。相反,光探测与测距(LiDAR)可捕获与光照无关的几何和强度属性,不过其生成的数据通常为单通道且稀疏,在应用基于RGB数据集预训练的视觉模型时会产生显著的模态差距。本文提出Cyclops框架,该框架将稀疏非重复扫描激光雷达(NRS-LiDAR)强度转换为RGB视频,支持全天候感知任务的无相机推理。我们的方法首先通过冻结的预训练密集化模块将稀疏LiDAR强度投影转换为密集表示,作为几何丰富的源条件;随后通过带有学习速度场的潜在桥接匹配(LBM),在少量常微分方程(ODE)积分步骤中将密集强度潜在特征迁移至目标RGB分布。为缓解帧间闪烁,我们通过时间注意力层注入前一帧上下文,并进一步将速度场表述为策略,该策略由可微终端奖励优化,该奖励通过沿ODE轨迹的反向传播鼓励终端保真度。大量实验表明,合成的RGB(包括近暗条件下生成的)能使标准基于RGB的感知模型在不同光照条件下的语义分割、车道检测和点云着色任务中,显著优于LiDAR基线和传统相机。
英文摘要
Conventionally, robotic perception relies heavily on cameras due to the rich semantic texture they provide. However, their performance degrades significantly in low-light or high-dynamic-range environments. Conversely, while Light Detection and Ranging (LiDAR) captures illumination-invariant geometric and intensity properties, the resulting data are typically single-channel and sparse, creating a significant modality gap when applying vision models pre-trained on RGB datasets. In this paper, we propose Cyclops, a framework that translates sparse Non-Repetitive Scanning LiDAR (NRS-LiDAR) intensity into RGB video, enabling camera-free inference for all-day perception tasks. Our approach first converts sparse LiDAR intensity projections into dense representations via a frozen pre-trained densification module, serving as a geometrically rich source condition. The dense intensity latent is then transported toward the target RGB distribution through Latent Bridge Matching (LBM) with a learned velocity field in a few ODE integration steps. To mitigate inter-frame flickering, we inject prior-frame context via temporal attention layers and further formulate the velocity field as a policy optimized by a differentiable terminal reward that encourages terminal fidelity through backpropagation along the ODE trajectory. Extensive experiments demonstrate that the synthesized RGB, including those generated under near-dark conditions, enable standard RGB-based perception models to substantially outperform both LiDAR baselines and conventional cameras on semantic segmentation, lane detection, and point cloud colorization across diverse lighting conditions.
发表机构
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- University of Michigan(密歇根大学)
- University of Macau(澳门大学)
机构由 AI 辅助整理,请以论文原文为准。