arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Noctif3R:面向光子受限场景的嵌入式硬件上的前馈单目实时SLAM

Noctif3R: Feed-Forward Monocular Real-Time SLAM for Photon-Limited Scenes on Embedded Hardware

Mihir Chauhan, Aditya Uday Abhang, Kevin Biju Mathew, Aniket Bera

arXiv 2609.21114首次发表:更新:

发表机构

Purdue University(普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对黑暗环境中单目实时SLAM在低光下失效的问题,提出基于低光前馈点图前端和显式匹配门的SYS流程,在嵌入式硬件上实现实时、低误差的定位,并显著降低资源消耗。

AI 中文摘要

在黑暗环境中执行任务的机器人需要从单个RGB相机进行定位,此时光照极低,每像素信号接近传感器自身噪声,且需在功耗受限的机载计算机上实时运行。这些约束各自都有成熟的流程,但它们的交集尚未解决。离线低光重建现在可以恢复低于-4 dB的结构,但速度太慢无法实时运行,而机器人实际可携带的实时单目系统(如DROID-SLAM、DPV-SLAM等)在信噪比降低时会退化或失败。我们测量了它们的失败方式:在我们场景的九个最低黑暗级别中,DROID-SLAM在所有九个级别上返回全长度轨迹但不携带任何关于相机运动的信息,VGGT-SLAM和CUT3R在八个级别上如此,pi^3在七个级别上,DPV-SLAM在四个级别上。我们提出了SYS,一个基于低光前馈点图前端和显式匹配门的单目流程,它返回三条有跟踪的轨迹且没有无信息轨迹,在跟踪时误差最低(为无信息上限的24-47%,而最强基线为56-73%),且覆盖范围最窄。在一个真实机器人视频中,86.5%的帧完全为黑色,我们所有配置在亮起的开头后停止,而DROID-SLAM和DPV-SLAM各自为全部1178帧输出位姿。我们的方法贡献是为Jetson AGX Orin提供嵌入式执行路径:以384像素运行地图、关键帧和后端,以256像素进行跟踪,加上对每帧位姿求解的两处修复,是一个可复现的Pareto改进,在一个场景上吞吐量提升1.28倍且误差降至0.964倍,在第二个场景上提升1.42倍且误差降至0.68倍,峰值GPU内存减少47%,每个位姿能耗降低29%。我们在校准的、位精确可再生的噪声阶梯、重新标注的真实世界暗光曝光以及从波士顿动力Spot机器人录制的新的暗室视频阶梯上进行了评估。

英文摘要

Robots carrying out tasks in dark environments need to localize from a single RGB camera, in light so low that the per-pixel signal approaches the sensor's own noise, on a power-constrained onboard computer, in real time. Each of these constraints has matured pipelines, but the intersection does not. Offline low-light reconstruction now recovers structure below -4 dB but is far too slow to run in real time, while the real-time monocular systems a robot can actually carry (DROID-SLAM, DPV-SLAM, etc.) degrade or fail when SNR gets low. We measured how they fail: across the nine lowest darkness levels of our scenes, DROID-SLAM returns a full-length trajectory carrying no information about the camera's motion on all nine, VGGT-SLAM and CUT3R on eight, pi^3 on seven, and DPV-SLAM on four. We present SYS, a monocular pipeline built on a low-light feed-forward pointmap front end with an explicit match gate, which returns three tracked trajectories and no uninformative ones, at the lowest error of any method where it tracks (24-47% of the no-information ceiling against 56-73% for the strongest baseline), and at the narrowest coverage. On a real robot video take in which 86.5% of delivered frames are entirely black, every configuration of ours stops after the lit beginning, while DROID-SLAM and DPV-SLAM each emit a pose for all 1178 frames. Our method contribution is an embedded execution path for the Jetson AGX Orin: running the map, keyframes and backend at 384 pixels with tracking at 256, together with two fixes to the per-frame pose solve, is a replicated Pareto improvement, 1.28x throughput at 0.964x error on one scene and 1.42x at 0.68x on a second, with 47% less peak GPU memory and 29% less energy per pose. We evaluate on a calibrated, bit-exact regenerable noise ladder, on relabelled real-world dark exposures, and on a new dark-room video ladder recorded from a Boston Dynamics Spot robot.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑