arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28183cs.DCcs.CV

用于实时非视距重建的内存高效型GPU流水线

Memory-efficient GPU pipelines for real-time non-line-of-sight reconstruction

Alfonso López-Ruiz, Diego Royo

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对NLOS成像的GPU重建瓶颈,优化两种波基算法的GPU流水线,实现速度大幅提升且内存占用极低,还提出适用于下一代NLOS视频处理的降噪策略。

中文摘要 AI 辅助

非视距(NLOS)成像是通过单光子雪崩二极管(SPAD)记录的间接光,重建拐角处隐藏场景的技术。单次重建是一个大型逆问题:需对数十亿个光子时间戳进行分箱、内存移动、变换与逆变换。随着SPAD阵列提升采集吞吐量,重建成为瓶颈。我们为两种已有的基于波的算法(f-k偏移和phasor-fields)重建了GPU执行流程,支持流式和离线处理。在phasor-fields方面,我们利用环形的解析傅里叶变换,一次性离线组装先前工作中的环形与半径核,使传播核在运行时永远不以稠密形式存在,从而减少内存与带宽消耗。我们通过融合核、warp级光子分箱、批量变换、CUDA图重放,以及仅在实际瓶颈处使用FP16存储,重新组织了两种算法的流水线。我们的实现比参考流式流水线快达42倍,比已发表的最快GPU基准快达14倍,同时仅使用少量内存(低至2.5%),使相同硬件上能进行更大、更精细的重建,或在低得多的内存预算内进行相当的重建。我们报告了各实现选择的 ablation 分析,并提出三种降噪策略,该策略由下一代NLOS视频处理的帧预算所支持。

英文摘要

Non-line-of-sight (NLOS) imaging reconstructs scenes hidden around a corner from indirect light recorded by a single-photon avalanche diode (SPAD). A single reconstruction is a large inverse problem: billions of photon timestamps must be binned, moved through memory, transformed and inverted. As SPAD arrays raise acquisition throughput, reconstruction becomes the limiting stage. We rebuild the GPU execution of two established wave-based algorithms, f-k migration and phasor-fields, for both streaming and offline processing. On the phasor-fields side we assemble the ring-and-radius kernels of previous work once and offline, using the analytic Fourier transform of a ring, so the propagation kernel never exists in dense form at runtime, reducing the memory and bandwidth. We reorganize the pipeline of both algorithms with fused kernels, warp-level photon binning, batched transforms, CUDA graph replay, and FP16 storage applied only where it reduces the actual bottleneck. Our implementations are up to 42x faster than the reference streaming pipeline and up to 14x faster than the fastest published GPU baseline, all while using a fraction of the memory (down to 2.5%), enabling vastly larger and finer reconstructions on the same hardware, or comparable ones within a much lower memory budget. We report an ablation of each implementation choice and propose three denoising strategies enabled by the resulting frame budget for next-generation NLOS video processing.

↑