AI 中文总结
本文提出了纯整数流核心的SweepLSD线段检测器,其内存为O(宽度),实现了1080p30实时FPGA检测,CPU速度优于ELSED等,在消失点和相机姿态任务上表现出色。
AI 中文摘要
我们提出了SweepLSD,一种线段检测器,它仅读取图像一次,并在扫描线经过其最后一个像素的几行内输出每个线段。包括连通分量标记和最终线段测试在内的每个阶段都以行流形式处理图像:中间内存为O(宽度)而非O(像素),且每个像素的核心为纯整数运算。我们首次完整描述了该算法,该算法设计于作者2014年的硕士论文中但从未发表,同时提供了开源C++17实现和FPGA实现——其硬件配置与软件的比特级完全一致——该实现可在2009年的芯片上实时检测1080p30视频中的线段,无需帧缓冲区或外部内存。在结构丰富的公开4K照片下采样至全高清的数据集上,单CPU线程检测线段耗时约11毫秒,比原作者实现的ELSED、EDLines和LSD分别快4.6倍、5.2倍和25倍,且具有最紧凑的帧时间分布和四个检测器中最佳的单线段方向精度,还具备固有曲线剔除能力,仅在合成真值上的F值略逊于ELSED。在对York Urban和NYU-VP数据集的曼哈顿框架消失点研究中,采用每个检测器单独选择/评估最佳估计器的协议对所有检测器进行评分,在此协议下,SweepLSD在NYU-VP上领先约0.3度,在York Urban上落后0.1度,且是四个检测器中两者端到端流水线最快的。单帧相机姿态应用在具有精确真值的合成场景以及EuRoC和TUM-VI数据集上进行评估,其精度与基线相当但内存远小于基线,可实现4K地平线锁定,中位数姿态误差为0.06度,每帧中位数耗时32毫秒。
英文摘要
We present SweepLSD, a line segment detector that reads the image exactly once and emits each segment within a few rows of its last pixel passing the scan line. Every stage, including connected-component labeling and the final line test, processes the image as a row stream: intermediate memory is O(width) rather than O(pixels), and the per-pixel core is integer-only. We give the first complete description of the algorithm, designed in the author's 2014 master's thesis but never published, together with an open-source C++17 implementation and an FPGA realization -- held bit-exact against the software in its hardware configuration -- detecting segments in live 1080p30 video on 2009-era silicon without frame buffer or external memory. On structure-rich public 4K photographs downscaled to Full-HD, one CPU thread detects segments in ~11 ms -- 4.6x/5.2x/25x faster than the original authors' implementations of ELSED, EDLines, and LSD -- with the tightest frame-time distribution and the best per-segment direction accuracy of the four detectors, and curve rejection by design, while trailing ELSED in F-score on synthetic ground truth. A Manhattan-frame vanishing-point study on York Urban and NYU-VP scores every detector under a selection/evaluation-separated best-estimator-per-detector protocol, under which SweepLSD leads on NYU-VP by ~0.3 degrees and trails by 0.1 degrees on York Urban, with the fastest end-to-end pipeline of the four detectors on both. A single-frame camera-attitude application, evaluated on synthetic scenes with exact ground truth and on EuRoC and TUM-VI, matches the baselines' accuracy at a fraction of their memory, and drives a 4K horizon lock to 0.06 degrees median attitude error at 32 ms median per frame.
Comments40 pages, 12 figures, 18 tables. Code, benchmarks, and evaluation harnesses (MIT): https://github.com/yosh-shimizu/sweeplsd