arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

手持手机的行人概率预测:世界坐标系热图、视觉惯性高度漂移及无真值评估

Probabilistic Pedestrian Forecasts from a Handheld Phone: World-Frame Heat Maps, Visual-Inertial Height Drift, and Evaluation without Ground Truth

Danial Safaei

arXiv 2610.04736首次发表:更新:

发表机构

WMG, University of Warwick(华威大学WMG)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文构建并评估了一个仅用手机单目相机和VIO的行人概率预测系统,通过射线-平面交点提升行人、U-Net生成每步概率图,并解决VIO高度漂移问题,在无真值情况下用自洽性评分验证了预测器的有效性。

AI 中文摘要

携带手机的行人可以看到附近的人在接下来几秒内可能的位置,前提是当手机移动时预测保持在地面上、经过校准,并且仅需单目相机和视觉惯性里程计(VIO)。我们构建并评估了这样一个系统。行人被检测出来,通过射线-平面交点提升到地面,在重力对齐的度量坐标系中跟踪,并由一个小型U-Net预测为每步概率图,该U-Net使用负对数似然(NLL)损失在鸟瞰图(SDD)和第一人称(EgoTraj-Bench)轨迹上训练。在手持ADVIO记录中,垂直VIO漂移和用户自身的高度变化会静默地重新缩放单目地面位置(在一个片段中90秒内缩放87%;在另一个片段中,所有轨迹在片段最后31%丢失);将相机相对于地面的高度保持在低通滤波高度下可以避免此问题,尽管在自动扶梯上会有滞后。在SDD和EgoTraj-Bench测试集上,最终预测器在4.8秒时的NLL相对于拟合的恒定速度高斯分布降低了1.51和1.37纳特。由于手持视频中缺乏行人真值,我们根据跟踪器自身的后续原始测量来评分预测。在一项内部预注册的评估中,针对七个保留片段,最终预测器在1.2、2.4和4.8秒时的NLL均低于基准拟合基线的NLL(分别低0.17、0.23和0.47纳特;行人上的95%区间排除零),但改善幅度不到开发片段的一半。探索性分析结果双向:对片段而非行人进行重采样会使区间在1.2和2.4秒时包含零,并且一旦两个预测器在开发片段上重新校准,网络仅在1.2秒时显著更好;但使用ADVIO参考位姿运行的两个片段,在改用手机自身位姿重新运行时,网络的优势更大。我们讨论了这种自洽性评分能证明什么和不能证明什么。

英文摘要

A pedestrian with a phone could be shown where nearby people will be in the next few seconds, if the forecast stays on the ground while the phone moves, is calibrated, and needs only a monocular camera and visual-inertial odometry (VIO). We build and evaluate such a system. People are detected, lifted onto the floor by ray-plane intersection, tracked in a gravity-aligned metric frame, and forecast as per-step probability maps by a small U-Net trained with a negative log-likelihood (NLL) loss on bird's-eye (SDD) and first-person (EgoTraj-Bench) trajectories. On handheld ADVIO recordings, vertical VIO drift and the user's own changes of level silently rescale monocular ground positions (by 87% within 90 s on one clip; on another, all tracks are lost for the last 31% of the clip); keeping the camera's height above the floor constant under a low-pass-filtered altitude avoids this, though it lags on escalators. On the SDD and EgoTraj-Bench test splits, the final forecaster lowers the NLL at 4.8 s by 1.51 and 1.37 nats relative to a fitted constant-velocity Gaussian. Lacking ground truth for people in handheld video, we score forecasts against the tracker's own later raw measurements. In an internally pre-registered evaluation on seven held-out clips, the final forecaster's NLL is lower than the benchmark-fitted baseline's at 1.2, 2.4 and 4.8 s (by 0.17, 0.23 and 0.47 nats; 95% intervals over people exclude zero), but by less than half as much as on the development clips. Exploratory analyses cut both ways: resampling clips instead of people widens the intervals to include zero at 1.2 and 2.4 s, and once both forecasters are recalibrated on the development clips the network is significantly better only at 1.2 s; but two clips run with ADVIO's reference poses favour the network much more when re-run with the phone's own poses. We discuss what such self-consistency scores can and cannot show.

Comments28 pages, 7 figures, 16 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑