arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

像人类一样移动:用于毫瓦级硬件上实时空中人员跟踪的自我运动归一化时间特征

Moving Like a Human: Ego-Motion-Normalized Temporal Signatures for Real-Time Aerial Person Tracking on Milliwatt-Class Hardware

Akbar Anbar Jafari, Cagri Ozcinar, Gholamreza Anbarjafari

arXiv 2607.16282首次发表:更新:

发表机构

University of Tartu; S Holding OÜ(塔尔图大学; 3S控股有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对无人机上人员跟踪难题,提出EMTS - Det五阶段系统,通过估计自我运动、转换帧等操作跟踪人员,在毫瓦级硬件上有良好表现,如在树莓派Zero 2W上有较高帧率和检测精度,能有效应对遮挡。

AI 中文摘要

跟随人员跟踪必须在无人机自身上运行,而价格实惠的配套计算机每秒仅提供少数有效的8位整数运算量。在典型的跟踪距离下,一个人占据10 - 60像素,与杂波难以区分,超出单帧外观检测器的能力范围。缺失的证据是时间性的,应在通过解析计算得到的输入表示中,而非在学习到的时间机制中。EMTS - Det是一个五阶段系统,它估计自我运动,将每一帧转换为自我运动归一化的残余运动通道,用一个22k参数、7.6兆次浮点运算的网络检测人员中心,在稳定坐标中用卡尔曼滤波器跟踪锁定目标,并用一维人类运动卷积分类器验证跟踪结果(ROC AUC为0.941)。训练使用合成运动课程,其运动通道由部署的自我运动代码生成。多种子消融实验确定了泛化的价值:在保留的VisDrone - DET数据集上,仅亮度变体的AP25降至0.051,而相同微调的YOLOv8n尽管计算量是其1100倍,也降至0.415,而部署的8位整数检测器在域内达到0.694 AP25,在此分割上为0.444。时间移位模块会降低准确性,所以部署的检测器是无状态的。记录了无声的8位整数校准失败情况;使用传播缓存的最小 - 最大校准与浮点值在0.008 AP内匹配。在树莓派Zero 2W上,该管道在1000多个真实世界无人机视频上以31.85 FPS运行,AP25为0.462,召回率为0.714,而YOLOv8n为1.95 FPS和0.172 AP25。一个57秒的现场序列显示在1.3秒自动锁定,锁定召回率为97.9%,并能从所有九次遮挡中恢复且无错误重新锁定。

英文摘要

Follow-me person tracking must run on the drone itself, where affordable companion computers offer only a few effective int8 GFLOP/s. At typical follow distances a person spans 10-60 pixels, indistinguishable from clutter and beyond the reach of single-frame appearance detectors. The missing evidence is temporal and belongs in the input representation, computed analytically, rather than in learned temporal machinery. EMTS-Det is a five-stage system that estimates ego-motion, converts each frame into ego-motion-normalized residual-motion channels, detects person centers with a 22k-parameter, 7.6-MFLOP network, tracks a locked target with a Kalman filter in stabilized coordinates, and verifies tracks with a 1-D convolutional classifier of human motion (ROC AUC 0.941). Training uses a synthetic-motion curriculum with motion channels generated by the deployed ego-motion code. Multi-seed ablations locate the value in generalization: on held-out VisDrone-DET a luminance-only variant collapses to 0.051 AP25 versus 0.415, as does YOLOv8n fine-tuned identically despite 1,100 times the compute, while the deployed int8 detector reaches 0.694 AP25 in-domain and 0.444 on this split. Temporal-shift modules lower accuracy, so the deployed detector is stateless. Silent int8 calibration failures are documented; min-max calibration with propagated caches matches float within 0.008 AP. On a Raspberry Pi Zero 2W the pipeline runs at 31.85 FPS with 0.462 AP25 and 0.714 recall over 1,000 real-world UAV videos, versus 1.95 FPS and 0.172 AP25 for YOLOv8n. A 57-second field sequence shows auto-lock at 1.3 s, 97.9% lock recall, and recovery from all nine occlusions with zero false re-locks.

Comments20 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑