发表机构
Conflux Laboratory(Conflux实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出一种实时边缘视觉流水线,结合YOLO26x、MiVOLO v2等技术,用于儿童检测与年龄估计,可提升童工监测系统的检测效率与准确性,经实地试点验证效果显著。
AI 中文摘要
据估计,全球仍有1.38亿儿童从事童工劳动,受影响行业采用的监测系统基于定期家访和访谈,系统地低估了童工数量。我们提出一种仅作为研究原型构建和运行的实时计算机视觉流水线,旨在研究为童工监测与补救系统(CLMRS)提供持续的、基于存在证据的渠道的可行性。该流水线结合了多任务人物与面部检测器(CerberusDet框架中采用YOLO26x主干)、MiVOLO v2与针对0-12岁儿童的专用模型配对的级联年龄估计、ByteTrack跟踪、ArcFace与DINOv2重识别,以及生成可审查的个人记录的跟踪级融合。该检测器将人物mAP@0.5从0.390提升至0.683,超过前代基线;儿童专用模型在仅含儿童的验证集上达到1.944年的平均绝对误差(MAE),而广泛使用的开源模型栈误差为18-23年。FP8 TensorRT编译带来1.77倍的加速,MAE仅增加0.002年,使流水线在嵌入式硬件上实现两倍以上实时性。在26.8小时的代理视频上,该系统检测到634名独特儿童候选,而前代仅为285名。我们还报告了在津巴布韦某农场开展的为期17天的无人值守实地试点(3870万帧、6台相机),对照每日考勤登记评估:软件调优使检测产量提升36倍,且在同时性否决规则下的身份合并将重复报告率从9.1倍降至1.8-3.9倍,且无经证实的错误合并。我们记录了训练和量化的成功与失败,以及此类系统所需的数据保护和人在回路的保障措施。
英文摘要
An estimated 138 million children remain in child labour worldwide, and the monitoring systems used by affected sectors, built on periodic household visits and interviews, systematically under-detect them. We present a real-time computer-vision pipeline, built and operated solely as a research prototype, that studies the feasibility of giving Child Labour Monitoring and Remediation Systems (CLMRS) a continuous, presence-based evidence channel. The pipeline combines a multi-task person and face detector (YOLO26x backbone in the CerberusDet framework), cascaded age estimation pairing MiVOLO v2 with a child-specialist model for ages 0-12, ByteTrack tracking, ArcFace and DINOv2 re-identification, and track-level fusion producing reviewable per-person records. The detector raises person mAP@0.5 from 0.390 to 0.683 over the previous-generation baseline; the child specialist reaches 1.944 years MAE on children-only validation, where widely used open-source stacks err by 18-23 years. FP8 TensorRT compilation yields a 1.77x speedup at +0.002 years MAE, bringing the pipeline above twice real-time on embedded hardware. On 26.8 hours of proxy video the system finds 634 unique child candidates versus 285 for its predecessor. We further report a seventeen-day unattended field pilot on a farm in Zimbabwe (38.7 million frames, six cameras) evaluated against a daily attendance register: software tuning improved detection yield 36-fold, and identity consolidation under a simultaneity veto cut over-reporting from 9.1x to 1.8-3.9x with zero proven-false merges. We document training and quantisation failures alongside successes, and the data-protection and human-in-the-loop safeguards such a system requires.
Comments39 pages, 1 figure, 13 tables