发表机构
Faculty of Artificial Intelligence, Universiti Teknologi Malaysia(马来西亚理工大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对吉隆坡热带城市交通场景的隐私保护数据集构建难题,提出结合Grounding DINO与空间车辆ROI包含引擎的自动化匿名化框架,在1266帧图像上实现约95%的匿名化成功率。
AI 中文摘要
智能交通系统与自动驾驶的快速发展高度依赖多模态城市交通数据集。然而,在马来西亚吉隆坡这类复杂热带城市环境中构建高保真视频图像数据集时,由于摩托车密度高、车牌为深色亚克力材质、相机动态倾斜以及强烈热带眩光,个人身份信息(PII)匿名化面临严峻挑战。本文针对通过移动骑行平台以2 FPS采集的吉隆坡道路数据集,提出一种自动化匿名化框架。研究发现传统Haar级联和YOLOv8在该场景下失效——对背景元素产生误检,同时遗漏旋转或被遮挡的目标。本架构通过将零样本开放集视觉-语言Transformer模型Grounding DINO与新型空间车辆感兴趣区域(ROI)包含引擎相结合解决上述问题。该流程要求车牌中心位于经验证的车辆边界内,从而抑制环境误检,同时自动模糊人脸、头部和车牌。对1266帧图像的初步评估显示,成功率约为95%,剩余失败案例仅限于小型、严重遮挡、倾斜或模糊的目标。结合时间持久性机制与自动化质量控制审计器,该框架在保留下游视觉任务场景上下文的同时,最大限度减少隐私相关的漏检。尽管正式法律合规性取决于更广泛的治理程序,但此公开可用的流程与演示笔记本为隐私感知的数据集构建提供了可审计的预处理阶段。
英文摘要
The rapid advancement of intelligent transportation systems and autonomous driving relies heavily on multi-modal urban traffic datasets. However, curating high-fidelity video imagery in complex tropical urban environments---specifically Kuala Lumpur, Malaysia---presents severe challenges for Personally Identifiable Information (PII) anonymization due to high motorcycle density, dark acrylic license plates, dynamic camera tilt, and extreme tropical glare. We propose an automated anonymization framework tailored for the Kuala Lumpur Road Dataset, captured via a mobile cycling platform at 2 FPS. We document how legacy Haar cascades and YOLOv8 fail under these conditions---generating false positives on background elements while missing rotated or occluded targets. Our architecture resolves this by integrating Grounding DINO---a zero-shot open-set vision-language transformer---with a novel Spatial Vehicle Region of Interest (ROI) Containment Engine. By requiring license plate centroids to reside within validated vehicle boundaries, the pipeline suppresses environmental false positives while automatically obfuscating faces, heads, and license plates. An initial evaluation on 1,266 frames demonstrates a $\sim$95\% success rate, with remaining failures restricted to small, heavily occluded, oblique, or ambiguous targets. Coupled with temporal persistence mechanisms and an automated quality-control auditor, the framework minimizes privacy-related false negatives while preserving scene context for downstream vision tasks. While formal legal compliance depends on broader governance procedures, this publicly available pipeline and demonstration notebook provide an auditable preprocessing stage for privacy-aware dataset curation.