能量比例型多模态上下文感知视觉物联网节点
An Energy-Proportional Multimodal and Context-Aware Vision IoT Node
浏览论文内容
中文总结 AI 辅助
该研究提出一种异构多模态双摄像头架构的视觉物联网节点,采用 TinyissimoYOLOv12 模型,实现节能的始终开启视觉监测,在低功耗下完成目标检测与遥测,可支持三个月运行寿命。
中文摘要 AI 辅助
尽管近年来 TinyML 的进展大幅降低了设备端视觉流水线的计算复杂度,但图像采集仍是系统级能耗和内存占用的主要贡献因素。在支持视觉功能的物联网平台中,图像传感器的能耗可与推理引擎相当,从而抵消了算法效率提升的效果。因此,当前设计面临一个根本性的权衡:持续且始终开启的感知会产生过高的能耗,而激进的 duty cycling(占空比循环)则会增加延迟并可能错过瞬态事件。本研究提出了一种能量比例型、上下文感知的视觉物联网节点,通过异构多模态双摄像头架构解决这一挑战。该架构将检测与识别解耦,结合以异步方式工作在节能的始终开启运动唤醒模式下的事件式成像仪,以及 RGB 成像仪。在低功耗微控制器上部署了新型 TinyissimoYOLOv12,用于高效且准确的目标检测。仅在稀疏视觉触发时激活高功耗的图像采集和处理阶段,所提架构提升了效率并降低了延迟,消除了冗余感知,同时保持了连续监测覆盖范围。实验结果显示,其能耗仅为 222μWh。在运动触发时,系统完成完整的感知到报告周期——RGB 采集、80 类目标检测以及 LoRa 遥测,总能耗为 28.7mJ。该网络在模型规模为 100 万参数时,达到了最高 32.3% 的 mAP。在每日活动率为 1% 的情况下,该平台使用 1.85Wh 电池可实现三个月的运行寿命,通过自主边缘智能实现了可放置后即忘场景下的始终开启视觉监测。
英文摘要
While recent advancements in TinyML have significantly reduced the computational complexity of on-device vision pipelines, image acquisition remains a dominant contributor to system-level energy consumption and memory footprint. In vision-enabled IoT platforms, the image sensor consumes energy comparable to the inference engine, thereby offsetting algorithmic efficiency gains. Consequently, current designs face a fundamental trade-off: continuous and always-on sensing incurs prohibitive energy consumption, whereas aggressive duty cycling increases latency and risks missing transient events. This work presents an energy-proportional, context-aware vision IoT node that addresses this challenge through a heterogeneous multimodal dual-camera architecture. Detection and recognition are decoupled by combining an event-based imager operating asynchronously in an energy-efficient always-on wake-on-motion mode together with an RGB imager. Deployed on a low-power microcontroller, a novel TinyissimoYOLOv12 is introduced for efficient and accurate object detection. By activating the high-power image acquisition and processing stages only upon sparse visual triggers, the proposed architecture improves efficiency and latency, eliminating redundant sensing while maintaining continuous monitoring coverage. Experimental results demonstrate an energy consumption of only 222$μ$Wh. Upon a motion trigger, the system completes a full sense-to-report cycle-RGB acquisition, object detection across 80 classes, and LoRa telemetry-with a total energy consumption of 28.7mJ. The network achieves up to 32.3% mAP with a model size of 1 million parameters. At a 1% daily activity ratio, the platform achieves a three-month operational lifetime with a 1.85Wh battery, enabling always-on visual monitoring in a place-and-forget scenario through autonomous edge intelligence.