发表机构
Case Western Reserve University; University of Louisiana at Lafayette; University of Southern California; The University of Texas at Austin; Apple; University of Wisconsin–Madison(凯斯西储大学; 路易斯安那大学拉法叶分校; 南加州大学; 德克萨斯大学奥斯汀分校; 苹果公司; 威斯康星大学麦迪逊分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
OASIS提出硬件感知的分布式感内视觉压缩框架,通过轻量编码器生成紧凑表示,结合量化/HDC实现18,816倍通信减少和2-4.5倍能耗降低。
AI 中文摘要
感内计算通过在传感器附近执行早期处理来降低传输高分辨率图像数据的成本。然而,与CMOS图像传感器(CIS)集成的逻辑芯片在计算和内存方面受到严格限制,限制了传统的深度神经网络划分。我们提出了OASIS,一个分布式感内视觉框架,使用轻量级编码器在片外传输前生成紧凑的、任务相关的表示。编码器使用任务、熵和重建目标进行端到端训练,而解码器仅在训练期间使用。OASIS支持两条互补的部署路径。第一条路径应用4比特量化和霍夫曼编码,同时保留分类和密集预测任务所需的空间结构。第二条路径使用基于Sobol的超维计算(HDC)将编码器潜在表示转换为固定维度的二进制超向量,用于联想记忆分类。对于基于SwinViT的VWW模型,将$3\ imes3\ imes8$的潜在表示映射到64维超向量,相对于128维配置,提供了额外的$1.77\ imes$通信减少,且准确率损失不到一个百分点,与原始8比特图像传输相比,实现了总计$18{,}816\ imes$的减少。我们在AMD Xilinx Zynq UltraScale+ FPGA上实现了数字近传感器流水线,并使用直接板级功率测量和Vivado实现后分析,结合电路仿真的CIS模型和7纳米ASIC投影进行了表征。在视觉唤醒词分类、手部跟踪和眼部跟踪任务中,OASIS将系统总能耗降低了约$2\ imes$-$4.5\ imes$,同时保持有竞争力的准确率,展示了通信高效的感内视觉的实用硬件-算法协同设计路径。
英文摘要
In-sensor computing reduces the cost of transmitting high-resolution image data by performing early-stage processing near the sensor. However, the logic chip integrated with a CMOS image sensor (CIS) is tightly constrained in compute and memory, limiting conventional deep neural network partitioning. We present OASIS, a distributed in-sensor vision framework that uses a lightweight encoder to generate compact, task-relevant representations before off-chip transmission. The encoder is trained end-to-end using task, entropy, and reconstruction objectives, while the decoder is used only during training. OASIS supports two complementary deployment paths. The first applies 4-bit quantization and Huffman coding while preserving the spatial structure required by classification and dense-prediction tasks. The second uses Sobol-based hyperdimensional computing (HDC) to transform the encoder latent into a fixed-dimensional binary hypervector for associative-memory classification. For the SwinViT-based VWW model, mapping a $3\times3\times8$ latent to a 64-dimensional hypervector provides an additional $1.77\times$ communication reduction with less than one percentage point of accuracy loss relative to the 128-dimensional configuration, yielding an overall $18{,}816\times$ reduction compared with raw 8-bit image transmission. We implement the digital near-sensor pipeline on an AMD Xilinx Zynq UltraScale+ FPGA and characterize it using direct board-level power measurements and Vivado post-implementation analysis, together with circuit-simulated CIS models and a 7-nm ASIC projection. Across visual wake-word classification, hand tracking, and eye tracking, OASIS reduces total system energy by approximately $2\times$-$4.5\times$ while maintaining competitive accuracy, demonstrating a practical hardware-algorithm co-design path for communication-efficient in-sensor vision.
CommentsUnder submission at IEEE Transactions on Emerging Topics in Computing