发表机构
Comexp Research Lab(Comexp研究实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出TAPe+ML v3,一种基于主动感知理论的结构化表示多任务视觉系统,以少于10万参数在COCO检测与分割、Imagenette和ImageNet-Real分类上取得高精度,证明结构化输入可降低资源需求。
AI 中文摘要
我们提出了TAPe+ML v3,一种基于TAPe(主动感知理论)的紧凑计算机视觉系统,TAPe是一种在识别之前编码感知元素之间关系的结构化表示。该系统不直接对像素张量进行操作,而是使用共享的TAPe表示和模块化识别架构,用于图像分类、目标检测和实例分割。TAPe+ML v3结合了背景与轮廓处理、局部目标定位、基于原型的分类以及用于专门子模型的协调器。在报告的实验中,它使用的参数少于100,000个。在COCO目标检测上,它获得了84.7的mAP50和65.3的mAP50-95。在COCO实例分割上,它获得了80.7的掩膜mAP50和58.4的掩膜mAP50-95。在分类实验中,在与原始像素基线进行相同训练设置的比较下,它在Imagenette上达到了92%的验证准确率,在ImageNet-Real上达到了89.9%的Top-1准确率。我们还在工业试点中评估了视频场景检测的紧凑性和分布偏移下的适应性。结果表明,将部分建模负担从网络参数转移到结构化输入表示上,可以支持紧凑的多任务视觉系统,并减少数据、内存和计算需求。
英文摘要
We present TAPe+ML v3, a compact computer vision system based on TAPe (Theory of Active Perception), a structured representation that encodes relations among perceptual elements before recognition. Instead of operating directly on pixel tensors, the system uses a shared TAPe representation and a modular recognition architecture for image classification, object detection, and instance segmentation. TAPe+ML v3 combines background and contour processing, local object localization, prototype-based classification, and a coordinator for specialized submodels. Across the reported experiments, it uses fewer than 100,000 parameters. On COCO object detection, it obtains 84.7 mAP50 and 65.3 mAP50-95. On COCO instance segmentation, it obtains 80.7 mask mAP50 and 58.4 mask mAP50-95. In classification experiments, it reaches 92 percent validation accuracy on Imagenette under an identical-training comparison with a raw-pixel baseline, and 89.9 percent Top-1 accuracy on ImageNet-Real. We also evaluate compactness in video scene detection and adaptation under distribution shift in an industrial pilot. The results suggest that shifting part of the modeling burden from network parameters to a structured input representation can support compact multi-task vision systems with reduced data, memory, and compute requirements.
Comments39 pages, 4 figures, 11 tables. Project page: https://ml.comexp.net Corrected the corresponding author's email address