发表机构
Zurich University of Applied Sciences (ZHAW); ETH Zürich; Politecnico di Torino; University of Tübingen; Czech Technical University(苏黎世应用科技大学; 苏黎世联邦理工学院; 都灵理工大学; 蒂宾根大学; 捷克技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出轻量主动视觉流水线,通过数字凹视实现高效语义理解,在ADE20K-Object数据集上仅用少量计算即可达到接近基线的准确率,为语义分割提供高效替代方案
AI 中文摘要
密集语义分割算法会在整幅图像上均匀分配计算资源,而不考虑场景复杂度或任务相关性。受生物视觉启发,本研究探索能否通过数字凹视感知更高效地实现语义理解。我们提出一种轻量主动视觉流水线,结合显著性驱动的注视点选择、高分辨率凹视观测、低分辨率上下文信息、语义累积及自适应计算。除传统密集预测指标外,我们采用对象级评估衡量稀疏观测下的语义理解能力。在ADE20K-Object数据集上,单次凹视观测达到基线Top-1准确率的95.9%、Top-3准确率的96.9%,仅需4.7%的计算成本;在场景级,语义累积恢复了基线对象召回率的90.6%,仅用58.6%的计算量。这些结果表明,当选择性分配计算资源时,稀疏观测可实现充分的语义理解,凸显主动视觉是均匀密集处理的高效替代方案,并推动超越传统像素级分割指标的评估协议发展。
英文摘要
Dense semantic segmentation allocates computational resources uniformly across the entire image, regardless of scene complexity or task relevance. Inspired by biological vision, we investigate whether semantic understanding can be achieved more efficiently through digital foveated perception. We introduce a lightweight active-vision pipeline that combines saliency-driven fixation selection, high-resolution foveal observations, low-resolution contextual information, semantic accumulation, and adaptive computation. Beyond conventional dense prediction metrics, we use object-level evaluation to measure semantic understanding under sparse observations. On ADE20K-Object, a single foveated observation achieves 95.9% of the baseline Top-1 accuracy and 96.9% of the baseline Top-3 accuracy while requiring only 4.7% of the computational cost. At the scene level, semantic accumulation recovers 90.6% of the baseline object recall while using 58.6% of the computation. These results suggest that substantial semantic understanding can emerge from sparse observations when computation is allocated selectively, highlighting active vision as an efficient alternative to uniform dense processing and motivating evaluation protocols beyond conventional pixel-wise segmentation metrics.
CommentsAccepted at the 3rd Human-inspired Computer Vision Workshop at ECCV 2026