面向具身主动视觉的、由人因因素引导的视觉空间认知复杂度模型
A Human-Factors Guided Cognitive Model of Visuospatial Complexity in Embodied Active Vision
查看机构详情
- Lund University(隆德大学)
- Örebro University(厄勒布鲁大学)
- Constructor University Bremen(不来梅 constructor 大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究提出一种多模态复杂度分析框架,以具身认知和主动视觉理论为基础,结合人因因素,为日常驾驶场景下视觉空间复杂度的表征及相关数据集构建提供理论支撑。
中文摘要 AI 辅助
我们提出了一种用于分析多模态数据的新型框架,涵盖视觉、听觉和空间刺激,突出动态自然环境下具身感知与交互中复杂度的作用。该框架以具身认知和主动视觉理论为基础,认为具身感知复杂度源于智能体与环境的动态互动,必须作为定性与定量属性的结合进行整体分析,例如涉及视觉空间和听觉特征的属性。在先前视觉复杂度研究的基础上,我们将其扩展为各类复杂度属性的分类,包括定量、结构、动态、听觉和交互属性,这些属性共同表征多模态复杂度。我们展示了该模型如何为表征视觉空间复杂度及其交互方面提供理论框架,尤其在日常驾驶场景中。我们还讨论了该模型在创建和评估基准数据集(例如驾驶领域)方面的实际应用,这些数据集以认知人因因素为核心,以及用于从视觉感知研究视角系统探究视觉空间复杂度对人类主动视觉影响的应用。该框架为从以人为中心的角度自动分析3D动态环境中的复杂度奠定了基础,作为语义模板,用于结合认知人因因素的分类重点对视觉空间复杂度进行可解释的计算分析。
英文摘要
We propose a novel framework for the analysis of multimodal data -- encompassing visual, auditory, and spatial stimuli -- foregrounding the role of complexity in embodied perception and interaction in dynamic, naturalistic settings. Grounded in theories of embodied cognition and active vision, we argue that embodied perceptual complexity emerges from an agent's dynamic engagement with the environment and must be analyzed holistically, as a combination of qualitative and quantitative attributes pertaining to, for instance, visuospatial and auditory features. Building on previous work on visual complexity, we expand this into a categorization of diverse complexity attributes -- quantitative, structural, dynamic, auditory, and interactional -- that together characterize multimodal complexity. We demonstrate how this model provides a theoretical framework for characterizing aspects of visuospatial complexity and their interactions, specifically in the context of everyday driving. We also discuss practical applications of the proposed model for creating and evaluating benchmark datasets (e.g., in driving) that centralize cognitive human factors, as well as applications aimed at systematically investigating the effect of visuospatial complexity on human active vision from the viewpoint of visual perception research. The proposed framework lays the foundation for automated methods that interpret complexity in 3D dynamic environments from a human-centered perspective, serving as a semantic template for explainable computational analysis of visuospatial complexity with a categorical focus on cognitive human factors.