arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越关键点提取:结构化图像分类中鲁棒几何特征构建的框架

Beyond Landmark Extraction: A Framework for Robust Geometric Feature Construction in Structured Image Classification

Saravana Mauree, Sakshi Arya

arXiv 2609.00634首次发表:更新:

发表机构

Case Western Reserve University(凯斯西储大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对结构化图像分类问题,提出鲁棒几何特征构建框架,以静态手势识别为案例,发现混合几何表示性能最优,将特征构建视为基础建模决策。

AI 中文摘要

现有结构化图像识别领域的文献大多过度聚焦于分类算法的比较,而非探究何种分类器在预测前应掌握的信息。在手势识别、面部表情分类、医学图像分析等结构化视觉问题中,判别性信息较少存在于单个像素,更多存在于语义部件间的空间关系中。原始像素空间维度高、易受干扰变异影响,且常混淆使视觉任务可解释的几何结构。关键点提取提供了一种降维方式,但本身无法确定保留的信息。本文将关键点后的特征图作为核心分析对象,提出一种系统框架,用于构建和解释源自关键点的表示,作为一种“知情的、基于特征的”降维步骤。以静态手势识别为案例研究,通过扰动和消融实验评估坐标、距离、角度及混合表示。结果表明,视觉可变数据凸显了原始坐标特征与其几何不变对应特征间的显著差距,而混合表示通过结合互补几何组件实现了最强的整体性能。这些发现将特征构建视为一项基础建模决策,并最终提出“分类器应学习何种表示”这一值得探究的问题。用于特征构建和评估的代码可在此 https URL 获取。

英文摘要

Much of the literature on structured image recognition has disproportionately focused on the comparison of classification algorithms. Rather than investigating which classifier performs best, this paper instead asks: what should a classifier know before it ever makes a prediction? In structured vision problems such as gesture recognition, facial expression categorization, and medical image analysis, discriminative information lies less in individual pixels and more in spatial relationships between semantic parts. Raw pixel spaces are high-dimensional, sensitive to nuisance variation, and often obfuscate the geometric structures that make visual tasks interpretable. Landmark extraction provides one form of dimension reduction, but it does not by itself determine the information preserved. This paper studies the post-landmark feature map as the central object of analysis and proposes a systematic framework for constructing and interpreting landmark-derived representations as an, informed, feature-based "dimension reduction" step. Using static hand gesture recognition as a case study, we evaluate coordinate, distance, angle, and hybrid representations through perturbation and ablation experiments. The results show that visually variable data exposes substantial gaps between raw coordinate features and their geometrically invariant counterparts, while hybrid representations achieve the strongest overall performance by combining complementary geometric components. These findings frame feature construction as a fundamental modeling decision and ultimately suggests that the question of what representation should a classifier learn from is one worth asking. The code used for feature construction and evaluation is available at https://github.com/ShivMaureeCWRU/Feature_based_dimension_reduction

CommentsUnder consideration at Pattern Recognition Letters

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑