arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00064cs.CV

可达性并非泛化:理解装配动作识别中的动词-名词分解

Reachability Is Not Generalization: Understanding Verb--Noun Decomposition in Assembly Action Recognition

Changyi Li, Yu Xiao

首次发表
浏览论文内容

中文总结 AI 辅助

本文系统分析装配动作识别中动词-名词分解的组合泛化,发现其仅部分超越原子上限,受共现结构、词汇不对称和组件纠缠限制,并提出可移植诊断框架。

中文摘要 AI 辅助

装配动作具有组合性:它们将操作与部件或工具相结合。在部署中,系统通常会遇到熟悉组件的新组合,然而原子动作分类器在构造上会将每个未见组合的概率恰好赋值为零。主流解决方案是动词-名词分解,即分别预测组件,再将其重新组合以达到未见动作。尽管该方法被广泛采用,但在组合偏移下分解如何泛化仍知之甚少。我们对三个装配数据集(MECCANO、HAViD和IMPACT)上的动词-名词分解进行了系统分析。尽管分解突破了原子上限,但其泛化仅部分超越该上限。未见组合的性能仍与训练数据的共现结构密切相关,表明观察到的增益大部分源于组合空间内密集支持区域的插值,而非不受约束的重新组合。跨数据集,失败一致地集中在词汇量较大的组件上,而IMPACT以动词为主的词汇将瓶颈从名词转向动词。我们进一步表明,共享编码器训练引入了组件纠缠,鼓励依赖对未见组合迁移性差的共现模式,其调和均值落后于独立重组最多达6.0倍。综合来看,这些发现解释了为何分解在实践中仅实现部分组合泛化。通过将原始支持、词汇不对称和组件纠缠识别为相互关联的错误来源,我们提供了一个可移植的诊断框架,用于研究超越聚合准确性的组合识别。代码:此https URL。

英文摘要

Assembly actions are compositional: they combine a manipulation with a part or tool. In deployment, systems routinely encounter novel combinations of familiar components, yet an atomic action classifier assigns every unseen combination exactly zero probability by construction. The prevailing solution is verb--noun decomposition, which predicts components separately and recombines them to reach unseen actions. While widely adopted, how decomposition generalizes under compositional shift remains poorly understood. We present a systematic analysis of verb--noun decomposition across three assembly datasets (MECCANO, HAViD, and IMPACT). Although decomposition escapes the atomic ceiling, its generalization extends only partially beyond it. Unseen-composition performance remains strongly tied to the co-occurrence structure of the training data, indicating that much of the observed gain arises from interpolation within densely supported regions of the compositional space rather than from unconstrained recombination. Across datasets, failures consistently concentrate on the larger-vocabulary component, and IMPACT's verb-heavy vocabulary reverses the bottleneck from nouns to verbs. We further show that shared-encoder training introduces component entanglement, encouraging reliance on co-occurrence patterns that transfer poorly to unseen compositions and trailing independent recombination by up to $6.0\times$ in harmonic mean. Taken together, these findings explain why decomposition achieves only partial compositional generalization in practice. By identifying primitive support, vocabulary asymmetry, and component entanglement as connected sources of error, we provide a portable diagnostic framework for studying compositional recognition beyond aggregate accuracy. Code: https://github.com/hisalaheiyo/assembly.

发表机构

  • Aalto University(阿尔托大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑