MHE-Former:基于熵最大化的多假设Transformer用于3D网格恢复
MHE-Former: Multi-Hypothesis Transformers via Entropy Maximization for 3D Mesh Recovery
- Communication University of China(中国传媒大学)
- Nankai University(南开大学)
- Beihang University(北京航空航天大学)
- East China Normal University(华东师范大学)
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对单目3D网格恢复中的遮挡和模糊问题,提出基于熵最大化的多假设Transformer框架MHE-Former,结合假设选择机制,实现高精度且多样的恢复,达到最先进性能。
AI中文摘要:
单目3D手部和身体网格恢复经常遭受严重的遮挡和模糊问题。传统的确定性方法通常回归单个最优解,导致过度自信的预测。在本文中,我们引入了一种探索-利用范式,用于通过多假设学习和选择来处理模糊的网格恢复。具体来说,在探索阶段,基于我们的概率公式和熵最大化,我们提出了一种新颖的多假设方法,称为MHE-Former。它是一个基于Transformer的多假设框架,确保高训练效率和标签友好性,同时生成合理且多样的假设。在利用阶段,我们提出了假设选择,一种针对多个预测的上下文感知过程。特别是利用VLM强大的视觉理解和推理能力,它允许用户通过额外的证据和自然语言意图选择最合理和期望的估计。大量实验表明,我们的框架在多个数据集上实现了准确性和多样性的最先进性能。用户偏好研究进一步展示了我们假设选择过程的实用性。
英文摘要:
Monocular 3D hand and body mesh recovery often suffers from severe occlusion and ambiguity. Traditional deterministic methods typically regress a single optimal solution, leading to overconfident predictions. In this paper, we introduce an exploration--exploitation paradigm for ambiguous mesh recovery with multi-hypothesis learning and selection. Specifically, during exploration, based on our probabilistic formulation and entropy maximization, we propose a novel multi-hypothesis method referred to as MHE-Former. It is a Transformer-based multi-hypothesis framework, ensuring high training efficiency and label friendliness while generating plausible and diverse hypotheses. During exploitation, we propose Hypothesis Selection, a context-aware process for multiple predictions. Especially leveraging VLM's powerful visual understanding and reasoning capabilities, it allows users to choose the most plausible and desired estimate with additional evidence and natural language intent. Extensive experiments demonstrate that our framework achieves state-of-the-art performance in accuracy and diversity across multiple datasets. The user preference study further shows the practicality of our hypothesis selection process.