发表机构
Keio University; JST Presto(庆应义塾大学; 日本科学技术振兴机构先驱计划)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对事件相机第一人称手部重建中背景干扰和手部实例缺失问题,提出EventEgoHands++框架,含手部检测器与自适应注意力,并构建最大真实数据集EEH-R,性能优于基线。
AI 中文摘要
3D手部网格重建对于包括人机交互和AR/VR在内的下游应用而言,是一项具有挑战性但又至关重要的任务。尽管传统相机已被广泛用于该任务,但依赖传统相机的方法在低光环境和严重运动模糊下表现不佳。为解决这些局限性,事件相机因其高动态范围和高时间分辨率而近来受到关注。然而,将事件相机应用于第一人称视角手部重建仍具挑战性,因为佩戴者的运动会产生密集的背景事件,从而掩盖手部特定信号。尽管首个基于事件的第一人称方法通过手部分割缓解了这一问题,但其二值手部掩码无法区分左手和右手。因此,该模型缺乏实例级手部信息,即使仅存在一只手或两只手都不存在时,也会预测出两只手。这一局限性导致手间关系错误和重建精度下降。在本文中,我们提出了EventEgoHands++,一个用于从第一人称视角进行基于事件的3D手部网格重建的框架。所提方法包含一个手部检测器,用于估计左右手的实例级边界框和掩码。此外,我们引入了自适应注意力,它根据这些检测结果动态门控注意力,以准确学习双手之间的空间关系和相互交互。为训练和评估我们的框架,我们扩展了合成N-HOT3D数据集,并新构建了EEH-R,这是迄今为止最大的真实世界基于事件的第一人称手部数据集,包含约100万帧在包括低光条件在内的环境中捕获的标注帧。在合成和真实数据集上的大量实验表明,我们的方法始终优于基线方法。
英文摘要
3D hand mesh reconstruction is a challenging yet essential task for downstream applications, including human-robot interaction and AR/VR. Although conventional cameras have been widely adopted for this task, methods that rely on them struggle in low-light environments and under severe motion blur. To address these limitations, event-based cameras have recently attracted attention for their high dynamic range and high temporal resolution. However, applying event cameras to egocentric hand reconstruction remains challenging because camera wearer's motion produces dense background events that obscure hand-specific signals. Although the first egocentric event-based approach mitigates this issue using hand segmentation, its binary hand mask does not distinguish between left and right hands. As a result, the model lacks instance-level hand information and predicts both hands even when only one or neither hand is present. This limitation leads to incorrect inter-hand relationships and degraded reconstruction accuracy. In this paper, we propose EventEgoHands++, a framework for event-based 3D hand mesh reconstruction from an egocentric viewpoint. The proposed method incorporates a Hand Detector that estimates instance-level bounding boxes and masks for both the left and right hands. Moreover, we introduce Adaptive Attention, which dynamically gates the attention based on these detection results to accurately learn the spatial relationship and mutual interactions between the hands. To train and evaluate our framework, we extend the synthetic N-HOT3D dataset and newly construct EEH-R, the largest real-world event-based egocentric hand dataset to date, comprising approximately 1M annotated frames captured in environments including low-light conditions. Extensive experiments on both synthetic and real datasets demonstrate that our method consistently outperforms the baselines.
CommentsAccepted to IEEE Access. Project Page: https://ryhara.github.io/EventEgoHandsV2/