发表机构
University of British Columbia; Google; Imperial College London; Vital Mechanics Research(不列颠哥伦比亚大学; 谷歌; 帝国理工学院; Vital Mechanics Research)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出emg2face,利用高密度表面肌电信号(HD-sEMG)在面部遮挡场景下实现非光学表情捕捉,通过亚毫秒同步与分阶段拟合,训练时空网络以100Hz预测融合变形参数,支持实时动画生成。
AI 中文摘要
面部运动传达了微妙而重要的信息,对于人类社会交流至关重要。当面部被头戴式设备(如VR头显)遮挡时,光学面部捕捉方法难以甚至无法使用。即使视线清晰,此类方法也会引发隐私担忧,并且需要将摄像头和照明从面部偏移安装的头戴式捕捉装置。我们证明,高密度表面肌电信号(HD-sEMG)提供了一种可行的非光学替代方案,解决了这些挑战。我们使用两个纺织肌电电极网格测量了64个肌电通道,其中32个来自前额(通常被头戴式设备遮挡),32个来自面部侧面。肌电数据以2048Hz的采样率数字化并进行滤波。同时记录面部运动,并使用MediaPipe的Face Landmarker估计478个三维面部地标。此类多模态记录中的一个主要挑战是同步肌电和视频数据,两者具有不同的采样频率和独立的时钟。我们开发了一种使用模拟音频突发的新型同步方法,能够实现亚毫秒级同步。我们还开发了一种分阶段拟合方法,将最近的高分辨率参数化头部模型(GNM,具有253个身份融合变形和383个表情融合变形)拟合到MediaPipe地标上,同时参与者做出不同的面部表情。我们训练了一个深度神经网络,包括每个网格的空间编码器,后接膨胀时间卷积网络(TCN),以100Hz的频率从HD-sEMG信号预测融合变形参数。训练完成后,该网络可以仅从HD-sEMG记录预测表情融合变形。输出可以使用标准的实时融合变形动画方法进行渲染。我们使用25名参与者的记录演示了这些方法,并直接将表情迁移到各种人脸和非人类角色上。
英文摘要
Facial movements convey subtle and important information that is critical for human social communication. Optical methods for face capture are difficult or impossible to use when the face is occluded by head-mounted devices (HMDs), such as VR headsets. Even with a clear line of sight, such methods raise privacy concerns and require head-mounted capture rigs that offset cameras and lighting from the face. We show that high-density surface electromyography (HD-sEMG) provides a viable non-optical alternative that addresses these challenges. We measured 64 EMG channels, using two textile EMG grids, with 32 from the forehead (typically occluded by an HMD) and 32 from the side of the face. EMG data were digitized at 2048 Hz and filtered. Facial movements were simultaneously recorded and used to estimate 478 3D facial landmarks using MediaPipe's Face Landmarker. A major challenge in such multimodal recordings is synchronizing EMG and video data, which have different sampling frequencies and independent clocks. We developed a novel synchronization method using analog audio bursts that is capable of sub-millisecond synchronization. We also developed a staged fitting method that fits a recent high-resolution parametric head model (GNM), with 253 identity blendshapes and 383 expression blendshapes, to the MediaPipe landmarks as participants performed different facial expressions. We trained a deep neural network comprising per-grid spatial encoders followed by a dilated temporal convolutional network (TCN) to predict blendshape parameters from HD-sEMG signals at 100 Hz. Once trained, the network can predict expression blendshapes solely from HD-sEMG recordings. The output can be rendered using standard real-time blendshape animation methods. We demonstrate the methods using recordings from 25 participants, and direct expression transfer to a variety of human faces and non-human characters.
Comments12 pages plus supplmentary material