发表机构
Institute for AI Industry Research (AIR), Tsinghua University; Beijing Academy of Artificial Intelligence (BAAI); Institute for Interdisciplinary Information Sciences (IIIS), Tsinghua University; Beihang University(清华大学人工智能产业研究院(AIR); 北京人工智能研究院(BAAI); 清华大学交叉信息研究院(IIIS); 北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在实现对话式仿人机器人面部机构自动合成,提出参数化模板及分层算法,开发对话面部运动合成框架,经多方面实验验证,可快速生成可物理实现的面部机构,推动大规模个性化发展。
AI 中文摘要
仿人机器人的面部是社交互动机器人的核心组成部分,能通过面部表情实现丰富的非语言交流。然而,现有仿人机器人面部通常是定制系统,新面部几何形状需大量手动机械重新设计,大规模个性化成本高且速度慢。本文致力于自动化、可扩展的机械面部合成,引入参数化、连杆驱动的机械面部模板,在此基础上提出分层自动设计算法。还开发了双身份对话面部运动合成框架。通过大量实验验证了系统,包括对面部几何形状自动机构合成的定量评估、与手动机械设计的比较等。
英文摘要
Animatronic faces are a central component of socially interactive robots, enabling rich nonverbal communication through facial articulation. However, state-of-the-art animatronic faces are typically tailored systems: each new facial geometry requires extensive manual mechanical redesign, making large-scale personalization prohibitively slow and costly. In this work, we pursue automated and scalable mechanical face synthesis, aiming to rapidly generate a physically realizable facial mechanism for a wide range of facial geometries. We introduce a parametric, linkage-driven mechanical face template whose topology and actuator layout are explicitly parameterized to support systematic scaling and retargeting across diverse facial morphologies. Building on this template, we propose a hierarchical automatic design algorithm that takes a single 2D portrait as input, reconstructs a target 3D face, and synthesizes a collision-free, manufacturable internal mechanism. The algorithm combines anatomy-guided feasible motion volumes, Action Unit (AU)-derived trajectory-based expressiveness objectives, and a collision-driven outer-loop refinement strategy. Beyond hardware synthesis, we argue that future mechanical faces deployed at scale must engage in bidirectional, multi-turn conversation rather than functioning solely as speaking or listening heads. To this end, we develop a dual-identity conversational facial motion synthesis framework that jointly models speaking and listening behaviors from audio, producing temporally coherent 3D facial motion suitable for physical execution. We validate our system through extensive experiments, including (i) quantitative evaluation of automatic mechanism synthesis across diverse facial geometries, (ii) comparisons against manual mechanical design, (iii) benchmarks on conversational facial motion synthesis and real-time deployment, and (iv) perceptual user studies.
CommentsAccepted by RSS 2026. Project page: https://zzongzheng0918.github.io/automated-facial-mechanisms-synthesis/