发表机构
Tsinghua University; Nanyang Technological University; SparcAI Inc.(清华大学; 南洋理工大学; SparcAI公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Bunraku系统,可从单张插图端到端生成可编辑Live2D角色,构建分层扩散与联合位移预测方法,发布首个相关基准Live2D-Bench及8884模型语料库,解决了Live2D模型手动构建效率低的问题。
AI 中文摘要
Live2D是动漫角色和虚拟化身的主流2D角色动画格式,它将每个角色表示为一组由逐层网格变形驱动的RGBA图层堆叠。尽管其广泛应用于虚拟直播、手机游戏和互动角色创作,但构建Live2D模型仍需数周的手动图层分离、遮挡补全、网格放置和关键帧制作,且现有生成式方法无法端到端生成此类结构化资产。本文提出首个系统,可从单张插图生成Live2D运行时所需的全部结构化信息:有序RGBA图层、每层的变形网格,以及驱动角色动作的参数驱动关键姿势顶点偏移。第一阶段将分层分解转化为感知Live2D的器官级分类下的分层扩散过程,生成带有隐藏区域补全的有序RGBA堆叠。第二阶段仅从图层的Alpha通道为每个图层构建内容一致的三角形网格,随后联合预测所有图层的关键姿势位移场:每个图层的每个顶点为一个token,自注意力跨图层边界,且每个位移被分解为有界方向和对数幅度。联合而非独立预测是使结果成为连贯角色而非各自合理部件的关键,也是本文最大的改进;将网络规模扩大112倍则无增益。在50个未见过的角色上,采用无教师强制的真实生成方式,第二阶段达到每顶点方向余弦0.768(中位数0.828)。由于图层的网格源自其Alpha通道,服装图层可根据自然语言指令重新纹理化,同时网格和预测动画可逐字节复用。本文还贡献了该任务的首个标准化基准Live2D-Bench,以及包含图层和动画监督的8884个模型的Live2D语料库。
英文摘要
Live2D is the dominant 2D character-animation format for anime characters and virtual avatars, representing each character as a stack of RGBA layers driven by per-layer mesh deformation. Despite its wide use in virtual streaming, mobile games, and interactive characters, authoring a Live2D model still demands weeks of manual layer separation, occlusion completion, mesh placement, and keyframing, and no prior generative method produces such a structured asset end-to-end. We present the first system that, from a single illustration, generates all the structured information a Live2D runtime consumes: ordered RGBA layers, a deformation mesh per layer, and the parameter-driven keypose vertex offsets that make the character move. Stage 1 casts layered decomposition as a layered diffusion process under a Live2D-aware organ-level taxonomy, producing an ordered RGBA stack with hidden-region completion. Stage 2 builds a content-conforming triangle mesh for each layer from its alpha channel alone, then predicts the keypose displacement field of all layers jointly: every vertex of every layer is one token, self-attention spans layer boundaries, and each displacement is factorised into a bounded direction and a log-magnitude. Joint rather than independent prediction is what makes the result a coherent character instead of separately plausible parts, and is our largest gain; scaling the network 112x yields none. On 50 held-out characters, under true generation with no teacher forcing, Stage 2 attains a per-vertex direction cosine of 0.768 (median 0.828). Because a layer's mesh derives from its alpha channel, a clothing layer can be re-textured from a natural-language instruction while the mesh and predicted animation are reused byte-for-byte. We further contribute Live2D-Bench, the first standardized benchmark for the task, and an 8,884-model Live2D corpus with layer and animation supervision.
CommentsProject page: https://bunraku-live2d.github.io/