发表机构
Simon Fraser University; Adobe(西蒙弗雷泽大学; 奥多比)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
我们提出模态动态字体设计,利用字形振动模态驱动动画,在保持可读性同时表达语义,实现更平滑、更少撕裂的运动,并获人类评估者青睐。
AI 中文摘要
我们提出了模态动态字体设计(modal kinetic typography),该方法在保持矢量字形可读性的同时,通过动画形式表达语义概念。我们的核心思想是从字形的自然振动模态中构建运动。具体而言,由矢量轮廓组装而成的有限元特征值问题可产生字形最柔和的模态,分别针对整个字母及其各个部分,使其能够弯曲。该问题的零能量解,即刚体平移和旋转,以闭式形式应用于每个部分,使各部分也能作为整体块移动。为了对字形进行动画处理,一个冻结的视频扩散模型仅监督模态的振幅和相位。我们的模态方法解决了先前工作的两个弱点。在视频分数蒸馏(SDS)下的自由形式点优化会沿着噪声梯度独立移动每个点和每一帧,撕裂轮廓并导致抖动。相比之下,我们的模态沿轮廓平滑,并由少数整周期谐波驱动,从而将这些梯度限制为平滑、无缝循环的运动。另一方面,结构化替代方案依赖于来自类别特定先验的骨架或关键点,而我们的模态来自字形本身;唯一的先验是一个由语言模型为整个字母表一次性生成的、列出每个字母运动部分的列表。在模态动态字体设计中,形状和运动通过构造而解耦:一个基础轮廓被雕刻以趋向概念,而模态驱动无法改变它,因此字母也可以在不重塑的情况下进行动画。我们的方法在概念对齐相当或更优的情况下,比Dynamic Typography和AniClipart产生更清晰、更平滑的运动,且比Dynamic Typography更少出现字形撕裂,并受到人类评估者的青睐,包括在仅靠运动传达概念的冻结形状设置中。我们的结果也被人类评估者认为优于Astra(GPT-6)。
英文摘要
We introduce modal kinetic typography, which animates a vector glyph to express a semantic concept while keeping it legible. Our key idea is to build motion from the glyph's natural vibration modes. Specifically, a finite-element eigenproblem assembled from the vector outline yields the glyph's softest modes, for the whole letter and for each of its parts, allowing it to bend. The problem's zero-energy solutions, i.e., rigid translations and rotations, are applied in closed form to each part, allowing parts to also move as blocks. To animate the glyph, a frozen video diffusion model supervises only the modes' amplitudes and phases. Our modal approach addresses two weaknesses of prior work. Free-form point optimization under video score distillation (SDS) moves each point and frame independently along noisy gradients, tearing the outline and causing jitter. In contrast, our modes are smooth along the outline and driven by a few whole-cycle harmonics, which restricts these gradients to smooth, seamlessly looping motion. On the other hand, structured alternatives rely on skeletons or keypoints from category-specific priors, whereas our modes come from the glyph itself; the only prior is a list naming each letter's moving parts, generated once for the whole alphabet by a language model. In modal kinetic typography, shape and motion are disentangled by construction: a single base outline is sculpted toward the concept, and the modal drive cannot alter it, so a letter can also be animated without being reshaped. Our method produces more articulated and smoother motion than Dynamic Typography and AniClipart at comparable or better concept alignment, with less glyph tearing than Dynamic Typography, and is preferred by human raters, including in a frozen-shape setting where motion alone must carry the concept. Our results were also preferred over Astra (GPT-6) by human raters.
CommentsProject page: https://strikeachordkt.github.io/strikeachord/