通过可微物理渲染连接在线与离线手写文字
Bridging Online and Offline Handwriting via Differentiable Physical Rendering
浏览论文内容
中文总结 AI 辅助
该研究提出统一在线-离线手写生成框架,通过物理画笔模型与可微渲染模块解决两类手写生成范式的缺陷,经实验验证实现了结构与视觉保真度。
中文摘要 AI 辅助
逼真的手写文字生成在字体设计、生物特征认证、机器书法等众多应用中发挥重要作用。现有方法通常分为两个独立范式:在线方法估计手写轨迹,离线方法合成逼真的手写图像。在线模型能捕捉结构与时间动态,但常缺乏细粒度纹理;离线模型能复现逼真外观,但会丢失笔画顺序。然而,统一在线与离线模型仍具挑战性,原因包括:1)缺乏将运动学与像素级外观关联的显式物理模型;2)缺乏配对的轨迹-图像数据集。此外,实现端到端学习需要跨运动与外观域的可微渲染过程。为应对这些挑战,我们提出一种紧凑的物理画笔模型,用于连接笔画动态与视觉外观,同时提出可微渲染模块,将笔画轨迹转换为风格化图像。通过整合这些组件,我们提出一种基于可微画笔渲染的统一在线-离线手写生成框架,该框架包含四个核心模块:1)文本到笔画生成器,基于给定文本和风格图像预测目标笔画;2)画笔参数观测器,从风格参考中提取画笔模型参数;3)可微画笔渲染器,将笔画序列与物理画笔参数映射为手写图像;4)零样本图像细化器,通过扩散模型细化渲染图像。大量实验与真实世界机器书法演示验证了我们的方法,实现了结构与视觉保真度。
英文摘要
Realistic handwritten text generation plays an important role in numerous applications, such as font design, biometric authentication, and robotic calligraphy. Existing methods are typically divided into two independent paradigms: online approaches that estimate handwriting trajectories and offline approaches that synthesize realistic handwriting images. While online models capture structural and temporal dynamics, they often lack fine-grained textures, whereas offline models reproduce realistic appearance but discard stroke order. However, unifying online and offline models remains challenging due to (1) the lack of an explicit physical model linking stroke kinematics to pixel-level appearance and (2) the absence of paired trajectory-image datasets. Moreover, enabling end-to-end learning requires a differentiable rendering process across motion and appearance domains. To address these challenges, we propose a compact physical brush model that bridges stroke dynamics and visual appearance, together with a differentiable rendering module that converts stroke trajectories into stylized images. By integrating these components, we propose a unified online-offline handwriting generation framework via differentiable brush rendering. The proposed framework consists of four core modules: 1) a text-to-stroke generator that predicts the target stroke conditioned on the given text and style image, 2) a brush parameter observer that extracts brush model parameters from style references, 3) a differentiable brush renderer that maps a stroke sequence and physical brush parameters into a handwritten image, and 4) a zero-shot image refiner that refines rendered images via diffusion models. Extensive experiments and real-world robotic calligraphy demonstrations validate our approach, achieving both structural and visual fidelity.