arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

轨迹对齐字形渲染增强视频文本编辑

Enhanced Video Text Editing with Trajectory-Aligned Glyph Rendering

Shulian Zhang, Xiangyu Shu, Wenbo Li, Jian Chen, Yong Guo

arXiv 2609.34178首次发表:更新:

发表机构

South China University of Technology; The Chinese University of Hong Kong; Huawei(华南理工大学; 香港中文大学; 华为)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视频扩散模型生成文本笔画错误的问题,提出轨迹对齐字形渲染与深度归一化识别器监督,构建VTEdit基准,实现高文本准确率与背景保持。

AI 中文摘要

视频文本编辑旨在替换或添加视频中的文本,同时保持视频其余部分不变,这要求编辑后的文本在每一帧中正确无误,并与场景连贯移动。尽管视频扩散模型取得了显著进展,但它们在重现精确的笔画结构方面仍存在困难,经常产生乱码或错误的字符,尤其是对于笔画复杂的字符。为解决这一问题,我们提出了一种轨迹对齐字形渲染参考,它根据文本的位置和透视提供明确的逐帧字形指导,以及一种深度归一化识别器特征监督,该监督在冻结的文本识别器的多深度特征上监督生成的文本,并采用逐深度归一化误差,针对扩散损失所忽视的笔画错误。我们进一步构建了VTEdit,一个包含288个真实场景片段和440条标注文本轨迹的基准,涵盖文本替换和文本添加,该基准将公开发布以促进未来研究。在VTEdit上的实验表明,我们的方法在文本准确性和背景保持方面优于图像文本编辑方法、视频编辑方法和商业模型,达到了0.9408的句子准确率,并在用户研究中获得了最高偏好。

英文摘要

Video text editing aims to replace or add text in a video while keeping the rest of the video unchanged, which requires the edited text to be correct in every frame and to move coherently with the scene. Despite the remarkable progress of video diffusion models, they struggle to reproduce exact stroke structures and often produce garbled or wrong characters, especially for characters with complex strokes. To address this, we propose a trajectory-aligned glyph rendering reference that provides explicit per-frame glyph guidance following the position and perspective of the text, and a depth-normalized recognizer feature supervision that supervises the generated text on multi-depth features of a frozen text recognizer with per-depth normalized errors, targeting stroke errors overlooked by the diffusion loss. We further build VTEdit, a benchmark of 288 real-scene clips with 440 annotated text trajectories covering text replacement and text addition, which will be publicly released to facilitate future research. Experiments on VTEdit show that our method outperforms image text editing methods, video editing methods, and commercial models in text accuracy and background preservation, achieving a sentence accuracy of 0.9408, and receives the highest preference in a user study.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑