TransAnyText:通过结构化视觉生成实现电商图像中任意文本的翻译
TransAnyText: Translating Arbitrary Text in E-commerce Images via Structured Visual Generation
浏览论文内容
中文总结 AI 辅助
本文提出TransAnyText框架,将电商图像文本翻译转化为生成可渲染HTML补丁,经三阶段后训练优化,在多基准上表现优于对比方法,为跨境电商图像翻译提供有效可控的解决方案。
中文摘要 AI 辅助
跨境电商图像翻译对全球零售至关重要,需针对产品图像、横幅、详情页生成不同语言版本。现有方法难以同时实现准确翻译、忠实视觉特征保留和易编辑输出。为解决这些挑战,本文提出TransAnyText,一种结构化视觉代码框架,将图像文本翻译重新表述为从源图像和目标语言生成可渲染HTML补丁。该框架将语义生成与像素渲染解耦:视觉语言模型(VLM)负责视觉理解、跨语言翻译和结构化视觉生成,扩散模型执行背景修复和像素级细化,随后通过确定性渲染合成最终图像。基于此表述,本文开发了三阶段后训练框架:监督微调(SFT)建立图像到代码的映射,特权差距加权自蒸馏(PWSD)改进风格和布局标记的学习,带可验证奖励的强化学习(RLVR)进一步优化任务级性能。本文还推出TransAnyDataset和TransAnyBench,这是用于电商图像翻译的多语言数据集和基准。大量实验表明,该方法在级联流水线、开源端到端模型和闭源图像编辑系统中表现出竞争力,为跨境电商图像翻译提供了有效、可控且可编辑的解决方案。
英文摘要
Cross-border e-commerce image translation is essential for global retail, where product images, banners, and detail pages need to be produced in different languages. Existing methods struggle to achieve accurate translation, faithful visual identity preservation, and easy-to-edit outputs, simultaneously. To address these challenges, we introduce TransAnyText, a structured visual code framework that reformulates image text translation as generating renderable HTML patches from source images and target languages. Our framework decouples semantic generation from pixel rendering: a vision-language model (VLM) handles visual understanding, cross-lingual translation, and structured visual generation, while a diffusion model performs background inpainting and pixel-level refinement, followed by deterministic rendering to synthesize the final image. Based on this formulation, we develop a three-stage post-training framework, where supervised fine-tuning (SFT) establishes the image-to-code mapping, privilege-gap weighted self-distillation (PWSD) improves the learning of style and layout tokens, and reinforcement learning with verifiable rewards (RLVR) further optimizes task-level performance. We further introduce TransAnyDataset and TransAnyBench, a multilingual dataset and benchmark for e-commerce image translation. Extensive experiments demonstrate competitive performance against cascaded pipelines, open-source end-to-end models, and closed-source image editing systems, providing an effective, controllable, and editable solution for cross-border e-commerce image translation.