arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11434cs.CLcs.CV

通过多模态基于人类反馈的强化学习偏好对齐实现汉喃手稿到现代越南语的直接图像翻译

Direct Image-to-Modern Vietnamese Translation of Han-Nom Manuscripts via Multimodal RLHF Preference Alignment

Thi Kim Trang Vo, Nghia Hieu Nguyen, Ha Minh Tan

首次发表
浏览论文内容

中文总结 AI 辅助

针对汉喃手稿到现代越南语翻译的挑战,提出多模态基于人类反馈的强化学习偏好对齐框架,结合四个流生成越南语,比较PPO、DPO和KTO,结果显示各有优劣,且多模态偏好优化能补充监督学习提升翻译质量。

中文摘要 AI 辅助

将汉喃手稿翻译成现代越南语具有挑战性,因为历史页面常退化、文字包含罕见表意字符且平行监督有限。我们提出了一个多模态基于人类反馈的强化学习偏好对齐框架,该框架根据手稿图像和对齐的汉喃源文本生成越南语。模型结合了四个流:用于视觉特征的CLIP ViT-L/14@336、用于汉喃表示的bert-base-chinese、用于越南语表示的vinai/phobert-base和T5-small编码器状态。模态特定投影和融合块将得到的2048维拼接压缩为共享的512维表示。从相同的监督微调策略开始,我们在匹配的工作级宏观平均评估下比较了近端策略优化(PPO)、直接偏好优化(DPO)和知识蒸馏优化(KTO)。DPO在BLEU-4、ROUGE-L、BERTScore、语义相似度、字符错误率(CER)、词错误率(WER)和令牌准确率方面表现最佳,而PPO获得了最高的精确率、召回率和F1值。KTO通过其期望-不期望效用目标保持竞争力。所有偏好对齐策略都提高了可用于监督微调基线的BLEU-4和语义相似度分数。这些结果表明,多模态偏好优化通过提高低资源历史翻译中的词汇和语义质量来补充监督学习。

英文摘要

Translating Han-Nom manuscripts into modern Vietnamese is challenging because historical pages are often degraded, the script contains rare logographic characters, and parallel supervision is limited. We propose a multimodal RLHF preference-alignment framework that conditions Vietnamese generation on manuscript images and aligned Han-Nom source text. The model combines four streams: CLIP ViT-L/14@336 for visual features, bert-base-chinese for Han-Nom representations, vinai/phobert-base for Vietnamese representations, and T5-small encoder states. Modality-specific projections and a fusion block compress the resulting 2,048-dimensional concatenation into a shared 512-dimensional representation. Starting from the same supervised fine-tuned policy, we compare PPO, DPO, and KTO under matched work-level macro-averaged evaluation. DPO achieves the best BLEU-4, ROUGE-L, BERTScore, semantic similarity, CER, WER, and token accuracy, whereas PPO obtains the highest precision, recall, and F1. KTO remains competitive through its desirable-undesirable utility objective. All preference-aligned policies improve the BLEU-4 and semantic-similarity scores available for the SFT baseline. These results indicate that multimodal preference optimization complements supervised learning by improving lexical and semantic quality in low-resource historical translation.

补充信息

↑