提高多模态机器翻译中LLMs视觉敏感性的基于度量的损失加权方法
Improving Visual Sensitivity of LLMs on Multimodal Machine Translation with Metric-based Loss Weighting
- Maastricht University(马斯特里赫特大学)
- Warsaw University of Technology(华沙理工大学)
- IDEAS Research Institute(IDEAS研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出基于度量的损失加权方法,利用PCXMI指标识别受益于图像的令牌并提高其损失,增强多模态翻译的视觉敏感性,在CoMMuTE数据集上准确率提升超7个百分点。
AI中文摘要:
多模态机器翻译旨在通过解决歧义来利用非文本模态的额外信号改进翻译。虽然模型通过多模态融合能够接受与源文本相关的图像,但它们可能会忽略这些信息。因此,提高其视觉敏感性仍是一个活跃的研究领域。在这项工作中,我们引入了一种训练方法,即基于度量的损失加权,通过增加对受益于伴随图像的令牌的损失函数来提高翻译的视觉接地性。我们使用点式交叉互信息(PCXMI)指标识别这些令牌,该指标比较模型在有和没有视觉上下文时的输出概率。我们引入了一种基于一致性的PCXMI指标,并通过实验表明,这两种指标结合使用能产生最佳结果。我们通过在三语言方向上微调三个预训练的多模态大语言模型进行图像引导机器翻译任务来评估我们的方法。基于度量的损失加权在CoMMuTE对比数据集上优于其他测试方法,与标准微调相比,准确率提高了最多超过7个百分点,同时保持了强大的通用翻译性能。
英文摘要:
Multimodal Machine Translation aims to incorporate additional signal from non-textual modalities to improve translations by resolving ambiguities. While models, through multimodal fusion, are able to accept images related to the source text, they can ignore this information. Therefore, increasing their visual sensitivity remains an active research area. In this work, we introduce a training method, Metric-based Loss Weighting, that improves visual grounding of translations by increasing the loss function for tokens that benefit from the accompanying image. We identify these tokens using the Point-wise Cross-mutual Information (PCXMI) metric, which compares the model's output probabilities with and without visual context. We introduce a Congruency-based PCXMI metric and experimentally show that both metrics working in combination yield the best results. We evaluate our method by fine-tuning three pretrained Multimodal Large Language Models on the task of Image-guided Machine Translation for three language directions. Metric-based Loss Weighting outperforms other tested methods on the CoMMuTE contrastive dataset, improving accuracy by up to more than 7 percentage points compared to standard fine-tuning, while maintaining strong general translation performance.