arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37287cs.CVcs.AI

VISTA-Bench:基于图像特定评分标准的多语言图像翻译基准

VISTA-Bench: Benchmarking Multilingual Image Translation with Image-Specific Rubrics

  • Tencent(腾讯)

机构由 AI 辅助整理,请以论文原文为准。

Bo Lv, Mao Zheng, Zheng Li, Fangxu Liu, Mingrui Sun, Tao Chen

AI总结:

本文提出VISTA-Bench,一个覆盖22种语言和10个领域的多语言图像翻译基准,采用图像特定评分标准协议,评估16个主流模型,以系统衡量翻译质量与视觉信息保持能力。

AI中文摘要:

图像翻译是多模态模型在多语言应用中的一项基本能力,需要跨语言的视觉理解和意义保持。然而,现有基准的语言覆盖范围有限,且往往缺乏明确的图像特定评估标准,使得全面评估该能力变得困难。为了系统性地评估这一能力,我们引入了VISTA-Bench,覆盖22种语言和10个领域,并开发了一种图像特定的评分标准评估协议。该基准结合了语言和场景覆盖的采样,以及模型辅助、人工验证的注释,将相关文本分组为连贯的语义单元,并提供多语言参考翻译。评分标准规定了必要内容、语义关系和可接受的翻译变体,从而为翻译质量以及视觉和知识依赖信息的保持分别生成基于输出的分数。我们对16个主流模型进行了广泛评估,包括12个多模态模型和4个文本输入模型,并跨语言、领域和评估维度提供了系统分析。

英文摘要:

Image translation is a fundamental capability of multimodal models for multilingual applications, requiring visual understanding and meaning preservation across languages. However, existing benchmarks have limited language coverage and often lack explicit image-specific evaluation criteria, making it difficult to comprehensively assess this capability. To systematically evaluate this capability, we introduce VISTA-Bench, covering 22 languages and 10 domains, and develop an image-specific rubric evaluation protocol. The benchmark combines sampling for language and scenario coverage with model-assisted, human-verified annotations that group related text into coherent semantic units and provide multilingual reference translations. The rubrics specify essential content, semantic relations, and acceptable translation variants, yielding separate output-based scores for translation quality and the preservation of visual and knowledge-dependent information. We conduct extensive evaluations of 16 mainstream models, including 12 multimodal models and four text-input models, and provide systematic analyses across languages, domains, and evaluation dimensions.

↑