arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FinMMEval 2026任务1概述:多语言金融多项选择题问答

Overview of FinMMEval 2026 Task 1: Multilingual Financial Multiple-Choice Question Answering

Zhuohan Xie, Yuyang Dai, Rania Elbadry, Vanshikaa Jani, Georgi Georgiev, Dimitar Dimitrov, Fan Zhang, Xueqing Peng, Lingfei Qian, Jimin Huang, Jiahui Geng, Yankai Chen, Ye Yuan, Haolun Wu, Yuxia Wang, Ivan Koychev, Veselin Stoyanov, Mingzi Song, Yu Chen, Xue Liu, Preslav Nakov

arXiv 2607.19856首次发表:更新:

发表机构

Mohamed bin Zayed University of Artificial Intelligence; INSAIT, Sofia University “St. Kliment Ohridski”; University of Arizona; FMI, Sofia University “St. Kliment Ohridski”; The Fin AI; The University of Tokyo; Linköping University; McGill University; Mila, Quebec AI Institute; Meiji Gakuin University(穆罕默德·本·扎耶德人工智能大学; 索非亚大学“圣克莱门特·奥赫里德斯基”INSAIT学院; 亚利桑那大学; 索非亚大学“圣克莱门特·奥赫里德斯基”FMI学院; The Fin AI公司; 东京大学; 林雪平大学; 麦吉尔大学; 魁北克Mila人工智能研究所; 明治学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FinMMEval 2026任务1测试多语言金融多项选择题问答,最终测试集含800题,每种语言200题,按准确率排名。最高准确率在92.0%(印地语)到97.5%(英语和阿拉伯语)之间,领先团队相近。系统采用检索增强等多种方法。

AI 中文摘要

FinMMEval 2026任务1评估英语、中文、阿拉伯语和印地语的多语言金融多项选择题问答。该任务测试系统能否跨语言和脚本选择涉及领域术语、数值解释和概念性金融推理的金融问题的正确答案。最终测试集包含800个问题,每种语言200个;提交时隐藏黄金答案,每种语言按准确率独立排名。最终排行榜包含13份英语、11份中文、11份阿拉伯语和10份印地语的排名提交。最高准确率从印地语的92.0%到英语和阿拉伯语的97.5%不等,所有四种语言中领先团队相近。记录的系统使用了检索增强、直接答案选项评分、特定语言提示、选择性自一致性、置信度检查和基于大语言模型的审查阶段。

英文摘要

FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering in English, Chinese, Arabic, and Hindi. The task tests whether systems can select the correct answer to finance questions involving domain terminology, numerical interpretation, and conceptual financial reasoning across languages and scripts. The final-test set contains 800 questions, with 200 questions per language; gold answers were withheld during submission, and each language was ranked independently by accuracy. The final leaderboards contain 13 English, 11 Chinese, 11 Arabic, and 10 Hindi ranked submissions. Top accuracies range from 92.0% in Hindi to 97.5% in English and Arabic, with the same leading teams appearing near the top across all four languages. The documented systems used retrieval augmentation, direct answer-option scoring, language-specific prompting, selective self-consistency, confidence checks, and LLM-based review stages.

Comments9 pages. Task overview paper for CLEF 2026 Working Notes (CEUR Workshop Proceedings)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑