arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08139cs.CLcs.AI

IGT @ FinMMEval 2026 任务2:面向多语言金融问答的题型提示与定向抽取

IGT @ FinMMEval 2026 Task 2: Question-Type Prompting with Targeted Extraction for Multilingual Financial QA

  • Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Yuwen Chiu

AI总结:

提出IGT系统,针对多语言金融问答,按题型区分结构化数值抽取与综合新闻选择,开发集ROUGE-1达0.395,测试集排名第三。

AI中文摘要:

我们提出了用于CLEF 2026 FinMMEval实验室PolyFiQA任务2的IGT系统,这是一个多语言金融问答任务,涉及四家公司的英文SEC文件和多语言新闻文章(英文、中文、日文、西班牙文、希腊文)。我们的核心观察是,344个开发问题可分为两类,需要根本不同的方法:结构化数值类型(研发占比、现金流、资本支出)最适合通过在申报文件文本上进行直接关键词抽取来回答,而综合类型(投资策略、资本配置、前三大收入重点)则需要基于规则的多语言新闻段落选择。数据集分析显示,每个综合类型的19个真实参考答案中有17-18个共享一个确切的证据标签前缀,其unigram标记直接贡献于ROUGE-1重叠。最终系统在开发集上达到ROUGE-1约0.395,相比通用RAG基线(约0.247)相对提升60%,并在官方测试集上排名12支队伍中的第3名,ROUGE-1=0.3071,精确率=0.2821,召回率=0.4044。

英文摘要:

We present the IGT system for PolyFiQA Task 2 of the FinMMEval Lab at CLEF 2026, a multilingual financial question answering task over English SEC filings and multilingual news articles (English, Chinese, Japanese, Spanish, Greek) for four companies. Our central observation is that the 344 development questions divide into two families requiring fundamentally different approaches: structured numeric types (R&D ratio, cash flow, capital expenditure) are best answered by direct keyword extraction on filing text, while synthesis types (investment strategy, capital allocation, top-three revenue focuses) require rule-based multilingual news passage selection. A dataset analysis reveals that 17-18 of 19 ground-truth reference answers per synthesis type share an exact evidence label prefix, whose unigram tokens contribute directly to ROUGE-1 overlap. The final system achieves development ROUGE-1 approximately 0.395, a 60% relative improvement over a generic RAG baseline (approximately 0.247), and ranks 3rd of 12 teams on the official test set with ROUGE-1 = 0.3071, Precision = 0.2821, and Recall = 0.4044.

补充信息

↑