arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

开源权重模型能在金融文本理解领域与(闭源)模型竞争吗?

Can Open-Weight Models Compete on Financial Text Comprehension?

Jan Spörer

arXiv 2608.08634首次发表:更新:

发表机构

University of St. Gallen(圣加仑大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究更新了Financial Touchstone金融文本理解基准,测试了20款模型,发现开源权重模型Kimi K2.6准确率排第三,挑战了推理架构或闭源权重是金融理解前提的假设,还发现中国模型存在地缘政治内容过滤器拒绝合法金融问题的现象。

AI 中文摘要

近几个月来,来自中国AI实验室的开源权重语言模型在基准测试上已追赶上闭源前沿模型,但它们在现实世界金融任务中的可靠性仍未得到充分测试。我们更新了Financial Touchstone基准,该基准现在包含495份国际年度报告中的2967个问题-上下文-答案三元组。我们还将一组新模型应用于该基准,覆盖范围从10家提供商的11个模型扩展到20个模型,其中包括GLM 4.7、GLM 5、Kimi K2.6、DeepSeek V3.2等近期开源权重模型,以及阿里巴巴的闭源旗舰模型Qwen3-Max。Anthropic的Claude Opus 4.6达到最高准确率(88.4%),而Google的Gemini 2.5 Pro保持最低幻觉率(0.08%)。值得注意的是,开源权重模型Kimi K2.6在准确率排名第三,非推理模型GLM 5和Mistral 3分别排名第四和第五,这挑战了推理架构或闭源权重是出色金融理解的前提的假设。信息检索仍是主要瓶颈,占所有失败案例的48.9%。我们还记录了一项新发现:中国模型中的地缘政治内容过滤器会拒绝合法的金融问题(占尝试次数的0.08%),有时没有明确原因,且拒绝行为既取决于访问途径也取决于模型。完整数据集和评估框架已公开提供。

英文摘要

Open-weight language models from Chinese AI labs caught up on benchmarks relative to proprietary frontier models in recent months. Yet their reliability on real-world financial tasks remains largely untested. We updated the Financial Touchstone benchmark, which now has 2,967 question context-answer triplets across 495 international annual reports. We also apply a new set of models on the benchmark, expanding coverage from eleven to twenty models across ten providers, including recent open-weight models such as GLM 4.7, GLM 5, Kimi K2.6, and DeepSeek V3.2, as well as Alibaba's proprietary flagship Qwen3-Max. Anthropic's Claude Opus 4.6 achieves the highest accuracy (88.4%), while Google's Gemini 2.5 Pro maintains the lowest hallucination rate (0.08%). Notably, the open-weight Kimi K2.6 ranks third in accuracy, and the non-reasoning models GLM 5 and Mistral 3 rank fourth and fifth, challenging the assumption that reasoning architectures or proprietary weights are a prerequisite for strong financial comprehension. Information retrieval remains the primary bottleneck, accounting for 48.9% of all failures. We also document a new finding: geopolitical content filters in Chinese models refuse legitimate financial questions (0.08% of attempts), sometimes without clear reason, and the refusal behavior depends on the access route as much as on the model. The complete dataset and evaluation framework are publicly available.

CommentsTo be presented at the workshop International Symposium on Large Language Models for Financial Services (FinLLM@IJCAI2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑