arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.01388cs.CL

RusFinChain:面向金融领域可验证思维链推理的俄语基准与模糊对齐评估

RusFinChain: A Russian Benchmark for Verifiable Chain-of-Thought Reasoning in Finance with Fuzzy-Aligned Evaluation

Mullosharaf K. Arabov

首次发表
浏览论文内容

中文总结 AI 辅助

提出首个俄语金融可验证思维链推理基准RusFinChain,包含5280个参数化示例和模糊对齐评估指标,揭示模型在步骤对齐与最终答案正确性之间存在显著差距。

中文摘要 AI 辅助

多步符号推理对于稳健的金融分析至关重要,但大多数基准忽略了中间推理步骤。FINCHAIN引入了可验证的思维链(CoT)评估,但仅限于英语。FINESSE-Bench包含一个俄语模块,但依赖于多项选择题,缺乏步骤级监督。我们提出了RusFinChain,这是首个用于金融领域可验证CoT推理的俄语符号基准。它涵盖17个领域、172个主题,包含来自可执行Python模板的5280个参数化示例,确保无污染评估。每个示例都包含一个带有中间数值的金标准推理链,用于自动验证。我们还引入了增强指标:模糊数值对齐和软注意力对齐。我们在分层样本上评估了8个开源大语言模型,生成了8100个响应。结果揭示了显著的推理差距:模型在步骤对齐上的Hard F1约为0.65,但最终答案的正确率仅为约29%。我们的模糊和软指标与最终答案正确性的相关性(Spearman rho约0.48)优于原始ChainEval(rho约0.38-0.46),显示出更强的诊断能力。我们发布了数据集、代码和评估框架,以促进面向俄语社区的可验证金融AI。

英文摘要

Multi-step symbolic reasoning is essential for robust financial analysis, yet most benchmarks neglect intermediate reasoning steps. FinChain introduced verifiable Chain-of-Thought (CoT) evaluation but is limited to English. FINESSE-Bench includes a Russian block but relies on multiple-choice questions without step-level supervision. We present RusFinChain, the first Russian-language symbolic benchmark for verifiable CoT reasoning in finance. It spans 17 domains, 172 topics, and comprises 5,280 parameterized examples from executable Python templates, ensuring contamination-free evaluation. Each example includes a gold-standard reasoning chain with intermediate numeric values for automatic verification. We also introduce enhanced metrics: Fuzzy Numeric Alignment and Soft-Attention Alignment. We evaluate 8 open-weight LLMs on a stratified sample, generating 8,100 responses. Results reveal a substantial reasoning gap: models achieve Hard F1 of ~0.65 for step alignment, but only ~29% of final answers are correct. Our fuzzy and soft metrics show stronger correlation with final-answer correctness (Spearman rho approx 0.48) than the original ChainEval (rho approx 0.38-0.46), demonstrating superior diagnostic power. We release dataset, code, and evaluation framework to foster verifiable financial AI for the Russian-speaking community.

发表机构

  • Kazan Federal University(喀山联邦大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑