arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26114cs.AIcs.CLq-fin.CP

CIFQA:一种用于金融查询问答的确定性工具基多智能体大语言模型框架

CIFQA: A Deterministic Tool-Grounded Multi-Agent LLM Framework for Financial Query Answering

Kunjesh Parekh, Anil Kumar Tiwari, Divya Saxena

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对LLM多步骤金融计算易出错的问题,提出CIFQA多智能体框架,在定期存款查询基准上准确率达95.54%,且17B开源模型在该框架内表现优于更大前沿模型,证明架构设计对数值可靠性更关键。

中文摘要 AI 辅助

计算密集型金融问答需要对结构化利率、时间条件、数值公式和基于规则的约束进行精确推理。尽管大语言模型(LLMs)在自然语言任务中表现出色,但在解决多步骤金融计算时,它们常常产生数值错误却看似合理的答案。为解决这一局限,我们引入CIFQA(Calculation-Intensive Financial Query Answering,计算密集型金融查询问答),一种用于金融查询问答的确定性工具基多智能体大语言模型框架。CIFQA通过分配专门智能体分别负责查询解释、路由、参数提取、计算规划和响应生成,将语言理解与数值执行分离,同时基于Python的确定性工具执行金融计算和规则应用。我们将CIFQA实例应用于定期存款查询问答,并在精心整理的定期存款查询基准上对其进行评估。CIFQA在计算密集型查询上达到95.54%的准确率,整体准确率为90.87%,即使在提供完整公式、利率表和基准说明的情况下,也显著优于直接使用LLM的基线模型。消融研究表明,确定性组件(如精确利率查询、期限计算、滚动年限调整和提前支取逻辑)是性能的关键贡献因素。值得注意的是,在CIFQA框架内运行的17B参数开源主干模型,在使用相同金融信息评估时,表现显著优于更大规模的前沿模型,这表明架构设计比模型规模更能决定数值可靠性。尽管在定期存款查询上进行了评估,但CIFQA为计算密集型金融推理任务提供了可泛化的框架。

英文摘要

Calculation-intensive financial question answering requires exact reasoning over structured rates, temporal conditions, numerical formulas, and rule-based constraints. Although Large Language Models (LLMs) perform strongly on natural language tasks, they often produce numerically incorrect yet plausible answers when solving multi-step financial calculations. To address this limitation, we introduce CIFQA (Calculation-Intensive Financial Query Answering), a deterministic tool-grounded multi-agent LLM framework for financial question answering. CIFQA separates language understanding from numerical execution by assigning specialized agents to query interpretation, routing, parameter extraction, computation planning, and response generation, while deterministic Python-based tools perform financial calculations and rule application. We instantiate CIFQA for fixed deposit query answering and evaluate it on a curated benchmark of fixed deposit queries. CIFQA achieves 95.54% accuracy on calculation-intensive queries and 90.87% overall accuracy, substantially outperforming direct LLM baselines even when provided with complete formulas, rate cards, and benchmark instructions. Ablation studies show that deterministic components such as exact rate lookup, tenure computation, rolling-year adjustment, and premature-withdrawal logic are critical contributors to performance. Notably, a 17B open-source backbone operating within CIFQA outperforms substantially larger frontier models evaluated with the same financial information, demonstrating that architectural design is a more important determinant of numerical reliability than model scale. While evaluated on fixed deposit queries, CIFQA provides a generalizable framework for calculation-intensive financial reasoning tasks.

发表机构

  • School of Artificial Intelligence and Data Science, Indian Institute of Technology Jodhpur(印度理工学院焦特布尔分校人工智能与数据科学学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑