arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03960cs.AIcs.CL

FiMI Banking:印度零售银行的主权模型

FiMI Banking: A Sovereign Model for Indian Retail Banking

  • NPCI AI Research Team(NPCI AI研究团队)

机构由 AI 辅助整理,请以论文原文为准。

NPCI AI Research Team, Aman Kumar, Asit Desai, Chandra Bhushan, Harsh Sharma, Harshit Bhushan, Hrithik Kadam, Keyur Doshi, Kolisetty Sai Kapardheeswar, Krishanu… 展开作者

NPCI AI Research Team, Aman Kumar, Asit Desai, Chandra Bhushan, Harsh Sharma, Harshit Bhushan, Hrithik Kadam, Keyur Doshi, Kolisetty Sai Kapardheeswar, Krishanu Adhikary, Nadeem Shaik, Navya Prakash, Nitin Kukreja, Prashant Devadiga, Shamanth MH, Shantanu Pandey, Suvradip Paul, Yatharth Dedhia

AI总结:

该研究针对通用语言模型无法满足银行对话系统需求的问题,构建了受控的印度零售银行场景FiMI Banking,通过偏好优化与带可验证奖励的强化学习提升了银行智能体的安全行为及任务性能。

AI中文摘要:

银行需要能够回答产品问题、协助处理账户相关请求、并在严格运营与监管约束下安全运行的对话系统。通用语言模型无法可靠满足这些需求,在任务需要真实信息、正确工具使用或谨慎处理银行特定敏感场景时表现不足。我们推出FiMI Banking,这是一个受控的印度零售银行场景,基于经审核的银行文档、结构化真实数据、合成客户背景及银行工具构建。我们评估了两种后训练方法:针对响应级行为的偏好优化,以及针对多轮工具使用任务的带可验证奖励的强化学习。偏好优化大幅提升了安全行为:越界拒绝率从52%升至80%;强化学习将边缘案例性能从0.509提升至0.718、顺序敏感任务性能从0.590提升至0.679,同时减少了29%的生成token。这些结果表明,偏好优化与带可验证奖励的强化学习可满足可靠银行智能体的互补需求。

英文摘要:

Banks need conversational systems that can answer product questions, assist customers with account-related requests, and operate safely within strict operational and regulatory constraints. General-purpose language models do not reliably meet these requirements. They fall short when a task requires grounded information, correct tool use, or cautious handling of bank-specific sensitive situations. We introduce FiMI Banking, a controlled Indian retail-banking setting. We build it from vetted banking documents, structured ground truth, synthetic customer backgrounds, and banking tools. We evaluate two post-training approaches: preference optimization for response-level behavior, and reinforcement learning with verifiable rewards for multi-turn tool-use tasks. Preference optimization improves safe behavior substantially: out-of-scope refusal rises from 52% to 80%. Reinforcement learning improves edge-case performance from 0.509 to 0.718 and order-sensitive task performance from 0.590 to 0.679, while using 29% fewer generated tokens. These results show that preference optimization and verifiable-reward reinforcement learning address complementary requirements for reliable banking agents.

↑