IndicBankBench:评估印度零售银行业语言模型助手的安全性与可靠性
IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking
- National Payments Corporation of India(印度国家支付公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对银行助手仅评估最终响应会遗漏错误的问题,提出IndicBankBench基准,含799个案例、四阶段评估及严格pass^3指标,揭示至少一次成功率高估可靠性,助力提升印度零售银行助手安全与可靠性。
AI中文摘要:
银行助手必须使用账户特定信息来回答请求,并在许多情况下通过工具执行操作。仅评估最终响应会遗漏重要错误。助手可能会询问其已掌握的信息、依赖过时的上下文、选择错误的账户,或在陈述正确数值后写入无效值。我们推出IndicBankBench,一个包含799个案例的基准,覆盖印度零售银行业的五个运营领域、一个能力/拒绝领域以及二十个主要评估轴。案例在四个阶段进行评估:安全性、操作与工具使用、响应充分性以及咨询质量。工具使用和大多数安全检查是确定性的。一个狭窄的解析器仅处理模糊的写前确认案例,而一个独立的LLM评判器评估语义响应充分性。我们每个案例运行三次,并报告严格的pass^3,即要求所有试验均成功。在评估的十一个模型中,严格可靠性范围为43.7%至58.2%,而至少一次成功率为60%至74%。这一差距表明,至少一次成功率可能高估了可靠的银行行为。案例级诊断还能区分那些提出不必要问题的系统与那些采取行动但未能协调客户上下文或完全解决请求的系统。我们发布了案例、模拟环境和评估框架。
英文摘要:
Banking assistants must use account-specific information to answer requests and, in many cases, take actions through tools. Evaluating only the final response misses important errors. An assistant may ask for information it already has, rely on stale context, select the wrong account, or write an invalid value after stating the correct one. We introduce IndicBankBench, a 799-case benchmark for Indian retail banking spanning five operational domains, a capability/refusal domain, and twenty primary axes. Cases are evaluated at four stages: safety, action and tool use, response adequacy, and advisory quality. Tool use and most safety checks are deterministic. A narrow resolver handles only ambiguous confirmation-before-write cases, while a separate LLM judge evaluates semantic response adequacy. We run every case three times and report strict pass^3, which requires success on all trials. Across the eleven evaluated models, strict reliability ranges from 43.7% to 58.2%, whereas at-least-once success ranges from 60% to 74%. This gap shows that at-least-once success can overstate dependable banking behavior. The case-level diagnostics also distinguish systems that ask unnecessary questions from those that act but fail to reconcile customer context or fully resolve the request. We release the cases, mock environment, and evaluation harness.