发表机构
leading UK bank(英国领先银行)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出基于合成客户智能体(SCA)数字孪生的两部分验证方法,用于解决银行等领域LLM聊天机器人的大规模安全验证问题,已成功应用于英国某头部银行的客户聊天机器人验证。
AI 中文摘要
基于大语言模型(LLM)的聊天机器人正在改变银行等受监管领域的客户服务,但可扩展且具成本效益的验证仍是安全部署的关键障碍。本文提出两部分贡献用于大规模聊天机器人验证:其一,引入创建高保真合成客户智能体(SCA)作为数字孪生的方法,该方法基于真实交易与对话数据,支持自动生成及行为调节,以模拟多样化客户画像与互动风格。评估显示,SCA与真实客户语义对齐度高、幻觉率低,且可通过可控干预成功复现人格特质;其二,开发基于SCA的验证框架,结合自动LLM作为评判者的评估、人类专家测试及对抗性探测。针对情绪状态、人口统计群体及语言因素的场景化验证证实其性能稳健。本文方法已用于验证英国某领先银行面向客户的聊天机器人,为金融机构提供了通往合规性的可扩展路径。
英文摘要
LLM-based chatbots are transforming customer service in regulated domains such as banking, but scalable and cost-effective validation remains a critical barrier to safe deployment. We present a two-part contribution for large-scale chatbot validation. First, we introduce a methodology for creating high-fidelity synthetic customer agents (SCAs) as digital twins, grounded in real transactional and conversational data, that enables automatic generation and behavioral conditioning to simulate diverse customer profiles and interaction styles. Evaluation demonstrates that SCAs achieve high semantic alignment with real customers, low hallucination rates, and successful personality trait reproduction with controllable interventions. Second, we develop an SCA-based validation framework combining automated LLM-as-a-Judge evaluation, human expert testing, and adversarial probing. Scenario-based validation across emotional states, demographic groups, and linguistic factors confirms robust performance. Our approach was used to validate a customer facing chatbot at a leading UK bank, providing financial institutions with a scalable pathway toward regulatory compliance.