发表机构
Prometeia S.p.A.(普罗米特亚公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究指出金融大语言模型不能仅靠基准测试验证,提出需从系统全栈进行多层验证,还讨论了LLM-as-a-judge的应用与控制措施,强调要研究系统感知基准等相关方向。
AI 中文摘要
大型语言模型正越来越多地被部署到结合了检索、专有数据、工具使用、编排逻辑、监控和人工上报流程的金融应用中。然而,评估往往仍以模型为中心:基准测试分数、任务准确率或一次性定性评估被视为模型可投入使用的证据。在金融场景中,这是不够的。我们认为,不能仅根据基准测试性能就批准金融大语言模型系统投入生产,它们需要覆盖应用全栈的系统级验证证据:数据、模型设计、检索与生成性能、智能体行为、治理及实现。基于在金融机构验证生成式AI应用的行业经验,我们概述了多层验证视角,并解释了为何混合评估是必要的。我们讨论了LLM-as-a-judge方法的适用场景,以及为何它们需要多重裁判、评分规则、一致性和可审计性检查等控制措施。我们还强调了静态基准测试难以捕捉的失败模式,包括检索失败、生成内容不忠实、工具误用、上报错误和操作不稳定。我们的立场是,金融大语言模型验证应是持续的系统级工作,而非一次性的模型评分练习,验证应产出可用于决策的证据,而非仅分数。最后,我们提出了系统感知基准测试、智能体轨迹验证、裁判对齐协议和生命周期验证标准的研究议程。
英文摘要
Large language models are increasingly deployed in financial applications that combine retrieval, proprietary data, tool use, orchestration logic, monitoring, and human escalation. Yet evaluation often remains model-centric: benchmark scores, task accuracy, or one-off qualitative reviews are treated as evidence of readiness. In financial settings, this is insufficient. We take the position that financial LLM systems should not be approved for production based on benchmark performance alone. They require system-level validation evidence across the application stack: data, model design, retrieval and generation performance, agent behavior, governance, and implementation. Drawing on industry experience validating GenAI applications in financial institutions, we outline a multi-layer validation view and explain why hybrid evaluation is necessary. We discuss where LLM-as-a-judge methods are useful and why they require controls such as multiple judges, rubrics, agreement, and auditability checks. We also highlight failure modes poorly captured by static benchmarks, including retrieval failures, unfaithful generation, tool misuse, escalation errors, and operational instability. Our position is that financial LLM validation should be an ongoing system discipline rather than a one-time model scoring exercise. Validation should produce decision-ready evidence, not only scores. We conclude with a research agenda for system-aware benchmarks, agent trace validation, judge alignment protocols, and lifecycle validation standards.