LLM输出漂移:金融工作流的跨提供商验证与缓解
LLM Output Drift: Cross-Provider Validation & Mitigation for Financial Workflows
- IBM -- Financial Services Market(IBM金融服务市场)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对金融工作流中LLM输出漂移问题,量化五种模型在金融任务上的漂移,提出金融校准测试工具、任务特定不变性检查等方法,为合规部署提供支持。
AI中文摘要:
金融机构将大型语言模型(LLMs)用于对账、监管报告和客户沟通,但非确定性输出(输出漂移)会削弱可审计性和信任度。我们在受监管的金融任务上量化了五种模型架构(7B-120B参数)的漂移情况,发现了明显的反比关系:较小的模型(Granite-3-8B、Qwen2.5-7B)在T=0.0时达到100%的输出一致性,而GPT-OSS-120B无论配置如何仅表现出12.5%的一致性(95%置信区间:3.5-36.0%,Fisher精确检验p<0.0001)。这一发现挑战了大型模型在生产部署中普遍更优的传统假设。我们的贡献包括:(i)一个金融校准的确定性测试工具,结合贪婪解码(T=0.0)、固定种子和SEC 10-K结构感知检索排序;(ii)使用金融校准的重要性阈值(±5%)和SEC引用验证对RAG、JSON和SQL输出进行任务特定的不变性检查;(iii)一个三层模型分类系统,支持风险适配的部署决策;(iv)一个具备双提供商验证的可审计认证系统。我们在三项受监管金融任务上评估了五种模型(通过Ollama的Qwen2.5-7B、通过IBM watsonx.ai的Granite-3-8B、Llama-3.3-70B、Mistral-Medium-2505和GPT-OSS-120B)。在480次运行中(每个条件n=16),结构化任务(SQL)即使在T=0.2时仍保持稳定,而RAG任务则出现漂移(25-75%),显示出任务依赖的敏感性。跨提供商验证证实确定性行为在本地和云部署之间可转移。我们将框架映射到金融稳定委员会(FSB)、国际清算银行(BIS)和商品期货交易委员会(CFTC)的要求,展示了合规AI部署的实用路径。
英文摘要:
Financial institutions deploy Large Language Models (LLMs) for reconciliations, regulatory reporting, and client communications, but nondeterministic outputs (output drift) undermine auditability and trust. We quantify drift across five model architectures (7B-120B parameters) on regulated financial tasks, revealing a stark inverse relationship: smaller models (Granite-3-8B, Qwen2.5-7B) achieve 100% output consistency at T=0.0, while GPT-OSS-120B exhibits only 12.5% consistency (95% CI: 3.5-36.0%) regardless of configuration (p<0.0001, Fisher's exact test). This finding challenges conventional assumptions that larger models are universally superior for production deployment. Our contributions include: (i) a finance-calibrated deterministic test harness combining greedy decoding (T=0.0), fixed seeds, and SEC 10-K structure-aware retrieval ordering; (ii) task-specific invariant checking for RAG, JSON, and SQL outputs using finance-calibrated materiality thresholds (plus or minus 5%) and SEC citation validation; (iii) a three-tier model classification system enabling risk-appropriate deployment decisions; and (iv) an audit-ready attestation system with dual-provider validation. We evaluated five models (Qwen2.5-7B via Ollama, Granite-3-8B via IBM watsonx.ai, Llama-3.3-70B, Mistral-Medium-2505, and GPT-OSS-120B) across three regulated financial tasks. Across 480 runs (n=16 per condition), structured tasks (SQL) remain stable even at T=0.2, while RAG tasks show drift (25-75%), revealing task-dependent sensitivity. Cross-provider validation confirms deterministic behavior transfers between local and cloud deployments. We map our framework to Financial Stability Board (FSB), Bank for International Settlements (BIS), and Commodity Futures Trading Commission (CFTC) requirements, demonstrating practical pathways for compliance-ready AI deployments.