发表机构
Stanford University; University of Notre Dame(斯坦福大学; 圣母大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出MERCI Cards,一个基于加权多目标优化的数学框架,用于高风险领域(如刑事司法)中LLM的迭代评估与部署,以指导系统改进并聚焦开发者与用户注意力于性能短板及输出验证。
AI 中文摘要
随着LLM越来越多地部署于高风险的专业工作流程中,工程师和研究人员需要有条理的协议来系统地跟踪、监控和改进模型在部署周期中的性能。我们提出了一个用于迭代式LLM评估与部署的数学框架,并展示了其在刑事司法领域AI系统中的应用。我们的框架通过加权多目标优化,形式化了LLM在高风险、高危险和资源受限领域中的集成,涵盖模型选择、规则设计、评估和部署。我们证明,MERCI Cards可以指导系统在多次部署迭代中的改进,将开发者的注意力引向表现不佳的领域,并引导用户关注对LLM输出的验证和错误纠正。
英文摘要
As LLMs are increasingly deployed in high-stakes professional workflows, engineers and researchers require principled protocols to systematically track, monitor, and improve model performance across deployment cycles. We present a mathematical framework for iterative LLM evaluation and deployment, and demonstrate its application to AI systems used in criminal justice. Our framework formalizes LLM integration in high-stakes, high-risk, and resource-constrained domains across model selection, rubric design, evaluations and deployment via a weighted multi-objective optimization. We demonstrate that MERCI Cards can guide improvements of the system across deployment iterations, direct developer attention toward under-performing areas, and focus user attention on validation and error-correction in the LLM's outputs.