arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MERCI Cards:面向高风险领域的LLM评估与部署框架

MERCI Cards: An LLM Evaluation and Deployment Framework for High-Stakes Domains

Aparna Komarla, Annalisa Szymanski

arXiv 2610.04430首次发表:更新:

发表机构

Stanford University; University of Notre Dame(斯坦福大学; 圣母大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出MERCI Cards,一个基于加权多目标优化的数学框架,用于高风险领域(如刑事司法)中LLM的迭代评估与部署,以指导系统改进并聚焦开发者与用户注意力于性能短板及输出验证。

AI 中文摘要

随着LLM越来越多地部署于高风险的专业工作流程中,工程师和研究人员需要有条理的协议来系统地跟踪、监控和改进模型在部署周期中的性能。我们提出了一个用于迭代式LLM评估与部署的数学框架,并展示了其在刑事司法领域AI系统中的应用。我们的框架通过加权多目标优化,形式化了LLM在高风险、高危险和资源受限领域中的集成,涵盖模型选择、规则设计、评估和部署。我们证明,MERCI Cards可以指导系统在多次部署迭代中的改进,将开发者的注意力引向表现不佳的领域,并引导用户关注对LLM输出的验证和错误纠正。

英文摘要

As LLMs are increasingly deployed in high-stakes professional workflows, engineers and researchers require principled protocols to systematically track, monitor, and improve model performance across deployment cycles. We present a mathematical framework for iterative LLM evaluation and deployment, and demonstrate its application to AI systems used in criminal justice. Our framework formalizes LLM integration in high-stakes, high-risk, and resource-constrained domains across model selection, rubric design, evaluations and deployment via a weighted multi-objective optimization. We demonstrate that MERCI Cards can guide improvements of the system across deployment iterations, direct developer attention toward under-performing areas, and focus user attention on validation and error-correction in the LLM's outputs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑