发表机构
University of Michigan; University of North Carolina at Charlotte(密歇根大学; 北卡罗来纳大学夏洛特分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出SCALE策略,结合选择性AI评分与序贯人工升级,在控制统计检验错误的同时最小化成本,并通过信息论下界验证其最优性。
AI 中文摘要
大型语言模型越来越多地被用作廉价的评判者来评估输出、标注数据,以及判断系统是否达到期望的质量标准。然而,将AI判断用于正式的统计推断与简单地将它们视为真实标签有着根本性的不同:AI评估可能存在偏差或噪声,而严格的假设检验需要显式控制第一类错误和第二类错误。我们研究如何利用AI判断,结合选择性人工验证,以最低成本进行有效的假设检验。我们考虑一个具有隐藏二值标签的项目总体。在选定一个固定的项目池后,决策者可以选择性地查询AI,将项目直接发送给人工,在观察到AI报告后将AI评分的项目升级给人工,或者在累积到足够证据后停止。我们推导了一个信息论下界,该下界刻画了在达到规定检验错误下实现最低成本,并通过一个依赖于报告的信息前沿来表征AI信息和人工验证的价值。受此表征的启发,我们开发了SCALE,一种序贯成本感知策略,它结合了选择性AI评分与自适应人工升级。SCALE在有限样本量下是有效的,并且当目标错误概率趋于零时,它与下界在一阶意义上匹配。我们进一步将框架扩展到使用成对AI-人工试点数据的未知AI输出模型。数值上,当单一来源明显占优时,SCALE接近纯人工或纯AI测试,而当廉价AI判断和选择性人工验证都有价值时,它实现最大的节省。
英文摘要
Large language models are increasingly used as inexpensive judges to evaluate outputs, label data, and assess whether a system meets a desired quality standard. Yet using AI judgments for formal statistical inference is fundamentally different from simply treating them as ground-truth labels: AI evaluations can be biased or noisy, and rigorous hypothesis testing requires explicit control of type-I and type-II errors. We study how to use AI judgments, together with selective human verification, to conduct a valid hypothesis test at minimum cost. We consider a population of items with hidden binary labels. After choosing a fixed pool of items, the decision maker can selectively query AI, send an item directly to a human, escalate an AI-scored item to a human after observing the AI report, or stop once sufficient evidence has accumulated. We derive an information-theoretic lower bound that captures the minimum cost of achieving prescribed testing errors and characterizes the value of AI information and human verification through a report-dependent information frontier. Motivated by this characterization, we develop SCALE, a sequential cost-aware policy that combines selective AI scoring with adaptive human escalation. SCALE is valid at finite sample sizes and matches the lower bound to first order as the target error probabilities vanish. We further extend the framework to an unknown AI-output model using paired AI-human pilot data. Numerically, SCALE approaches Human-only or AI-only testing when one source clearly dominates, while achieving its largest savings when inexpensive AI judgments and selective human verification are both valuable.