分组随机机:将精度而非能力作为AI系统的前沿度量标准
Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Metric for AI Systems
浏览论文内容
中文总结 AI 辅助
本文提出将精度而非能力作为AI系统的前沿度量,指出当前基准测试未测量精度,定义了分组度量方法,其测量结果可指导决策,且该方法需通过实际工作测量验证规范价值。
中文摘要 AI 辅助
前沿语言模型的比较、推广与基准测试均围绕“能力”展开,即其最优或平均输出所能达成的效果。本文认为这一衡量维度是错误的:这类模型的准确率已趋于饱和,其平均输出已能命中目标。当前实践中区分不同系统的核心是精度,即模型在重复、相同请求下,输出围绕该目标的集中程度。借用射手的概念,能力是平均射击落点,而可靠性是弹着点的密集程度。本文提出三项主张:第一,精度而非能力是区分不同系统的前沿指标,而基准测试文化系统地未能测量精度,仅报告中心趋势而非离散程度;第二,精度可通过固定温度下多次运行一组经确定性评分的任务来廉价且无循环性地测量,无需模型内部 grader,即可计算任务间结果的一致性;第三,该测量不仅具有描述性,还能指导决策:它可区分两类失败——集中性失败(弹着点密集但偏离中心,可通过《论文1》的操作规范(即瞄准调整)修正)与分散性失败(弹着点分散,仅能通过更换模型或调整采样(即步枪问题)修正)。本文定义了分组度量标准,指定了测试框架,并展示了追踪人机对的分组随时间变化如何产生《论文1》实地研究所需的复合信号。首次实际运行(后被复制)既说明了该方法,也揭示了其最重要的局限:一个测得的差距可通过单条规则完全消除(0/5→5/5),而一套由规则本身生成的任务未发现任何价值,因为前沿模型已体现了明确的良好实践——这确立了规范的价值需通过实际工作中的测量来验证,而非从自身规则手册中构建。
英文摘要
Frontier language models are compared, marketed, and benchmarked on capability -- what their best or average output can achieve. I argue this measures the wrong axis. The models have saturated accuracy: their mean output lands on the target. What now separates one system from another in practice is precision: how tightly concentrated their outputs are around that target across repeated, identical requests. Borrowing the marksman's distinction, capability is where the average shot lands; reliability is the size of the group. I make three claims. First, precision, not capability, is the frontier differentiator between systems, and benchmark culture systematically fails to measure it, reporting central tendency rather than spread. Second, precision is measurable, cheaply and without circularity, by running a fixed suite of deterministically scored tasks many times at fixed temperature and computing the per-task consistency of outcomes -- no model-in-the-loop grader required. Third, the measurement is not merely descriptive but decision-guiding: it separates consistent failures (a tight group off-centre, correctable by the operating discipline of Paper 1 -- a sight adjustment) from scattered failures (a wide group, correctable only by changing the model or its sampling -- a rifle problem). I define a grouping metric, specify a harness, and show how tracking a human-AI pair's grouping over time yields the compounding signal that Paper 1's field study requires. A first real run, since replicated, illustrates both the method and its most important limit: one measured gap was closed completely by a single rule (0/5 -> 5/5), while a suite of tasks authored from the rules themselves found no value, because a frontier model already embodies explicit good practice -- establishing that a discipline's worth is found by measurement on real work, not constructed from its own rulebook.