IBIB:一种按服务路由而非模型标识符度量企业AI系统的协议
IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier
浏览论文内容
中文总结 AI 辅助
针对企业AI系统按标识符评分导致的测量误差,提出IBIB协议,通过预检、可靠性评分和盲裁决度量路由能力,实验表明其可区分系统并改变结论。
中文摘要 AI 辅助
企业部署的是系统,而非检查点。可用能力共同取决于权重、服务路由、精度、输出契约和测试框架,然而所有18个经审计的基准都只对广告宣传的模型标识符进行评分。我们将此视为测量误差,并给出一种使其可报告的协议。该协议包含三个部分。一个金标盲的能力绑定预检在任何任务到达之前验证路由能否执行评估契约;一个包含可靠性的首轮评分规则将失败计入得分,同时将不支持的能力排除在外;裁决在结构上对评分保持盲态。我们将该协议称为IB2,并发布其算法、分类表、请求契约和清单模式。其参考实现——128个锁定任务和987个断言,涵盖文档、电子表格、图表、工具和数据库工作——保持密封:程序本身即工件,而非语料库。在十一个系统上,有四个结果。能力可用性是可测量的:两次在相同权重上的完整单路由运行后来未能通过最终绑定门的不同谓词,而第三次在重新运行前通过了该门。广告宣传的标识符未暴露任一限制。区分度并非均匀:七个测试套件中有四个在六系统带宽下饱和,差异几乎完全来自受治理的数据库工作和多标签页连接,因此我们报告基于区间的分辨率组,而非排名;名义五标签输出的四个切分中有两个未能通过多重性调整。服务臂的选择使一个声明的版本和精度从77.38变为82.54,配对区间为[0.11,10.60],尽管这些臂在访问模式、测试框架生成和服务工具调用解析器上有所不同,且测试框架生成是我们评估器的属性,而非任何端点。将失败响应从分母中排除会改变点排序,因此包含可靠性会改变结论,而不仅仅是措辞。
英文摘要
Enterprises deploy systems, not checkpoints. Usable capability depends jointly on weights, serving route, precision, output contract, and harness, yet all 18 audited benchmarks score advertised model identifiers. We treat this as measurement error and give a protocol that makes it reportable. It has three parts. A gold-blind capability-binding preflight verifies that a route can execute the evaluation contract before any task reaches it; a reliability-inclusive first-pass scoring rule keeps failure in the score while keeping unsupported capability out; and adjudication is structurally score-blind. We call the protocol IB2 and release its algorithms, classification tables, request contract, and manifest schemas. Its reference instantiation, 128 locked tasks and 987 assertions over document, spreadsheet, chart, tool and database work, stays sealed: the procedure is the artifact, not the corpus. Across eleven systems, four results. Capability availability is measurable: two complete single-route runs on identical weights later failed distinct predicates of the finalized binding gate, while a third passed that gate before a fresh run. The advertised identifier exposed neither limit. Discrimination is not uniform: four of seven suites saturate under a six-system band, with the spread almost entirely from governed database work and multi-tab joins, so we report interval-backed resolution groups, not ranks; two of the nominal five-label output's four cuts fail multiplicity adjustment. Serving-arm choice moved one declared revision and precision from 77.38 to 82.54, paired interval [0.11,10.60], though the arms differ in access mode, harness generation, and the serving tool-call parser, and harness generation is a property of our evaluator, not any endpoint. Excluding failed responses from denominators changes the point ordering, so reliability inclusion changes a conclusion, not its wording.