发表机构
Stony Brook University(石溪大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MECHVAR提出一种基于后验加权方差的轻量级获取规则,用于有限库自主机器学习实验选择,在保持可审计性的同时实现高效评分,并在多个中等错误设定下优于先确认策略。
AI 中文摘要
基准测试的提升往往具有机制模糊性:复现一个改进本身并不能识别其产生的原因。我们研究了有限库机制判别问题,其中后验加权候选机制、可执行探针和有限的实验预算定义了一个序贯实验选择问题。MECHVAR通过最大化其预测响应的后验加权方差来选择下一个探针。在共享高斯预测模型下,该分数恰好与经典的Box-Hill后验加权成对KL准则成正比,但它允许O(KE)向量化重新评分,并能对机制对进行透明的加性审计。局部展开进一步将该分数与预测响应分离较小时的期望信息增益(EIG)联系起来。在25块压力审计中,MECHVAR在几个中等错误设定情形下优于先确认策略,而其与EIG的主要比较在统计上仍未解决。在留出法Digits循环中,MECHVAR的归一化机制识别AUC为0.8975,分数贪婪策略为0.7825,EIG为0.9092。在K=100、E=200时,记录环境中MECHVAR的单线程全库评分中位数为10.36微秒,而六节点求积EIG为57.69毫秒。因此,当共享预测尺度是合理近似时,MECHVAR为有限库实验选择提供了一种轻量级、可审计的获取规则。
英文摘要
Benchmark gains are often mechanism-ambiguous: reproducing an improvement does not by itself identify why it occurs. We study finite-library mechanism discrimination, where posterior-weighted candidate mechanisms, executable probes, and a limited experimental budget define a sequential experiment-selection problem. MECHVAR selects the next probe by maximizing the posterior-weighted variance of its predicted responses. Under a shared-Gaussian predictive model, this score is exactly proportional to the classical Box--Hill posterior-weighted pairwise-KL criterion, yet it admits O(KE) vectorized rescoring and a transparent additive audit over mechanism pairs. A local expansion further links the score to expected information gain (EIG) when predicted response separations are small. In a 25-block stress audit, MECHVAR outperforms confirmation-first in several moderate misspecification regimes, while its primary comparisons with EIG remain statistically unresolved. In a held-out Digits loop, normalized mechanism-identification AUC is 0.8975 for MECHVAR, 0.7825 for a score-greedy policy, and 0.9092 for EIG. At K = 100, E = 200, median single-thread full-library scoring is 10.36 microseconds for MECHVAR versus 57.69 ms for six-node quadrature EIG in the recorded environment. MECHVAR therefore provides a lightweight, auditable acquisition rule for finite-library experiment selection when a shared predictive scale is a defensible approximation.
Comments17 pages, 7 figures