arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于传感器的AI在分布偏移下的可问责与不确定性感知评估:设备、受试者与近三年地下实验

Accountable and uncertainty-aware evaluation of sensor-based AI under distribution shift: devices, subjects, and nearly three years underground

Benny Platte, Rico Thomanek, Christian Roschke, Marc Ritter

arXiv 2609.09257首次发表:更新:

发表机构

Mittweida University of Applied Sciences(米特韦达应用科学大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对传感器AI部署中设备、人员和时间变化导致的性能下降,提出分阶段可问责评估协议,以5%分位数而非均值决策,在地下矿井定位中验证,揭示均值误导性。

AI 中文摘要

基于传感器的AI系统很少在其训练条件下运行:设备、人员和记录时期会发生变化,而每种变化都会以随机训练-测试划分无法揭示的方式降低性能。我们提出了一种分阶段的、可问责的评估协议,该协议将部署模型的评估视为具有声明参考水平和量化不确定性的测量。四个累积泛化阶段分别留出设备、受试者和时间。每个阶段根据重复训练的分位数与具有正确类别数量的随机参考进行比较来判断,超出当前范围率暴露了相对于训练,部署中不再存在的类别的静默误导,而明确的决策规则将推出决策不基于均值而是基于5%分位数。我们在两个真实地下矿井中,使用基于智能手机的循环分类器进行无基础设施地磁定位,演示了该协议,包括在第二个地点复制该方案的训练阶段。未更改的模型在训练活动34个月后记录的数据上重新评估,该数据来自训练时未知的设备代次,并由留出的测量员采集。在那里,其当前条件精度的5%分位数在299次重复训练中为0.39,是随机水平的16.5倍;在42个可达位置类别的组成中,该数值变化±0.08,是重复运行之间差异的数倍。单一配置的重复训练显示了均值为何具有误导性:双峰配置通过基于均值的测试决定性地,而其5%分位数低于随机水平一个数量级以上。

英文摘要

Sensor-based AI systems are rarely operated under the conditions under which they were trained: devices, personnel and recording epochs change, and each change degrades performance in ways a random train-test split cannot reveal. We propose a staged, accountable evaluation protocol that treats the evaluation of a deployed model as a measurement with declared reference levels and a quantified uncertainty. Four cumulative generalisation stages hold out devices, subjects and time. Each stage is judged on quantiles of repeated trainings against chance references with the correct class count, an out-of-present-scope rate exposes silent misdirection towards classes that are no longer present in deployment relative to training, and an explicit decision rule ties roll-out decisions not to means but to 5% quantiles. We demonstrate the protocol on infrastructure-free geomagnetic localisation with smartphone-based recurrent classifiers in two real underground mines, including a replication of the scheme's training stages at the second site. Unchanged models are re-evaluated on data recorded 34 months after the training campaign, on a device generation unknown at training time and with a held-out surveyor. The 5% quantile of their present-conditioned precision there is 0.39 over 299 repeated trainings, 16.5 times the chance level; across the composition of the 42 reachable location classes the figure varies by +/-0.08, several times the spread between repeated runs. Repeated trainings of a single configuration show why means mislead: a bimodal configuration passes a mean-based test decisively while its 5% quantile lies more than an order of magnitude below chance.

Comments24 pages, 8 figures, 8 tables. Manuscript prepared for Measurement Science and Technology

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑