arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.31108cs.LG

对高效负责任AI评估的压力测试:当计算节省改变基准结论时

Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark Conclusions

  • Vector Institute(向量研究所)
  • University of Calgary(卡尔加里大学)

机构由 AI 辅助整理,请以论文原文为准。

Ahmed El Kady, Aravind Narayanan, Rehana Riaz, Yani Ioannou, Shaina Raza

AI总结:

该研究通过在BBQ和BBQ-V上对3种模型开展7种条件下的测试,验证高效负责任AI评估的结论稳健性,发现不同高效评估方式对性能、能耗等的影响,提出需检查高效评估的有效性。

AI中文摘要:

高效评估会改变用于支撑模型行为相关结论的协议,但很少有人测试这些结论在评估本身变得更廉价后是否依然稳定。我们通过在涵盖批处理、量化、基准缩减及其组合的7种条件下,于BBQ和BBQ-V上评估3个密集模型和混合专家模型,对负责任AI基准测试中的结论稳健性进行压力测试。我们不将保留的总体准确率视为足够,而是将准确率、偏差严重程度与流行度、推理质量、子群行为、子集成员稳定性、运行时间及测得的GPU能耗与全基准BF16基线进行对比。较大的批处理使准确率保持在基线的0.35个百分点以内,且产生相对较小的子群变化,在6种模型-数据集设置中的5种里降低了能耗。INT8在很大程度上保留了质量,但能耗是基线的1.79至4.26倍;INT4则导致更大的、依赖模型和上下文的变化。缩减后的基准提供了最一致的节省,但极小的子集对保留哪些项目明显更敏感。因此,高效评估应被视为一种测量干预,其有效性必须针对基准旨在支撑的所有结论进行检查。本项目网址为this https URL,代码可在this https URL获取。

英文摘要:

Efficient evaluation changes the protocol used to support claims about model behavior, yet it is rarely tested whether those claims remain stable after the evaluation itself is made cheaper. We stress-test conclusion robustness in responsible-AI benchmarking by evaluating three dense and mixture-of-experts models on BBQ and BBQ-V under seven conditions spanning batching, quantization, benchmark reduction, and their combinations. Rather than treating preserved aggregate accuracy as sufficient, we compare accuracy, bias severity and prevalence, reasoning quality, subgroup behavior, subset-membership stability, runtime, and measured GPU energy against a full-benchmark BF16 baseline. Larger batching keeps accuracy within 0.35 percentage points of baseline and produces comparatively small subgroup changes, while reducing energy in five of six model--dataset settings. INT8 largely preserves quality but uses 1.79--4.26$\times$ baseline energy. INT4 causes larger, model- and context-dependent changes. Reduced benchmarks provide the most consistent savings, but very small subsets are substantially more sensitive to which items are retained. Efficient evaluation should therefore be treated as a measurement intervention whose validity must be checked across the conclusions the benchmark is intended to support. Our project website is https://vectorinstitute.github.io/sustainable-rai-evaluation/ and the code is available at https://github.com/VectorInstitute/sustainable-rai-evaluation.

↑