Benchmarking is Broken -- Don't Let AI be its Own Judge
机构 * Princeton University(普林斯顿大学) ; CISPA Helmholtz Center for Information Security(CISPA海德堡信息安全中心) ; Michigan State University(密歇根州立大学) ; Ohio State University(俄亥俄州立大学) ; Brown University(布朗大学) ; Massachusetts Institute of Technology(麻省理工学院) ; University of California, Los Angeles(加州大学洛杉矶分校) ; University of Tübingen(图宾根大学) ; Old Dominion University(旧 Dominion 大学) ; Technical University of Munich(慕尼黑技术大学) ; Cornell University(康奈尔大学) ; Forest AI(森林AI)
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
Comments 14 pages; Accepted to NeurIPS 2025. Link to poster: https://neurips.cc/virtual/2025/poster/121919; Link to project website: https://www.peerbench.ai/