GrandJury: A Collaborative Machine Learning Model Evaluation Protocol for Dynamic Quality Rubrics
Arthur Cho
机构
*
Memoirji LLC(Memiorji公司)
专题命中
推理评测
:reasoning(abstract);分类 cs.AI、cs.LG
Comments14 pages (incl. arXiv cover), 1 table, code & dataset links inside. Open-source implementation available on PyPI (grandjury package) and GitHub. Dataset available on Hugging Face under CC-BY-4.0 license
Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
Andy Zou, Maxwell Lin, Eliot Jones, Micha Nowak, Mateusz Dziemian, Nick Winter, Alexander Grattan, Valent Nathanael, Ayla Croft, Xander Davies, Jai Patel, Robert Kirk, Nate Burnikell, Yarin Gal, Dan Hendrycks, J. Zico Kolter, Matt Fredrikson
机构
*
Department of Computer Science and Engineering, The Chinese University of Hong Kong(中国香港中文大学计算机科学与工程系)
;
School of Electronic Science and Engineering, Nanjing University(南京大学电子科学与工程学院)
;
School of Integrated Circuits, Peking University(北京大学集成电路学院)
;
School of Intergrated Circuits, Southeast University(东南大学集成电路学院)
;
School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
;
Department of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术系)
;
National Center of Technology Innovation for EDA(EDA技术创新国家中心)
专题命中
推理评测
:reasoning(abstract);分类 cs.AI、cs.LG
Comments10 pages, 1 figure, 5 tables. To appear in ICCAD 2025
2048: Reinforcement Learning in a Delayed Reward Environment
Prady Saligram, Tanvir Bhathal, Robby Manihani
专题命中
推理评测
:planning(abstract);分类 cs.AI、cs.LG
CommentsWe found an issue with our result aggregation scripts: some evaluation logs were incomplete and others duplicated, causing incorrect numbers in tables and figures. Because these graphs and tables underpin key comparisons, we are withdrawing the paper to regenerate verified results
机构
*
HeartVoice Medical Technology(HeartVoice医疗科技)
;
Saw Swee Hock School of Public Health and Institute of Data Science(Saw Swee Hock公共卫生学院和数据科学研究所)
;
National University of Singapore(新加坡国立大学)
;
Department of Cardiology, Peking University People’s Hospital(北京大学人民医院心内科)
;
College of Integrative Chinese and Western Medicine, Anhui University of Chinese Medicine(安徽中医药大学整合中西医学学院)
;
National Institute of Health Data Science, Peking University(北京大学国家健康数据科学研究院)
;
Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院)
机构
*
Indian Institute of Technology Kharagpur(印度理工学院Khargapur分校)
;
Rochester Institute of Technology(罗切斯特理工学院)
;
Researcher, Fatima Fellowship(研究员,Fatima fellowship)
;
AI Institute, University of South Carolina(人工智能研究所,南卡罗来纳大学)
;
Amazon GenAI(亚马逊生成人工智能)
;
Stanford University(斯坦福大学)
;
James Silberrad Brown Center for AI, San Diego State University(詹姆斯·西伯拉德·布朗人工智能中心,圣地亚哥州立大学)
专题命中
推理评测
:reasoning(abstract);分类 cs.CL、cs.AI
Comments18 pages, 9 figures, KDD workshop on Prompt Optimization 2025
Retrieval-Augmented Clinical Benchmarking for Contextual Model Testing in Kenyan Primary Care: A Methodology Paper
Fred Mutisya, Shikoh Gitau, Christine Syovata, Diana Oigara, Ibrahim Matende, Muna Aden, Munira Ali, Ryan Nyotu, Diana Marion, Job Nyangena, Nasubo Ongoma, Keith Mbae, Elizabeth Wamicha, Eric Mibuari, Jean Philbert Nsengemana, Talkmore Chidede
专题命中
推理评测
:reasoning(abstract);分类 cs.CL、cs.AI
Comments29 pages, 6 figs, 6 tables. Companion methods paper forthcoming