An Experimental Design Approach to Evaluating Agentic AI's Autonomous Model Discovery
一种评估智能体人工智能自主模型发现的实验设计方法
Hao He, Xueying Liu, Chris J. Kuhlman, Xinwei Deng
机构
*
Department of Statistics, Virginia Tech(统计学系,弗吉尼亚理工学院)
;
Department of Statistical Science, Baylor University(统计科学系,贝勒大学)
;
Advanced Research Computing, Virginia Tech(高级研究计算,弗吉尼亚理工学院)
机构
*
Hong Kong Polytechnic University(香港理工大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Shanghai AI Lab(上海人工智能实验室)
;
National University of Singapore(新加坡国立大学)
Comments15 pages. Extended preprint incorporating versions accepted at Agent Skills '26 (ACM CAIS 2026) and the KDD 2026 Workshop on Enterprise AI Agents: From Prototypes to Production (oral presentation). Open-source implementation: https://github.com/NVIDIA/SkillEvaluator
Machine Learning and ARIMA Model Averaging for Adaptive Public Health Forecasting: Comparative Evaluation and an Ontario COVID-19 Case Study
机器学习与ARIMA模型平均法用于自适应公共卫生预测:比较评估及安大略省COVID-19案例研究
Yushu Zou, Ye Li, Johra Moosa, Martin Grunnill, Samir N. Patel, Venkata R. Duvvuri
机构
*
Public Health Ontario(安大略省公共卫生局)
;
Dalla Lana School of Public Health, University of Toronto(多伦多大学达拉·拉纳公共卫生学院)
;
University of Toronto(多伦多大学)
;
York University(约克大学)
PLSQLBench: Benchmarking LLM Systems for Executable Procedural Database Programming
PLSQLBench:面向可执行过程式数据库编程的大语言模型系统基准测试
Marianne Menglin Liu, Leonid Boytsov, Daniel W. Peterson, Pramuditha Perera, Rongguang Wang, Sai Ashish Somayajula, Syed Hamza Rafique, Rohit Saini, Shubham Pathak, Sujeeth Bharadwaj, Tao Sheng, Graham Horwood, Fahad Shah, Ankan Bansal, Sujith Ravi, Dan Roth
机构
*
Oracle AI, Turing Enterprise Inc(Oracle AI 图灵企业公司)
Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams
Diagram-MMU:面向科学图表的多模态基准测试
Weihao Bo, Shan Zhang, Yanpeng Sun, Jie Liu, Yongke Yao, Jinhao Du, Wei He, Kai Zou, Zechao Li, Jingdong Wang
机构
*
Nanjing University of Science and Technology(南京理工大学)
;
Baidu Inc(百度公司)
;
AIML, Adelaide University(阿德莱德大学AIML)
;
SUTD(新加坡科技设计大学)
;
Southeast University(东南大学)
;
East China Normal University(华东师范大学)
;
University of Oxford(牛津大学)
The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes
LLM推理的周期表:推理范式、方法与失败模式的结构化综述
Avinash Anand, Mahisha Ramesh, Avni Mittal, Ashutosh Kumar, Rishitej Reddy Vyalla, Erik Cambria, Zhengkui Wang, Timothy Liu, Aik Beng Ng, Simon See, Rajiv Ratn Shah
机构
*
Singapore Institute of Technology(新加坡理工大学)
;
Nvidia AI Center (SNAIC)(英伟达人工智能中心(SNAIC))
;
MIDAS Lab, IIIT Delhi(IIIT德里MIDAS实验室)
;
MIDAS Lab, IIT Mandi(IIT曼迪MIDAS实验室)
;
Owl Autonomous Imaging, Inc.(Owl自主成像公司)
;
College of Computing & Data Science, NTU Singapore(新加坡南洋理工大学计算与数据科学学院)
;
NVIDIA AI Technology Centre, Singapore(英伟达新加坡人工智能技术中心)
;
Department of Computer Science and Engineering, IIT Kanpur(IIT坎普尔计算机科学与工程系)
SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents
SciVisAgentBench:用于评估科学数据分析和可视化代理的基准测试
Kuangshi Ai, Haichao Miao, Kaiyuan Tang, Nathaniel Gorski, Jianxin Sun, Guoxi Liu, Helgi I. Ingolfsson, David Lenz, Hanqi Guo, Hongfeng Yu, Teja Leburu, Michael Molash, Bei Wang, Tom Peterka, Chaoli Wang, Shusen Liu
机构
*
University of Notre Dame(圣母大学)
;
Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室)
;
University of Utah(犹他大学)
;
University of Nebraska–Lincoln(内布拉斯加大学林肯分校)
;
The Ohio State University(俄亥俄州立大学)
;
Argonne National Laboratory(阿贡国家实验室)
;
Anthropic PBC
MicroEvo: Knowledge-Guided LLM Sampling for Efficient Microarchitecture Design Space Exploration
MicroEvo:知识引导的大语言模型采样用于高效微架构设计空间探索
Jia Xiong, Runkai Li, Chenxu Niu, Guangyuan Gao, Changwen Xing, Yifan Zhang, Xinlai Wan, Jieran Cui, Chen Bai, Yusheng Hua, Ying Wang, Ming Ling, Xi Wang, Tao Xie
机构
*
Southeast University(东南大学)
;
National Center of Technology Innovation for EDA(国家EDA技术创新中心)
;
NVIDIA Corporation(英伟达公司)
;
Fudan University(复旦大学)
;
Nanjing University of Posts and Telecommunications(南京邮电大学)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
;
Peking University(北京大学)
Journal refIn Findings of the Association for Computational Linguistics: ACL 2026, pages 35478-35507, San Diego, California, United States. Association for Computational Linguistics, 2026