Yu Pan, Zhongze Cai, Guanting Chen, Huaiyang Zhong, Chonghuan Wang
机构
*
University of Sydney(悉尼大学)
;
Imperial College London(伦敦帝国理工学院)
;
University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
;
Virginia Tech(弗吉尼亚理工大学)
;
University of Texas at Dallas(德克萨斯大学达拉斯分校)
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
Amr Gomaa, Ahmed Salem, Sahar Abdelnabi
机构
*
German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI))
;
Microsoft(微软)
;
ELLIS Institute Tübingen and MPI for Intelligent Systems(图宾根ELLIS研究所和智能系统研究所)
;
Tübingen AI Center(图宾根人工智能中心)
Yu Bai, Yukai Miao, Dawei Wang, Li Chen, Fei Long, Rundi Zhai, Dan Li, Yanyu Ren, Tianfeng Liu, Hongtao Xie, Ce Yang, Xuhui Cai
机构
*
Zhongguancun Laboratory(中关村实验室)
;
Tsinghua University(清华大学)
;
China Mobile Communications Group Co., Ltd.(中国移动通信集团有限公司)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
Investigating U.S. Consumer Demand for Food Products with Innovative Transportation Certificates Based on Stated Preferences and Machine Learning Approaches
LEME: Open Large Language Models for Ophthalmology with Advanced Reasoning and Clinical Validation
Hyunjae Kim, Xuguang Ai, Sahana Srinivasan, Aidan Gilson, Maxwell B. Singer, Krithi Pushpanathan, Qianqian Xie, Jungwoo Park, Serina Applebaum, Gabriel Dawei Yang, Minjie Zou, David Ziyou Chen, Ke Zou, Soshian Sarrafpour, Ji Liu, Yu Yin, Jimin Huang, Quang Ngoc Nguyen, Erping Long, Peixing Wan, Dianbo Liu, Richard Hintz, W. Jim Zheng, Sophia Y. Wang, Lucila Ohno-Machado, Hua Xu, Ron A. Adelman, Luciano V. Del Priore, Yih-Chung Tham, Qingyu Chen
机构
*
School of Electronic and Electrical Engineering, Shanghai University of Engineering Science(上海工程技术大学电子与电气工程学院)
;
Tencent Youtu Lab(腾讯优图实验室)
;
Artificial Intelligence Innovation and Incubation Institute, Fudan University(复旦大学人工智能创新与孵化院)
XBreaking: Understanding how LLMs security alignment can be broken
Marco Arazzi, Vignesh Kumar Kembu, Antonino Nocera, Vinod P
机构
*
Department of Electrical, Computer
;
Biomedical Engineering, University of Pavia, Italy\ .
;
Department of Computer Applications,\ University of Science \& Technology, India\ .
GUARD: Role-playing to Generate Natural-language Jailbreakings to Test Guideline Adherence of Large Language Models
Haibo Jin, Ruoxi Chen, Peiyan Zhang, Andy Zhou, Haohan Wang
机构
*
School of Information Sciences University of Illinois at Urbana-Champaign(信息科学学院伊利诺伊大学厄巴纳-香槟分校)
;
Independent Researcher, Starc Institute(Starc研究所独立研究者)
;
Computer Science and Engineering HKUST(HKUST计算机科学与工程学院)
;
Computer Science Lapis Labs University of Illinois Urbana-Champaign(计算机科学Lapis Labs伊利诺伊大学厄巴纳-香槟分校)
;
School of Information Sciences University of Illinois Urbana-Champaign(信息科学学院伊利诺伊大学厄巴纳-香槟分校)
Beta Distribution Learning for Reliable Roadway Crash Risk Assessment
Ahmad Elallaf, Nathan Jacobs, Xinyue Ye, Mei Chen, Gongbo Liang
机构
*
Texas A&M University-San Antonio(德克萨斯A&M大学-圣安东尼奥分校)
;
Washington University in St. Louis(华盛顿大学圣路易斯分校)
;
University of Alabama(阿拉巴马大学)
;
University of Kentucky(肯塔基大学)
HugAgent: Benchmarking LLMs for Simulation of Individualized Human Reasoning
Chance Jiajie Li, Zhenze Mo, Yuhan Tang, Ao Qu, Jiayi Wu, Kaiya Ivy Zhao, Yulu Gan, Jie Fan, Jiangbo Yu, Hang Jiang, Paul Pu Liang, Jinhua Zhao, Luis Alberto Alonso Pastor, Kent Larson
机构
*
MIT Media Lab(MIT媒体实验室)
;
MIT EECS(MIT电子工程与计算机科学系)
;
MIT BCS(MIT生物科学系)
;
MIT IDSS(MIT国际设计系统研究所)
;
MIT CEE(MIT土木与环境工程系)
;
MIT DUSP(MIT设计学教授职位)
;
MIT Architecture(MIT建筑系)
;
Northeastern University(东北大学)
;
Brown University(布朗大学)
;
McGill University(麦吉尔大学)
Measuring what Matters: Construct Validity in Large Language Model Benchmarks
Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou, Franziska Sofia Hafner, Harry Mayne, Jan Batzner, Negar Foroutan, Chris Schmitz, Karolina Korgul, Hunar Batra, Oishi Deb, Emma Beharry, Cornelius Emde, Thomas Foster, Anna Gausen, María Grandury, Simeng Han, Valentin Hofmann, Lujain Ibrahim, Hazel Kim, Hannah Rose Kirk, Fangru Lin, Gabrielle Kaili-May Liu, Lennart Luettgau, Jabez Magomere, Jonathan Rystrøm, Anna Sotnikova, Yushi Yang, Yilun Zhao, Adel Bibi, Antoine Bosselut, Ronald Clark, Arman Cohan, Jakob Foerster, Yarin Gal, Scott A. Hale, Inioluwa Deborah Raji, Christopher Summerfield, Philip H. S. Torr, Cozmin Ududec, Luc Rocher, Adam Mahdi
机构
*
University of Oxford(牛津大学)
;
EPFL(苏黎世联邦理工学院)
;
Weizenbaum Institute Berlin(柏林Weizenbaum研究所)
;
Technical University Munich(慕尼黑技术大学)
;
Centre for Digital Governance, Hertie School(赫尔姆霍兹学院数字治理中心)
;
Stanford University(斯坦福大学)
;
UK AI Security Institute(英国人工智能安全研究所)
;
SomosNLP
;
Universdad Politécnica de Madrid(马德里理工大学)
;
Yale University(耶鲁大学)
;
Allen Institute for AI(人工智能研究所)
;
University of Washington(华盛顿大学)
;
Meedan
;
UC Berkeley(加州大学伯克利分校)
专题命中
安全评测
:safety(abstract);分类 cs.CL、cs.AI
Comments39th Conference on Neural Information Processing Systems (NeurIPS 2025) Track on Datasets and Benchmarks
LiveStar: Live Streaming Assistant for Real-World Online Video Understanding
Zhenyu Yang, Kairui Zhang, Yuhang Hu, Bing Wang, Shengsheng Qian, Bin Wen, Fan Yang, Tingting Gao, Weiming Dong, Changsheng Xu
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
ShanghaiTech University(上海科技大学)
;
Kuaishou Technology(快手科技)
;
Peng Cheng Laboratory(鹏城实验室)
GAITEX: Human motion dataset of impaired gait and rehabilitation exercises using inertial and optical sensors
Andreas Spilz, Heiko Oppel, Jochen Werner, Kathrin Stucke-Straub, Felix Capanni, Michael Munz
机构
*
AI for Sensor Data Analytics Research Group(人工智能传感器数据解析研究组)
;
Ulm University of Applied Sciences(乌尔姆应用科学大学)
;
Biomechatronic Research Group(生物机械研究组)
;
Institute of Computer Science(计算机科学研究所)
ExGra-Med: Extended Context Graph Alignment for Medical Vision-Language Models
Duy M. H. Nguyen, Nghiem T. Diep, Trung Q. Nguyen, Hoang-Bao Le, Tai Nguyen, Tien Nguyen, TrungTin Nguyen, Nhat Ho, Pengtao Xie, Roger Wattenhofer, James Zou, Daniel Sonntag, Mathias Niepert
机构
*
German Research Centre for Artificial Intelligence (DFKI)(德国人工智能研究中心)
;
Max Planck Research School for Intelligent Systems (IMPRS-IS)(马克斯·普朗克智能系统研究学校)
;
University of Stuttgart(斯图加特大学)
;
University Medical Center Gottingen(哥廷根大学医学中心)
;
Max Planck Institute for Multidisciplinary Sciences(马克斯·普朗克多学科科学研究所)
;
ARC Centre of Excellence for the Mathematical Analysis of Cellular Systems(细胞系统数学分析卓越中心)
;
School of Mathematical Sciences, Queensland University of Technology(昆士兰科技大学数学科学学院)
;
University of Oldenburg(奥尔登堡大学)
;
University of Texas at Austin(德克萨斯大学奥斯汀分校)
;
University of California San Diego(加州大学圣地亚哥分校)
;
MBZUAI(马克斯·普朗克人工智能研究所)
;
ETH Zurich(苏黎世联邦理工学院)
;
Stanford University(斯坦福大学)