DrugClaw and DrugAudit: A Primary-Source-Grounded Agent and Authority-Aware Benchmark for Drug-Information Question Answering
DrugClaw与DrugAudit:基于原始来源的智能体与权威感知基准用于药物信息问答
Qing Wang, Bo Li, Jialu Liang, Daling Shi, Bob Zhang, Qianqian Song
机构
*
Department of Health Outcomes and Biomedical Informatics, College of Medicine, University of Florida(佛罗里达大学健康结局与生物医学信息学系,医学院)
;
PAMI Research Group, Department of Computer and Information Science, Faculty of Science and Technology, University of Macau(澳门大学科学与技术学院计算机与信息科学系PAMI研究组)
VeriTrip: A Verifiable Benchmark for Travel Planning Agents over Unstructured Web Corpora
VeriTrip: 面向非结构化网络语料的旅行规划智能体可验证基准
Yuting Xu, Jiayi Tian, Jian Liang, Xin Xiong, Hang Zhang, Mu Xu, Xiao-Yu Zhang
机构
*
Institute of Information Engineering, CAS(中国科学院信息工程研究所)
;
School of Cyber Security, UCAS(中国科学技术大学网络安全学院)
;
Amap, Alibaba Group(阿里巴巴集团阿地图)
;
NLPR & MAIS, Institute of Automation, CAS(中国科学院自动化研究所神经信息处理实验室及机器智能研究所)
;
School of Artificial Intelligence, UCAS(中国科学技术大学人工智能学院)
Evidence-Grounded Frontier Mapping and Agentic Hypothesis Generation in Nanomedicine
基于证据的前沿映射与代理假设生成在纳米医学中
Christiaan G. A. Viviers, Koen de Bruin, Mirre M. Trines, Ayla M. Hokke, Roy van der Meel, Avi Schroeder, Twan Lammers, Willem J. M. Mulder, Fons van der Sommen
机构
*
ARIA Lab, Signal Processing Systems, Department of Electrical Engineering, Eindhoven University of Technology(ARIA实验室,信号处理系统,电气工程系,埃因霍温理工大学)
;
Laboratory of Chemical Biology, Department of Biomedical Engineering, Eindhoven University of Technology(化学生物学实验室,生物医学工程系,埃因霍温理工大学)
;
The Louis Family Laboratory for Targeted Drug Delivery and Personalized Medicine Technologies, Department of Chemical Engineering, Technion - Israel Institute of Technology(定向药物输送与个性化医学技术实验室,化学工程系,技术离子-以色列理工学院)
;
Department of Nanomedicine and Theranostics, Institute for Experimental Molecular Imaging (ExMI), RWTH Aachen University Hospital(纳米医学与诊疗学系,实验分子成像研究所(ExMI),亚琛工业大学医院)
;
Department of Internal Medicine and Radboud Center for Infectious Diseases (RCI), Radboud University Medical Center(内科学系和Radboud感染疾病中心(RCI),Radboud大学医学中心)
AgenticEval: Toward Agentic and Self-Evolving Safety Evaluation of Large Language Models
AgenticEval: 向大型语言模型的代理和自演化安全评估迈进
Yixu Wang, Xin Wang, Yang Yao, Xinyuan Li, Xibang Yang, Yan Teng, Xingjun Ma, Yingchun Wang
机构
*
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Fudan University(复旦大学)
;
The University of Hong Kong(香港大学)
;
East China Normal University(华东师范大学)
机构
*
Kavli Institute for Theoretical Physics, University of California Santa Barbara(加州大学圣芭芭拉分校凯文利理论物理研究所)
;
Center for Gravitational Physics, University of Texas at Austin(德克萨斯大学奥斯汀分校重力物理中心)
;
Department of Physics, University of California at Santa Barbara(加州大学圣芭芭拉分校物理系)
;
International Centre for Theoretical Sciences, Tata Institute of Fundamental Research(塔塔基础研究机构国际理论科学研究中心)
;
School of Natural Sciences, Institute for Advanced Study(高级研究院自然科学学院)
;
Chennai Mathematical Institute(钦奈数学研究所)
;
Kavli Institute for Cosmological Physics, The University of Chicago(芝加哥大学凯文利宇宙物理研究所)
;
Department of Particle Physics & Astrophysics, Weizmann Institute of Science(魏茨曼科学研究所粒子物理与天体物理系)
MDGYM: Benchmarking AI Agents on Molecular Simulations
MDGYM:在分子模拟上评估AI代理的基准测试
Vinay Kumar, Satyendra Rajput, Mausam, N. M. Anoop Krishnan
机构
*
Yardi School of Artificial Intelligence, Indian Institute of Technology Delhi(印度理工学院德里人工智能学院)
;
Department of Computer Science and Engineering, Indian Institute of Technology Delhi(印度理工学院德里计算机科学与工程系)
;
Department of Civil and Environmental Engineering, Indian Institute of Technology Delhi(印度理工学院德里土木与环境工程系)
机构
*
Department of Computer Science, University of British Columbia, Vancouver, Canada(英属哥伦比亚大学计算机科学系)
;
Department of Artificial Intelligence, IIT Hyderabad, Hyderabad, India(印度海得拉巴理工学院人工智能系)
;
Amazon, Delhi, India(印度德里亚马逊公司)
AI Identity: Standards, Gaps, and Research Directions for AI Agents
AI身份:面向AI代理的标准、缺口与研究方向
Takumi Otsuka, Kentaroh Toyoda, Alex Leung
机构
*
AIFT, Singapore(新加坡AIFT)
;
Graduate School of Fundamental Science and Engineering, Department of Computer Science and Communications Engineering, Waseda University, Japan(日本早稻田大学基础科学与工程研究生院计算机科学与通信工程系)
;
Keio Global Research Institute (KGRI), Japan(日本庆应大学全球研究机构)
AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum
AIT Academy:通过儒家三领域课程培养完整智能体
Jiaqi Li, Lvyang Zhang, Yang Zhao, Wen Lu, Lidong Zhai
机构
*
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China(中国科学院大学网络安全学院)
机构
*
Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
University of Science and Technology of China(中国科学技术大学)
;
Peking University(北京大学)
Governance-Constrained Agentic AI: Blockchain-Enforced Human Oversight for Safety-Critical Wildfire Monitoring
受治理约束的代理AI:区块链强制的人类监督用于安全关键的野火监测
Ali Akarma, Toqeer Ali Syed, Salman Jan, Hammad Muneer, Abdul Khadar Jilani
机构
*
AI Center, Faculty of Computer and Information Systems, Islamic University of Madinah(伊斯兰大学麦地那分校计算机与信息系统学院人工智能中心)
;
Faculty of Computer Studies, Arab Open University-Bahrain(阿拉伯开放大学巴林分校计算机研究学院)
;
Department of Computer Science, The Islamia University of Bahawalpur(巴哈瓦尔布尔伊斯兰大学计算机科学系)
;
College of Computer Studies, University of Technology Bahrain(巴林科技大学计算机研究学院)
Quantifying Trust: Financial Risk Management for Trustworthy AI Agents
量化信任:为可信AI代理的金融风险管理
Wenyue Hua, Tianyi Peng, Chi Wang, Ian Kaufman, Bryan Lim, Chandler Fang
机构
*
Microsoft Research(微软研究院)
;
Columbia University(哥伦比亚大学)
;
Google DeepMind(谷歌DeepMind)
;
Stanford University(斯坦福大学)
;
t54.ai
;
Virtuals ACP
;
University of California, Santa Barbara(加州大学圣塔芭芭拉分校)