机构
*
School of Interactive Computing, Georgia Institute of Technology(交互计算学院,佐治亚理工学院)
;
Amazon(亚马逊)
;
Siebel School of Computing and Data Science, University of Illinois Urbana-Champaign(Siebel计算与数据科学学院,伊利诺伊大学厄巴纳-香槟分校)
机构
*
Zhejiang University(浙江大学)
;
Zhongguancun Academy(中关村学院)
;
University of Science and Technology of China(中国科学技术大学)
;
National University of Singapore(新加坡国立大学)
Commentsv2: Adds an adaptive-attacker evaluation in which Ring 1 is fully evaded (500/500 documents, three corpora); scales to the full NQ, HotpotQA and MS-MARCO corpora; corrects Proposition 1, whose boundary is corpus-dependent (0.214/0.251/0.558) not 0.5; retracts a proposed closed form after a pre-registered prediction failed. The v1 headline 91%-to-13% result is withdrawn
ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents
ProvenanceGuard: 基于MCP的LLM智能体的源感知事实性验证
Ander Alvarez, Santhiya Rajan, Alessandro Genuardi, Oliver Wirjadi, Samuel Mugel, Román Orús
机构
*
Multiverse Computing
;
Parque Cientifico y Tecnológico de Gipuzkoa(吉普斯夸科技园)
;
Centre for Social Innovation(社会创新中心)
;
Donostia International Physics Center(多诺斯蒂亚国际物理中心)
;
Ikerbasque Foundation for Science(伊克尔巴斯克科学基金会)
CommentsThis submission is an iterative version of our previous work, **"AD^2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions"** ( arXiv:2506.09557 (https://arxiv.org/abs/2506.09557) ). We plan to consolidate the current submission with the earlier version into a unified manuscript. Therefore, we would like to withdraw this submission
Distinct Profiles of Run-to-Run Score Reliability and Expert-Panel Alignment Across Four LLM Evaluators of Simulated Japanese-Language AI-to-AI Counseling
机构
*
Japan National Institute of Occupational Safety and Health(日本国立职业安全卫生研究所)
;
Kaze To Taiyo(凯泽・太阳)
;
Saga Occupational Health Association(Saga职业健康协会)
;
Department of Pharmacy, Zikei Hospital/Zikei Institute of Psychiatry(药剂科,Zikei医院/Zikei精神医学研究所)
;
Department of Medical Welfare, Suzuka University of Medical Science(医疗福祉科, Suzuka医科大学)
;
Graduate School of Human Sciences, Ritsumeikan University(人类科学研究生院,立命馆大学)
;
Faculty of Nursing, National Defense Medical College(护理学部,国家防卫医疗大学)
;
Support Center for Students with Disabilities, Aoyama Gakuin University(残疾学生支持中心,上智大学)
Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning?
LLM能否准确评分医学诊断和临床推理?
Amy Rouillard, Sitwala Mundia, Linda Camara, Ziyaad Dangor, Michael Cameron Gramanie, Ismail Kalla, Shabir A. Madhi, Kajal Morar, Marlvin T. Ncube, Haroon Saloojee, Bruce A. Bassett
机构
*
Wits MIND Institute, University of the Witwatersrand, Johannesburg, South Africa(维特士心理研究所,沃斯兰德大学,约翰内斯堡,南非)
;
Grai Labs, Cape Town, South Africa(格雷实验室,开普敦,南非)
;
South African Medical Research Council Vaccines and Infectious Diseases Analytics Research Unit, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa(南非医学研究理事会疫苗和传染病分析研究组,健康科学学院,沃斯兰德大学,约翰内斯堡,南非)
;
Department of Internal Medicine, Charlotte Maxeke Johannesburg Academic Hospital, and Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa(内科学系,查理·马克斯凯约翰内斯堡学术医院,以及健康科学学院,沃斯兰德大学,约翰内斯堡,南非)
;
Department of Paediatrics and Child Health, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa(儿科学与儿童健康系,健康科学学院,沃斯兰德大学,约翰内斯堡,南非)
;
Wits MIND Institute, University of the Witwatersrand, Johannesbu(维特士心理研究所,沃斯兰德大学,约翰内斯堡)
机构
*
State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(人工智能安全国家重点实验室,计算技术研究所,中国科学院)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Baidu Inc(百度公司)