Comments15 pages, 4 figures, 5 tables. Proof-of-concept evaluation of a task-centric chemistry ontology and deterministic rule engine on 300 human-authored and manually validated school-level problems. The proposed LLM translation layer and token-efficiency hypothesis are not evaluated in the current study
Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities
大型视觉语言模型(LVLMs)能否揭示视觉错觉背后的真相?感知与推理能力分析
Liangjie Zhao, Jiaqing Lyu, Kexin Tang, Zecheng Fang, Rong Yin, Yulan Hu, Da Li, Jianing Li
机构
*
Adelaide University(阿德莱德大学)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Tsinghua University(清华大学)
;
Amap, Alibaba Group(阿里巴巴集团高德地图)
;
Beihang University(北京航空航天大学)
Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning?
LLM能否准确评分医学诊断和临床推理?
Amy Rouillard, Sitwala Mundia, Linda Camara, Ziyaad Dangor, Michael Cameron Gramanie, Ismail Kalla, Shabir A. Madhi, Kajal Morar, Marlvin T. Ncube, Haroon Saloojee, Bruce A. Bassett
机构
*
Wits MIND Institute, University of the Witwatersrand, Johannesburg, South Africa(维特士心理研究所,沃斯兰德大学,约翰内斯堡,南非)
;
Grai Labs, Cape Town, South Africa(格雷实验室,开普敦,南非)
;
South African Medical Research Council Vaccines and Infectious Diseases Analytics Research Unit, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa(南非医学研究理事会疫苗和传染病分析研究组,健康科学学院,沃斯兰德大学,约翰内斯堡,南非)
;
Department of Internal Medicine, Charlotte Maxeke Johannesburg Academic Hospital, and Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa(内科学系,查理·马克斯凯约翰内斯堡学术医院,以及健康科学学院,沃斯兰德大学,约翰内斯堡,南非)
;
Department of Paediatrics and Child Health, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa(儿科学与儿童健康系,健康科学学院,沃斯兰德大学,约翰内斯堡,南非)
;
Wits MIND Institute, University of the Witwatersrand, Johannesbu(维特士心理研究所,沃斯兰德大学,约翰内斯堡)
CommentsThis submission is an iterative version of our previous work, **"AD^2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions"** (arXiv:2506.09557). We plan to consolidate the current submission with the earlier version into a unified manuscript. Therefore, we would like to withdraw this submission
RuleWeaver: Benchmarking Rule-Centered Scenario Reasoning for Large Language Models
RuleWeaver:面向大型语言模型的以规则为中心的场景推理基准测试
Bohan Yu, Shi-Yang Li, Pengfei Cao, Jun Zhao, Kang Liu
机构
*
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉科学学院)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
Comments15 pages, 8 figures and tables, accepted at the 2026 TPCTC Conference and will be published at Performance Evaluation and Benchmarking (Springer)
BALMS: Benchmarking Agentic LLMs for Longitudinal Mental Health Sensing
BALMS:面向纵向心理健康感知的智能体大语言模型基准测试
Yu Yvonne Wu, Arvind Pillai, Yuliang Chen, Yuwei Zhang, Sudarshan Regmi, Tess Z. Griffin, Michael V. Heinz, Lisa A. Marsch, Nicholas C. Jacobson, Andrew Campbell
机构
*
Dartmouth College(达特茅斯学院)
;
University of Cambridge(剑桥大学)
MedFabric: Gold Evidence Hides the Difficulty of Word-Level Medical Fabrication Detection
MedFabric:黄金证据掩盖了词级医学编造检测的难度
Tung Sum Thomas Kwok, Qian Qian, Xiaofeng Lin, Dongxu Zhang, Jun Han, Zhichao Yang, Davin Hill, Tamer Soliman, Sanjit Singh Batra, Robert Tillman, Guang Cheng