CommentsRevised version following peer review. Expanded methodological details, practitioner-survey validation, statistical analyses, and discussion; main conclusions unchanged
机构
*
Nanjing University(南京大学)
;
The Hong Kong University of Science and Technology(香港科技大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
Comparative Validation of GPT-4o-mini and Teacher Mean Scores for Automated Scoring of Music Analysis Responses: Single-Pass Deployment, Repeatability, and Strategy-Specific Bias
Measuring in-context algorithmic reasoning in language models against an exact Bayes-optimal standard
基于精确贝叶斯最优标准测量语言模型的上下文内算法推理能力
Hector Zenil, Luan Ozelim
机构
*
Oxford Immune Algorithmics(牛津免疫算法学公司)
;
Oxford University Innovation(牛津大学创新公司)
;
London Institute for Healthcare Engineering(伦敦医疗工程研究所)
;
King’s College London(伦敦国王学院)
CommentsThis paper has been withdrawn by the authors due to a significant bug discovered in our data processing pipeline. This bug affects the validity of the experimental results, and we can no longer stand by the conclusions presented
SPARC-Rad: A Multimodal Benchmark Dataset and Evaluation Pipeline for Spatial and Anatomical Reasoning in Radiology Vision-Language Models
SPARC-Rad:面向放射学视觉语言模型的空间与解剖推理多模态基准数据集及评估流程
Satvik Tripathi, Mustafa Ege Seker, Kristian Quevada, Ebubechukwu D Enwerem, Pratham Khandelwal, Emine Meltem, Bera Koca, Shahriar Faghani, Jacinta Arnold, Dania Daye, Tessa S. Cook
机构
*
Department of Data Science and Hong Kong Institute of AI for Science, City University of Hong Kong(数据科学系和香港人工智能科学研究所,香港城市大学)
;
Li Auto Inc., China(中国利汽车公司)
;
Department of Statistics, University of Oxford(统计系,牛津大学)
机构
*
University of Isfahan(伊斯法罕大学)
;
University of Tehran(德黑兰大学)
;
University of Windsor(温莎大学)
;
Alzahra University(阿勒扎哈拉大学)
;
University of Texas at Dallas(德克萨斯大学达拉斯分校)