机构
*
Universidad de Buenos Aires, Facultad de Ciencias Exactas y Naturales, Departamento de Computación(布宜诺斯艾利斯大学,精确与自然科学学院,计算机系)
;
AI Safety Argentina (AISAR)(阿根廷人工智能安全组织 (AISAR))
;
Department of Computer Science, University of Oxford(牛津大学计算机科学系)
;
CONICET-Universidad de Buenos Aires, Instituto de Ciencias de la Computación (ICC)(阿根廷国家科学与技术研究理事会-布宜诺斯艾利斯大学,计算机科学研究所 (ICC))
From Knowing to Acting: Benchmarking Self-Awareness Capability of LLM Agents
从知道到行动:基准测试LLM代理的自我意识能力
Yifan Li, Shengbin Yue, Boyu Feng, Jinhu Qi, Bo Ke, Zixing Song, Hongru Wang, Zhongyu Wei, Irwin King
机构
*
The Chinese University of Hong Kong(香港中文大学)
;
Fudan University(复旦大学)
;
University of Edinburgh(爱丁堡大学)
;
Tencent(腾讯)
;
University of Bristol(布里斯托大学)
OGD4All: A Framework for Accessible Interaction with Geospatial Open Government Data Based on Large Language Models
OGD4All: 基于大语言模型的可访问地理空间开放政府数据交互框架
Michael Siebenmann, Javier Argota Sánchez-Vaquerizo, Stefan Arisona, Krystian Samp, Luis Gisler, Dirk Helbing
机构
*
Professorship of Computational Social Science, ETH Zurich(计算社会科学教授职位,苏黎世联邦理工学院)
;
Esri R&D Center Zurich(埃斯里苏黎世研发中心)
;
Complexity Science Hub(复杂性科学中心)
CommentsAuthor Accepted Manuscript (AAM). Proceedings of 2026 IEEE CAI (Granada, Spain). Update manuscript with final DOI. Code & data available at: https://github.com/ethz-coss/ogd4all
Journal ref2026 IEEE Conference on Artificial Intelligence (CAI), pp. 882-888
On the Adversarial Robustness of Multimodal LLM Judges
多模态大语言模型评判器的对抗鲁棒性
Zihan Wang, Guansong Pang, Zelin Liu, Wenjun Miao, Jin Zheng, Xiao Bai
机构
*
School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院)
;
State Key Laboratory of Virtual Reality Technology and System, Beihang University(北京航空航天大学虚拟现实技术与系统国家重点实验室)
;
State Key Laboratory of Software Development Environment, Jiangxi Research Institute, Beihang University(北京航空航天大学江西研究院软件开发环境国家重点实验室)
;
School of Computing and Information Systems, Singapore Management University(新加坡管理大学计算机与信息系统学院)
OpenMedReason: Scientific Reasoning Supervision for Medical Vision-Language Models
OpenMedReason: 医学视觉语言模型的科学推理监督
Negin Baghbanzadeh, Pritam Sarkar, Michael Colacci, Abeer Badawi, Adibvafa Fallahpour, Arash Afkanpour, Leonid Sigal, Ali Etemad, Elham Dolatabadi
机构
*
York University(约克大学)
;
Vector Institute(向量研究所)
;
University of British Columbia(不列颠哥伦比亚大学)
;
University of Toronto(多伦多大学)
;
Unity Health Toronto / St. Michael’s Hospital(多伦多联合健康/圣迈克尔医院)
;
University Health Network(大学健康网络)
;
Arc Institute(弧研究所)
;
Queen's University(女王大学)
PreAct-Bench: Benchmarking Predictive Monitoring in LLMs
PreAct-Bench:大语言模型中的预测性监控基准
Hainiu Xu, Italo Luis da Silva, Jiangnan Ye, Yuhao Wang, Wei Liu, Linyi Yang, Jonathan Richard Schwarz, Nicola Paoletti, Yulan He, Hanqi Yan
机构
*
King’s College London(伦敦国王学院)
;
National University of Singapore(新加坡国立大学)
;
Southern University of Science and Technology(南方科技大学)
;
Thomson Reuters Foundational Research(汤姆森路透基础研究)
;
Imperial College London(伦敦帝国学院)
;
The Alan Turing Institute(艾伦·图灵研究所)
CommentsICAHS, \c{opyright} 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
Subtitle-Aligned Fine-Tuning of Whisper for Swiss German ASR: Benchmark Contamination, Convention Mismatch, and an Honest Baseline at 25.6% WER (13.8% cWER)
CommentsWithdrawn by the authors due to pending intellectual property considerations. The authors have determined that the current version contains material that should not have been publicly disseminated at this stage
Changling Li, Terry Jingchen Zhang, Jie Zhang, Zhijing Jin, Sahar Abdelnabi, Maksym Andriushchenko
机构
*
ETH Zürich(苏黎世联邦理工学院)
;
ELLIS Institute Tübingen(图宾根ELLIS研究所)
;
Max Planck Institute for Intelligent Systems(智能系统马克斯·普朗克研究院)
;
Tübingen AI Center(图宾根人工智能中心)
;
University of Toronto & Vector Institute(多伦多大学及向量研究所)
;
EuroSafeAI(欧洲安全人工智能)
CommentsAccepted to RLEval @ ACM CAIS 2026 (Workshop on Methods and RL Environments for Evaluating AI Agents) and selected for an invited talk based on reviewer ratings. 4-page short paper + appendix