Comments20 pages, 3 figures, 5 tables. Accepted at the 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). To appear in LIPIcs Vol. 394. v2: corrected the bibliographic record of one reference (preprint, not a journal article) and added the related-version link to the published LIPIcs article
TimeSage-EV: A Live Benchmark for Agentic Time Series Analysis in Evolving Environments
TimeSage-EV:面向动态环境下智能体时间序列分析的实时基准测试集
Qingren Yao, Yaxuan Kong, Yuqi Nie, Yichen Li, Stefan Zohren, Anna Vettoruzzo, Qingsong Wen, Ming Jin, Joaquin Vanschoren
机构
*
Eindhoven University of Technology(埃因霍温理工大学)
;
University of Oxford(牛津大学)
;
VulpiVox Intelligence(VulpiVox智能公司)
;
Squirrel Ai Learning(Squirrel AI学习公司)
;
Griffith University(格里菲斯大学)
专题命中
评测与基准
:LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI
Foundation Models in Biomedical Imaging: Turning Hype into Reality
生物医学影像中的基础模型:从 hype 到现实
Amgad Muneer, Kai Zhang, Ibraheem Hamdi, Rizwan Qureshi, Muhammad Waqas, Shereen Fouad, Hazrat Ali, Syed Muhammad Anwar, Jia Wu
机构
*
Department of Imaging Physics, The University of Texas MD Anderson Cancer Center(影像物理系,德克萨斯大学MD安德森癌症中心)
;
Center for Secure Artificial Intelligence for Healthcare (SAFE), McWilliams School of Biomedical Informatics, UTHealth Houston(安全人工智能用于医疗保健中心(SAFE),麦威廉斯生物医学信息学学院,UTHealth休斯顿)
;
Female Medicine in Machine Learning, Massachusetts Institute of Technology(机器学习中的女性医学,麻省理工学院)
;
Department of Computer Science, Salim Habib University(计算机科学系,Salim Habib大学)
;
School of Computer Science and Digital Technologies, Aston Centre for Artificial Intelligence Research and Application, Aston University(计算机科学与数字技术学院,阿斯顿人工智能研究与应用中心,阿斯顿大学)
;
Division of Computing Science and Mathematics, University of Stirling(计算科学与数学系,斯特灵大学)
;
School of Medicine and Health Sciences, George Washington University(医学与健康科学学院,乔治·华盛顿大学)
;
Sheikh Zayed Institute for Pediatric Surgical Innovation, Children’s National Hospital(谢赫扎耶德儿童外科创新研究所,儿童医院)
;
Department of Thoracic/Head and Neck Medical Oncology, The University of Texas MD Anderson Cancer Center(胸腔/头颈医学肿瘤科,德克萨斯大学MD安德森癌症中心)
AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs
AnchorBench:针对大语言模型中锚定效应的多路径基准测试
Yiderigun Borjigin, Alexander Hermann, Christian Cyron, Roland Aydin
机构
*
Saarland University(萨尔大学)
;
Hamburg University of Technology(汉堡工业大学)
;
Helmholtz-Zentrum Hereon(亥姆霍兹中心赫伦)
;
German Research Centre for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI))
专题命中
评测与基准
:LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI