URSA: Chemistry-Aware Benchmark for Utilitarian Retrosynthesis Assessment
URSA:用于功利性逆合成评估的化学感知基准测试
Bogdan Zagribelnyy, Ivan Ilin, Nikita Bondarev, Anton Morgunov, Arkadii Lin, Maksim Kuznetsov, Rim Shayakhmetov, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov
机构
*
Insilico Medicine AI Limited(Insilico Medicine AI有限公司)
;
Independent researcher(独立研究者)
;
Insilico Medicine Canada Inc.(Insilico Medicine加拿大公司)
;
Insilico Medicine Hong Kong Ltd.(Insilico Medicine香港有限公司)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG
GuideMe: Multi-Domain Task Guidance and Intervention in Streaming Video
GuideMe:流视频中的多域任务指导与干预
Fang Liu, Jinpeng Chen, Ke Xu, Yuhao Liu, Huankang Guan, Xudong Lu, Bo Yang, Gerhard Hancke, Rui Liu, Rynson W. H. Lau
机构
*
City University of Hong Kong(香港城市大学)
;
Huawei Research(华为研究院)
;
University of Science and Technology of China(中国科学技术大学)
;
Chinese University of Hong Kong(香港中文大学)
;
City University of Hong Kong (Dongguan)(香港城市大学(东莞))
专题命中
评测与基准
:LLM(abstract);large language model(abstract);language model(abstract)
Comments28 pages, 2 figures, 13 tables. Benchmark, environment spec, and app contract released. First open-weight three-model sweep (k=5) on a 40-task oracle-validated executable suite; frontier-model leaderboard committed in the roadmap
Comments17 pages, 6 figures, 4 tables. Accepted at AIES 2026 (AAAI/ACM Conference on AI, Ethics, and Society). This version includes the supplementary appendix. Code and data: https://github.com/mbrcic/llm-political-steerability (Zenodo DOI 10.5281/zenodo.21489805)
Evaluating the Diagnostic Robustness of Vision-Language Models Under Visual and Textual Perturbations
评估视觉-语言模型在视觉和文本扰动下的诊断鲁棒性
Ali Khoramfar, Mohammad Javad Dousti, Alireza Mohamadian, Heshaam Faili
机构
*
University of Tehran(德黑兰大学)
;
Tehran University of Medical Sciences(德黑兰医科大学)
;
Advanced Diagnostic and Interventional Radiology Research Center (ADIR)(高级诊断与介入放射学研究中心(ADIR))