Towards Conversational Medical AI with Eyes, Ears and a Voice
面向有眼睛、耳朵和声音的对话式医疗AI
Meet Shah, Jason Gusdorf, Anil Palepu, Chunjong Park, Jack W. O'Sullivan, Vishnu Ravi, Tim Strother, Pavel Dubov, Aliya Rysbek, Toshiyuki Fukuzawa, Yana Lunts, Jan Freyberg, Michael B. Chang, Aniruddh Raghu, David Stutz, Devora Berlowitz, Eliseo Papa, Taylan Cemgil, JD Velasquez, Jack Chen, Arthur Chen, Doug Fritz, Charlie Taylor, Katya Tregubova, Jing Rong Lim, Richard Green, Sara Mahdavi, Mahvish Nagda, Jihyeon Lee, Craig Schiff, Liviu Panait, Sukhdeep Singh, Valentin Liévin, David G. T. Barrett, Hannah Gladman, Anna Cupani, Francesca Pietra, Uchechi Okereke, Katherine Tong, Clemens Meyer, Erwan Rolland, Mili Sanwalka, Michael D. Howell, Shixiang Shane Gu, Bibo Xu, Euan A. Ashley, S. M. Ali Eslami, Gregory Wayne, Pushmeet Kohli, Vivek Natarajan, Adam Rodman, Alan Karthikesalingam, Ryutaro Tanno
机构
*
Google DeepMind(谷歌DeepMind)
;
Google Research(谷歌研究)
;
Beth Israel Deaconess Medical Center, Harvard Medical School(贝塞斯达医院, 哈佛医学院)
;
Stanford University(斯坦福大学)
Generating Leakage-Free Benchmarks for Robust RAG Evaluation
生成无泄漏的基准以评估鲁棒的RAG
Jiayi Liu, Jiaxing Zhang, Bowen Jin, Jennifer Neville
机构
*
Department of Computer Science, Purdue University(普渡大学计算机科学系)
;
New Jersey Institute of Technology(新泽西理工学院)
;
University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Microsoft Research(微软研究院)
Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment
基准是否低估了大语言模型的性能?通过大语言模型优先的人类仲裁评估来评估幻觉检测
I. F. Atasoy, B. Mutlu, E. A. Sezer, A. Wahdan
机构
*
Department of Computer Engineering, Hacettepe University(哈切塔佩大学计算机工程系)
;
Department of Computer Engineering, Ankara University(安卡拉大学计算机工程系)
;
Zephlen AI and Information Technologies Inc.(泽夫伦人工智能与信息技术公司)
Commentshis is the version of the article accepted for publication in SUMMA 2025 after peer review. The final, published version is available at IEEE Xplore: https://doi.org/10.1109/SUMMA68668.2025.11302248
Journal ref2025 7th International Conference on Control Systems, Mathematical Modeling, Automation and Energy Efficiency (SUMMA), Lipetsk, Russian Federation, 2025, pp. 799-804
机构
*
National Chengchi University(国立中正大学)
;
Georgetown University(乔治城大学)
;
University of Michigan(密歇根大学)
;
Stevens Institute of Technology(史蒂文斯理工学院)
;
National Taiwan University of Science and Technology(台湾科技大学)
;
Far Eastern Memorial Hospital(东方纪念医院)
Evaluating Large Language Models in Scientific Discovery
评估大型语言模型在科学发现中的表现
Zhangde Song, Jieyu Lu, Yuanqi Du, Botao Yu, Thomas M. Pruyn, Yue Huang, Kehan Guo, Xiuzhe Luo, Yuanhao Qu, Yi Qu, Yinkai Wang, Haorui Wang, Jeff Guo, Jingru Gan, Parshin Shojaee, Di Luo, Andres M Bran, Gen Li, Qiyuan Zhao, Shao-Xiong Lennon Luo, Yuxuan Zhang, Xiang Zou, Wanru Zhao, Yifan F. Zhang, Wucheng Zhang, Shunan Zheng, Saiyang Zhang, Sartaaj Takrim Khan, Mahyar Rajabi-Kochi, Samantha Paradi-Maropakis, Tony Baltoiu, Fengyu Xie, Tianyang Chen, Kexin Huang, Weiliang Luo, Meijing Fang, Xin Yang, Lixue Cheng, Jiajun He, Soha Hassoun, Xiangliang Zhang, Wei Wang, Chandan K. Reddy, Chao Zhang, Zhiling Zheng, Mengdi Wang, Le Cong, Carla P. Gomes, Chang-Yu Hsieh, Aditya Nandy, Philippe Schwaller, Heather J. Kulik, Haojun Jia, Huan Sun, Seyed Mohamad Moosavi, Chenru Duan
机构
*
Deep Principle(深原则)
;
Department of Computer Science, Cornell University(计算机科学系,康奈尔大学)
;
Department of Computer Science and Engineering, The Ohio State University(计算机科学与工程系,俄亥俄州立大学)
;
Department of Chemical Engineering & Applied Chemistry, University of Toronto(化学工程与应用化学系,多伦多大学)
;
Department of Computer Science and Engineering, University of Notre Dame(计算机科学与工程系,圣母大学)
;
QuEra Computing Inc.(QuEra计算公司)
;
Department of Pathology, Department of Genetics, Cancer Biology Program, Stanford University School of Medicine(病理学系、遗传学系、癌症生物学项目,斯坦福大学医学院)
;
Harvard Law School(哈佛法学院)
;
Department of Computer Science, Tufts University(计算机科学系,塔夫茨大学)
;
School of Computational Science and Engineering, Georgia Institute of Technology(计算科学与工程学院,佐治亚理工学院)
;
Department of Computer Science, University of California, Los Angeles(计算机科学系,加州大学洛杉矶分校)
;
Department of Computer Science, Virginia Tech(计算机科学系,弗吉尼亚理工大学)
;
Department of Physics, Tsinghua University(物理系,清华大学)
;
Institute for Advanced Study, Tsinghua University(清华大学高级研究所)
;
Laboratory of Artificial Chemical Intelligence, Ecole Polytechnique Federale de Lausanne(人工化学智能实验室,瑞士联邦理工学院)
ITS-Mina: A Harris Hawks Optimization-Based All-MLP Framework with Iterative Refinement and External Attention for Multivariate Time Series Forecasting
Pourya Zamanvaziri, Amirhossein Sadr, Aida Pakniyat, Dara Rahmati
机构
*
Department of Computer Science and Engineering, Shahid Beheshti University, Iran(伊朗谢赫·贝赫什提大学计算机科学与工程系)
;
School of Computer Science, Institute for Research in Fundamental Sciences (IPM), Iran(伊朗基础科学研究院(IPM)计算机科学学院)
Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs
视频中的奉承:视频大语言模型中奉承行为的基准测试与分析
Wenrui Zhou, Mohamed Hendy, Shu Yang, Qingsong Yang, Zikun Guo, Yuyu Luo, Lijie Hu, Di Wang
机构
*
Provable Responsible AI and Data Analytics (PRADA) Lab(可证负责任人工智能与数据 analytics 实验室)
;
King Abdullah University of Science and Technology(国王阿卜杜勒阿齐兹科学与技术大学)
;
HKUST(香港科技大学)
;
MBZUAI(穆罕默德·本·拉希德人工智能研究所)
;
University of Science and Technology of China(中国科学技术大学)
;
Kyungpook National University(庆尚国立大学)
A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron?
对计算机使用代理的安全性和安全威胁的综述:贾维斯或乌tron?
Ada Chen, Yongjiang Wu, Junyuan Zhang, Jingyu Xiao, Shu Yang, Jen-tse Huang, Kun Wang, Wenxuan Wang, Shuai Wang
机构
*
Carnegie Mellon University(卡内基梅隆大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
KAUST(卡塔尔科技大学)
;
Johns Hopkins University(约翰霍普金斯大学)
;
Nanyang Technological University(南洋理工大学)
;
Renmin University of China(中国人民大学)
;
The Hong Kong University of Science and Technology(香港科学大学)
Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers
条件性偏差:通用干预可以隐藏新兴偏差背后的情境触发
Jan Dubiński, Jan Betley, Anna Sztyber-Betley, Daniel Tan, Owain Evans
机构
*
Warsaw University of Technology(华沙技术大学)
;
NASK National Research Institute(NASK国家研究院)
;
Constellation
;
Truthful AI
;
University College London(伦敦大学学院)
;
Center on Long-Term Risk(长期风险中心)
;
UC Berkeley(伯克利大学)