Has Multimodal Learning Delivered Universal Intelligence in Healthcare? A Comprehensive Survey
多模态学习是否在医疗领域实现了通用智能?一项全面的综述
Qika Lin, Yifan Zhu, Xin Mei, Ling Huang, Jingying Ma, Kai He, Zhen Peng, Erik Cambria, Mengling Feng
机构
*
Saw Swee Hock School of Public Health, National University of Singapore(新加坡国立大学公共健康学院)
;
School of Computer Science, Beijing University of Posts and Telecommunications(北京邮电大学计算机学院)
;
School of Automation, Northwestern Polytechnical University(西北工业大学自动化学院)
;
School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院)
BLUFF: Benchmarking the Detection of False and Synthetic Content across 58 Low-Resource Languages
BLUFF:跨58种低资源语言的虚假和合成内容检测基准测试
Jason Lucas, Matt Murtagh-White, Adaku Uchendu, Ali Al-Lawati, Michiharu Yamashita, Dominik Macko, Ivan Srba, Robert Moro, Dongwon Lee
机构
*
Penn State University(宾夕法尼亚州立大学)
;
Trinity College Dublin(都柏林圣三一学院)
;
MIT Lincoln Lab(麻省理工学院林肯实验室)
;
Visa Research(Visa研究)
;
Kempelen Institute of Intelligent Technologies(凯普勒智能技术研究所)
OPGAgent: An Agent for Auditable Dental Panoramic X-ray Interpretation
OPGAgent: 一种用于可审计牙科全景X光解读的智能体
Zhaolin Yu, Litao Yang, Ben Babicka, Ming Hu, Jing Hao, Anthony Huang, James Huang, Yueming Jin, Jiasong Wu, Zongyuan Ge
机构
*
AIM for Health Lab
;
Faculty of Information Technology, Monash University(信息科技学院,莫纳什大学)
;
Monash University(莫纳什大学)
;
Curae Health
;
Faculty of Dentistry, The University of Hong Kong(牙科学院,香港大学)
;
National University of Singapore(新加坡国立大学)
;
Southeast University(东南大学)
RTLocating: Intent-aware RTL Localization for Hardware Design Iteration
RTLocating:面向硬件设计迭代的意图感知RTL定位
Changwen Xing, Yanfeng Lu, Lei Qi, Chenxu Niu, Jie Li, Xi Wang, Yong Chen, Jun Yang
机构
*
School of Integrated Circuits, Southeast University, Nanjing, China(东南大学集成电路学院)
;
National Center of Technology Innovation for EDA, Nanjing, China(EDA技术创新国家中心)
;
School of Computer Science and Engineering, Southeast University, Nanjing, China(东南大学计算机科学与工程学院)
;
Department of Computer Science, Texas Tech University, Lubbock, USA(塔拉斯大学计算机科学系)
CommentsEqual Contribution: Xiaochuang Yuan and Hui Xu contributed equally to this work. All correspondence should be directed to yxc20098@gmail.com. Submitted to Agents in the Wild Workshop, ICLR2026
机构
*
University of Texas(德克萨斯大学)
;
Dell Children’s Medical Center(德尔儿童医学中心)
;
University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
;
University of Nevada, Reno(内华达大学里诺分校)
机构
*
Department of Informatics and Telematics, Harokopio University of Athens(信息与电信学系,雅典哈罗科波斯大学)
;
Department of Networks and Digital Media, Kingston University(网络与数字媒体系,金士顿大学)
;
Department of Electrical and Computer Engineering, University of Western Macedonia(电子与计算机工程系,西马其顿大学)
;
Archimedes, Athena Research Center(阿基米德,雅典研究中心)
GroundedSurg: A Multi-Procedure Benchmark for Language-Conditioned Surgical Tool Segmentation
GroundedSurg: 一种多手术流程的语言条件手术工具分割基准
Tajamul Ashraf, Abrar Ul Riyaz, Wasif Tak, Tavaheed Tariq, Sonia Yadav, Moloud Abdar, Janibul Bashir
机构
*
King Abdullah University of Science and Technology (KAUST)(卡奥尔大学科学与技术学院)
;
Thapar Institute of Engineering and Technology(塔帕尔工程与技术学院)
;
The University of Queensland(昆士兰大学)
;
Gaash Research Lab, National Institute of Technology Srinagar(加什研究实验室,锡纳加尔国家理工学院)
CommentsWe present a framework for evaluation of Multi-modal Agents consisting of Voice-to-voice model components viz. Text to Speech (TTS), Retrieval Augmented Generation (RAG) and Speech-to-text (STT)
Vision-Language Feature Alignment for Road Anomaly Segmentation
视觉-语言特征对齐用于道路异常分割
Zhuolin He, Jiacheng Tang, Jian Pu, Xiangyang Xue
机构
*
School of Computer Science, Fudan University(复旦大学计算机科学学院)
;
Institute of Science and Technology for Brain-Inspired Intelligence, Fudan University(复旦大学脑启发智能科学与技术研究院)
MLRecon: Robust Markerless Freehand 3D Ultrasound Reconstruction via Coarse-to-Fine Pose Estimation
MLRecon: 通过粗到细的姿态估计实现鲁棒的无标记自由手3D超声重建
Yi Zhang, Puxun Tu, Kun Wang, Yulin Yan, Tao Ying, Xiaojun Chen
机构
*
Institute of Biomedical Manufacturing and Life Quality Engineering, School of Mechanical Engineering, Shanghai Jiao Tong University, Shanghai, China(生物医学制造与生命质量工程学院,机械工程学院,上海交通大学,上海,中国)
;
Department of Ultrasound in Medicine, Shanghai Sixth People's Hospital Affiliated to Shanghai Jiao Tong University School of Medicine, Shanghai, China(医学超声科,上海第六人民医院(隶属于上海交通大学医学院))
;
Institute of Medical Robotics, Shanghai Jiao Tong University, Shanghai, China(医学机器人研究院,上海交通大学,上海,中国)
nchellwig at SemEval-2026 Task 3: Self-Consistent Structured Generation (SCSG) for Dimensional Aspect-Based Sentiment Analysis using Large Language Models
Wild-Drive: Off-Road Scene Captioning and Path Planning via Robust Multi-modal Routing and Efficient Large Language Model
Wild-Drive: 通过鲁棒多模态路由和高效大语言模型实现越野场景描述与路径规划
Zihang Wang, Xu Li, Benwu Wang, Wenkai Zhu, Xieyuanli Chen, Dong Kong, Kailin Lyu, Yinan Du, Yiming Peng, Haoyang Che
机构
*
School of Instrument Science and Engineering, Southeast University(东南大学仪器科学与工程学院)
;
Southeast University Nanjing Jiangbei New Area Innovation Research Institute(东南大学南京江滨新区创新研究院)
;
National Key Laboratory of Equipment State Sensing and Smart Support, National University of Defense Technology(国防科技大学装备状态感知与智能支撑国家重点实验室)
;
School of Transportation, Shandong University of Science and Technology(山东科技大学交通学院)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
专题命中
效率与部署
:large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI
What Do Visual Tokens Really Encode? Uncovering Sparsity and Redundancy in Multimodal Large Language Models
视觉标记究竟编码了什么?揭示多模态大语言模型中的稀疏性与冗余性
Yingqi Fan, Junlong Tong, Anhao Zhao, Xiaoyu Shen
机构
*
Institute of Digital Twin, Eastern Institute of Technology(数字孪生研究所,东部技术研究所)
;
Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生关键实验室)
;
Shanghai Jiao Tong University(上海交通大学)
;
The Hong Kong Polytechnic University(香港理工大学)
专题命中
效率与部署
:large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI