Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning
裁决式字幕生成:用于严格零样本图像字幕生成的多智能体对齐评分与共识蒸馏束仲裁
Duy Tran Thanh, Thien-Phuc Doan, Long Nguyen-Vu, Ngo Tan Vu Khanh
机构
*
AI Platform OneNexus, OneMount(OneMount AI平台OneNexus)
;
School of Electronic Engineering, Soongsil University(崇实大学电子工程学院)
;
MoAdata(MoAdata公司)
;
University of Economics Ho Chi Minh City (UEH)(胡志明市经济大学)
Comments12 pages, 13 figures. Accepted at EXPLIMED 2026 (Third Workshop on Explainable Artificial Intelligence for the medical domain), IJCAI-ECAI 2026
Self-supervision drives representational convergence in medical foundation models more than clinical supervision
自我监督比临床监督更能推动医学基础模型中的表征趋同
Soroosh Tayebi Arasteh, Sebastian Ziegelmayer, Mahshad Lotfinia, Lisa Adams, Sven Nebelung, Jakob Nikolas Kather, Daniel Truhn
机构
*
RWTH Aachen University(亚琛工业大学)
;
University Hospital RWTH Aachen(亚琛工业大学附属医院)
;
Technical University of Munich(慕尼黑工业大学)
;
Friedrich-Alexander-Universität Erlangen-Nürnberg(埃尔朗根-纽伦堡弗里德里希-亚历山大大学)
;
Technical University Dresden(德累斯顿工业大学)
;
University Hospital Dresden(德累斯顿大学附属医院)
;
University Hospital Heidelberg(海德堡大学附属医院)
机构
*
School of Intelligence Science and Engineering, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)智能科学与工程学院)
;
School of Information Science and Engineering, Ningbo University(宁波大学信息科学与工程学院)
;
School of Information Science and Technology, University of Science and Technology of China(中国科学技术大学信息科学与技术学院)
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
XR-1:通过学习统一的视觉-运动表示实现多功能的视觉-语言-动作模型
Shichao Fan, Kun Wu, Zhengping Che, Xinhua Wang, Di Wu, Fei Liao, Ning Liu, Yixue Zhang, Zhen Zhao, Zhiyuan Xu, Meng Li, Qingjie Liu, Shanghang Zhang, Min Wan, Jian Tang
机构
*
Beijing Innovation Center of Humanoid Robotics, Beijing, China(北京人形机器人创新中心,北京,中国)
;
School of Mechanical Engineering and Automation, Beihang University, Beijing, China(北京航空航天大学机械工程及自动化学院,北京,中国)
;
State Key Laboratory of Virtual Reality Technology and Systems, SCSE, Beihang University, Beijing, China(虚拟现实技术与系统国家重点实验室,SCSE,北京航空航天大学,北京,中国)
;
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University, Beijing, China(多媒体信息处理国家重点实验室,计算机科学学院,北京大学,北京,中国)
MEDLAYXPLAIN: Benchmarking the Expert-Lay Gap in Medical Vision-Language Models
MEDLAYXPLAIN: 医学视觉语言模型中的专家-外行差距基准测试
Han Jang, Junhyeok Lee, Songsoo Kim, Chae Young Lim, Hyeonjin Goh, Heeseong Eum, Kyu Sung Choi
机构
*
Seoul National University(首尔大学)
;
Seoul National University Hospital(首尔大学医院)
;
Seoul National University College of Medicine(首尔大学医学院)
;
Sungkyunkwan University School of Medicine(成均馆大学医学院)
Vision-language models for chest radiography do not always need the image
胸部X光片的视觉-语言模型并不总是需要图像
Mahshad Lotfinia, Sebastian Ziegelmayer, Lisa Adams, Daniel Truhn, Andreas Maier, Soroosh Tayebi Arasteh
机构
*
Pattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg(弗里德里希-亚历山大-埃尔朗根-纽伦堡大学模式识别实验室)
;
Department of Diagnostic and Interventional Radiology, TUM University Clinic, School of Medicine and Health, Klinikum rechts der Isar, Technical University of Munich(慕尼黑工业大学医学院与健康学院伊萨尔河右岸医院诊断与介入放射学系)
;
Lab for AI in Medicine, RWTH Aachen University(亚琛工业大学医学人工智能实验室)
;
Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen(亚琛工业大学医院诊断与介入放射学系)
TEVI: Text-Conditioned Editing of Visual Representations via Sparse Autoencoders for Improved Vision-Language Alignment
TEVI: 基于稀疏自编码器的文本条件视觉表示编辑以改进视觉-语言对齐
Sweta Mahajan, Sukrut Rao, Jiahao Xie, Alexander Koller, Bernt Schiele
机构
*
Max Planck Institute for Informatics, Saarland Informatics Campus, Saarbrücken, Germany(马克斯·普朗克研究所信息学院,萨尔兰信息学院,德国萨尔布吕肯)
;
Department of Language Science and Technology, Saarland University, Saarbrücken, Germany(语言科学与技术系,萨尔兰大学,德国萨尔布吕肯)
Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics
视觉语言模型能预测未来状态吗?从逆动力学引导世界模型
Yifu Qiu, Yftah Ziser, Anna Korhonen, Shay B. Cohen, Edoardo M. Ponti
机构
*
Institute for Language, Cognition and Computation, University of Edinburgh(语言、认知与计算研究所,爱丁堡大学)
;
Language Technology Lab, University of Cambridge(语言技术实验室,剑桥大学)
;
NVIDIA(NVIDIA公司)
;
University of Groningen(格罗宁根大学)
Skip-It? Theoretical Conditions for Layer Skipping in Vision-Language Models
跳过层?视觉-语言模型中层跳过的理论条件
Max Hartman, Vidhata Jayaraman, Moulik Choraria, Akhil Bhimaraju, Lav R. Varshney
机构
*
Department of Electrical and Computer Engineering, The University of Illinois, Champaign, IL, United States(电气与计算机工程系,伊利诺伊大学香槟分校)
;
Department of Mathematics, The University of Illinois, Champaign, IL, United States(数学系,伊利诺伊大学香槟分校)
;
AI Innovation Institute, Stony Brook University, Stony Brook, NY, United States(人工智能创新研究所,石溪大学)
OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model
OMIBench:用于大型视觉-语言模型在奥林匹克级多图像推理中的基准测试
Qiguang Chen, Chengyu Luan, Jiajun Wu, Qiming Yu, Yi Yang, Yizhuo Li, Jingqi Tong, Xiachong Feng, Libo Qin, Wanxiang Che
机构
*
Research Center for Social Computing and Interactive Robotics(社会计算与交互机器人研究室)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
Central South University(中南大学)
;
Fudan University(复旦大学)
;
The University of Hong Kong(香港大学)
;
Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
;
Text Computing and Cognitive Intelligence Ministry of Education Engineering Research Center(教育部文本计算与认知智能工程研究中心)
;
Guizhou University(贵州大学)
CLASP: Closed-loop Asynchronous Spatial Perception for Open-vocabulary Desktop Object Grasping
CLASP: 闭环异步空间感知用于开放词汇桌面物体抓取
Yiran Ling, Wenxuan Li, Siying Dong, Yize Zhang, Xiaoyao Huang, Jing Jiang, Ruonan Li, Jie Liu
机构
*
Harbin Institute of Technology(哈尔滨工业大学)
;
National Key Laboratory of Smart Farm Technologies and Systems(智慧农场技术与系统全国重点实验室)
;
Peng Cheng Laboratory(鹏城实验室)
;
Northeastern University(东北大学)