CommentsPreprint. Submitted to EMNLP 2026. 21 pages, including appendices; 5 figures Under review. Yanqiu Zhao and Dongying Zheng contributed equally to this work
Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey
面向具身操作的高效视觉-语言-动作模型:系统综述
Weifan Guan, Qinghao Hu, Aosheng Li, Jian Cheng
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
AiRiA
;
Nanjing University of Information Science and Technology(南京信息科学技术大学)
What Makes LVLMs Hallucinate Less? Unveiling the Architectural Factors Behind Hallucination Robustness
什么使LVLMs更少产生幻觉?揭示影响幻觉鲁棒性的架构因素
Yusheng He, Jizhe Zhou, Xia Du, Zheng Lin, Jun Luo, Jiancheng Lv
机构
*
School of Computer Science, Engineering Research Center of Machine Learning and Industry Intelligence, Sichuan University(计算机科学学院,机器学习与产业智能工程研究中心,四川大学)
;
School of Computer and Information Engineering, Xiamen University of Technology(计算机与信息工程学院,厦门理工大学)
;
Department of Electrical and Computer Engineering, University of Hong Kong(电气与计算机工程系,香港大学)
;
College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学)
Cross-Modal Attention Calibration for LVLM Hallucination Mitigation
跨模态注意力校准用于LVLM幻觉缓解
Jiaming Li, Jiacheng Zhang, Zequn Jie, Lin Ma, Guanbin Li
机构
*
Sun Yat-sen University(中山大学)
;
The University of Hong Kong(香港大学)
;
Meituan(美团)
;
Inspur Database Technology(Inspur数据库技术)
;
Guilin University of Electronic Technology(桂林电子科技大学)
;
Shenzhen Loop Area Institute(深圳河套学院)
;
Guangdong Key Laboratory of Big Data Analysis and Processing(广东大数据分析与处理重点实验室)
A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models
策展人引导的多语言艺术描述对盲人和低视力观众的小型视觉语言模型试点研究
Iosif Tsangko, Andreas Triantafyllopoulos, George Margetis, Ioana Crihana, Björn W. Schuller
机构
*
Technical University of Munich(慕尼黑技术大学)
;
Foundation for Research and Technology -- Hellas(希腊研究与技术基金会)
;
National University of Science and Technology Politehnica Bucharest(布加勒斯特政治技术科学与技术国家大学)
Beyond Classification: Dynamic Adapter Routing for Continual Multimodal Retrieval
超越分类:面向持续多模态检索的动态适配器路由
Alicja Dobrzeniecka, Filip Szatkowski, Sebastian Cygert, Szymon Lukasik, Bartlomiej Twardowski
机构
*
NASK National Research Institute(NASK国家研究院)
;
IDEAS Research Institute(IDEAS研究所)
;
Warsaw University of Technology(华沙技术大学)
;
Universitat Autonoma de Barcelona(巴塞罗那自治大学)
Variational Adapter for Cross-modal Similarity Representation
变分适配器用于跨模态相似性表示
WenZhang Wei, Zhipeng Gui, Dehua Peng, Tiandi Ye, Huayi Wu
机构
*
School of Remote Sensing and Information Engineering(遥感与信息工程学院)
;
Wuhan University(武汉大学)
;
School of Data Science and Engineering(数据科学与工程学院)
;
East China Normal University(华东师范大学)
;
State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing(测绘遥感信息工程国家重点实验室)
IsoCLIP: Decomposing CLIP Projectors for Efficient Intra-modal Alignment
IsoCLIP: 分解CLIP投影器以实现高效的模态内对齐
Simone Magistri, Dipam Goswami, Marco Mistretta, Bartłomiej Twardowski, Joost van de Weijer, Andrew D. Bagdanov
机构
*
Media Integration and Communication Center (MICC), University of Florence, Italy(意大利佛罗伦萨大学媒体集成与通信中心)
;
Department of Computer Science, Universitat Autònoma de Barcelona, Spain(西班牙巴塞罗那自治大学计算机科学系)
;
Computer Vision Center, Barcelona, Spain(西班牙巴塞罗那计算机视觉中心)
;
IDEAS Research Institute, Warsaw, Poland(波兰华沙IDEAS研究所)
PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection
PRISM:免训练多模态数据选择的自剪枝内在选择方法
Jinhe Bi, Aniri, Zengjie Jin, Yifan Wang, Danqi Yan, Wenke Huang, Xiaowen Ma, Sikuan Yan, Artur Hecker, Mang Ye, Xun Xiao, Hinrich Schuetze, Volker Tresp, Yunpu Ma
机构
*
LMU Munich(慕尼黑大学)
;
Munich Research Center, Huawei Technologies(慕尼黑研究中心,华为技术)
;
METEOR
;
School of Computer Science, Wuhan University(武汉大学计算机学院)
;
Munich Center for Machine Learning(慕尼黑机器学习中心)
专题命中
VLM训练与架构
:multimodal large language model(abstract);分类 cs.CV、cs.AI
CommentsAccepted to ACL 2026 and selected for the Best Paper list; later desk-rejected due to an inadvertent manual bibliography-editing error. Previous versions are withdrawn due to an inadvertent manual bibliography-editing error; please refer to the latest corrected version