Speculative Decoding Reimagined for Multimodal Large Language Models
Luxi Lin, Zhihang Lin, Zhanpeng Zeng, Rongrong Ji
机构
*
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(多媒体可信感知与高效计算重点实验室、中华人民共和国教育部、厦门大学)
专题命中
VLM训练与架构
:multimodal large language model(title,abstract);LLaVA(abstract);MLLM(abstract);分类 cs.CV、cs.AI
Towards Interpreting Visual Information Processing in Vision-Language Models
Clement Neo, Luke Ong, Philip Torr, Mor Geva, David Krueger, Fazl Barez
机构
*
Nanyang Technological University(南洋理工大学)
;
University of Oxford(牛津大学)
;
Tel Aviv University(特拉维夫大学)
;
MILA(蒙特利尔人工智能研究院)
;
ERA-Krueger AI Safety Lab(ERA-Krueger人工智能安全实验室)
机构
*
Shanghai AI Laboratory(上海人工智能实验室)
;
Fudan University(复旦大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
South China University of Technology(华南理工大学)
;
Nanjing University(南京大学)
;
Xiamen University(厦门大学)
;
CUHK MMLab(香港中文大学MMLab)
专题命中
VLM训练与架构
:InternVL(title,abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
CommentsProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) Accepted to ACL 2026 (Oral presentation). Code available at https://github.com/mkimhi/CARES
A Hybrid Vision-Language Architecture for Automated Defect Reasoning and Report Generation in Industrial Inspection
一种用于工业检测中自动缺陷推理与报告生成的混合视觉-语言架构
Malikussaid, Imad Gohar
机构
*
School of Computing, Telkom University(Telkom大学计算机学院)
;
Faculty of Engineering and Technology, School of Computing and Artificial Intelligence(工程与技术学院,计算与人工智能学院)