LightEMMA: Lightweight End-to-End Multimodal Model for Autonomous Driving
Zhijie Qiao, Haowei Li, Zhong Cao, Henry X. Liu
机构
*
Department of Civil and Environmental Engineering, University of Michigan(土木与环境工程系,密歇根大学)
;
University of Michigan Transportation Research Institute(密歇根大学交通研究所)
Fighting Fire with Fire (F3): A Training-free and Efficient Visual Adversarial Example Purification Method in LVLMs
Yudong Zhang, Ruobing Xie, Yiqing Huang, Jiansheng Chen, Xingwu Sun, Zhanhui Kang, Di Wang, Yu Wang
机构
*
Tsinghua University, Tencent(清华大学,腾讯)
;
Tencent(腾讯)
;
University of Science and Technology Beijing(北京科技大学)
;
Tencent, University of Macau(腾讯,澳门大学)
;
Tsinghua University(清华大学)
机构
*
School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
;
Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University)(东南大学新一代人工智能技术及其交叉应用关键实验室)
;
Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
;
Wuhan University of Technology(武汉理工大学)
;
National University of Singapore(新加坡国立大学)
Integrating Prior Observations for Incremental 3D Scene Graph Prediction
Marian Renz, Felix Igelbrink, Martin Atzmueller
机构
*
DFKI Niedersachsen(德克萨斯联合研究所(北莱茵威斯特法伦))
;
Cooperative and Autonomous Systems, DFKI Niedersachsen(合作与自主系统,DFKI北莱茵威斯特法伦)
;
German Research Center for Artificial Intelligence(德国人工智能研究中心)
;
Semantic Information Systems, Osnabrück University(语义信息系统,奥斯纳布吕克大学)
专题命中
图文多模态
:multi-modal(abstract);分类 cs.CV、cs.AI
CommentsAccepted at 24th International Conference on Machine Learning and Applications (ICMLA'25)
机构
*
Carnegie Mellon University(卡内基梅隆大学)
;
University of Nottingham(诺丁汉大学)
;
The University of Hong Kong(香港大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
Columbia University(哥伦比亚大学)
;
University of California, Berkeley(加州大学伯克利分校)
专题命中
图文多模态
:multimodal(abstract);分类 cs.CL
Comments7pages, accepted by ICML TTODLer-FM workshop
Alexandre Sallinen, Stefan Krsteski, Paul Teiletche, Marc-Antoine Allard, Baptiste Lecoeur, Michael Zhang, Fabrice Nemo, David Kalajdzic, Matthias Meyer, Mary-Anne Hartley
机构
*
École Polytechnique Fédérale de Lausanne (EPFL), Switzerland(瑞士联邦理工学院洛桑校区)
;
ETH Zürich, Switzerland(瑞士苏黎世联邦理工学院)
;
T.H. Chan School of Public Health, Harvard University, USA(哈佛大学T.H. Chan公共卫生学院)
专题命中
音频语音多模态
:multimodal(title,abstract);分类 cs.AI
CommentsThis paper was originally submitted to the CODEML workshop for ICML 2025. 9 pages (including references and appendices)
Character-Centric Understanding of Animated Movies
Zhongrui Gui, Junyu Xie, Tengda Han, Weidi Xie, Andrew Zisserman
机构
*
Visual Geometry Group Dept.\ of Engineering Science University of Oxford, UK
;
School of Artifitial Intelligence\ Jiao Tong University Shanghai China
;
School of Artifitial Intelligence\ Jiao Tong University
机构
*
School of Economics and Finance, Shanghai International Studies University(经济金融学院,上海国际问题研究大学)
;
School of AI and Advanced Computing, Xi’an Jiaotong-Liverpool University(人工智能与先进计算学院,西安交通大学利物浦大学)
;
School of Foreign Studies, Shanghai University of Finance and Economics(外国语言学院,上海金融学院)
机构
*
School of Automation, Northwestern Polytechnical University(自动化学院,西北工业大学)
;
The University of Hong Kong(香港大学)
;
School of Software, Northwestern Polytechnical University(软件学院,西北工业大学)
;
Unmanned System Research Institute, Northwestern Polytechnical University(无人系统研究院,西北工业大学)
CogGNN: Cognitive Graph Neural Networks in Generative Connectomics
Mayssa Soussia, Yijun Lin, Mohamed Ali Mahjoub, Islem Rekik
机构
*
National Engineering School of Sousse, University of Sousse, LATIS- Laboratory of Advanced Technology and Intelligent Systems(突尼斯苏塞国立工程学校,苏塞大学,先进技术与智能系统实验室)
;
BASIRA Lab, Imperial-X(BASIRA实验室,Imperial-X)
;
Department of Computing, Imperial College London, UK(计算系,伦敦帝国学院,英国)
Trial-Level Time-frequency EEG Desynchronization as a Neural Marker of Pain
D. A. Blanco-Mora, A. Dierolf, J. Gonçalves, M. van Der Meulen
机构
*
Luxembourg Centre for Systems Biomedicine, University of Luxembourg(卢森堡系统生物医学研究中心,卢森堡大学)
;
Department of Behavioural and Cognitive Sciences, University of Luxembourg(行为与认知科学系,卢森堡大学)
机构
*
IEEE Publication Technology Department(IEEE出版技术部门)
;
State Key Laboratory of Integrated Services Networks, School of Telecommunications Engineering, Xidian University(信息服务网络国家重点实验室,电信工程学院,西安电子科技大学)
;
Department of Computer Science and Technology, Tongji University(计算机科学与技术系,同济大学)
What is the Visual Cognition Gap between Humans and Multimodal LLMs?
Xu Cao, Yifan Shen, Bolin Lai, Wenqian Ye, Yunsheng Ma, Joerg Heintz, Jintai Chen, Meihuan Huang, Jianguo Cao, Aidong Zhang, James M. Rehg
机构
*
Department of Computer Science, University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机科学系)
;
College of Computing, Georgia Institute of Technology(佐治亚理工学院计算机学院)
;
Department of Computer Science, University of Virginia(弗吉尼亚大学计算机科学系)
;
Digital Twin Lab, Purdue University(普渡大学数字孪生实验室)
;
HKUST (Guangzhou)(香港科技大学(广州))
;
Department of Rehabilitation Medicine, Shenzhen Children’s Hospital(深圳儿童医院康复医学系)
FairCoT: Enhancing Fairness in Text-to-Image Generation via Chain of Thought Reasoning with Multimodal Large Language Models
Zahraa Al Sahili, Ioannis Patras, Matthew Purver
机构
*
School of Electronic Engineering and Computer Science, Queen Mary University of London(伦敦女王学院电子工程与计算机科学学院)
;
Department of Knowledge Technologies, Jožef Stefan Institute(Jožef Stefan研究所知识技术系)