arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6918 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6918 篇

2508.13995 2025-08-20 cs.CV 57%

Self-Supervised Sparse Sensor Fusion for Long Range Perception

Edoardo Palladin, Samuel Brucker, Filippo Ghilotti, Praveen Narayanan, Mario Bijelic, Felix Heide

机构 * Torc Robotics(Torc机器人公司) Princeton University(普林斯顿大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12609 2025-08-19 cs.CV 57%

Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation

Lexiang Tang, Xianwei Zhuang, Bang Yang, Zhiyuan Hu, Hongxiang Li, Lu Ma, Jinghan Ru, Yuexian Zou

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01212 2025-08-19 cs.CV cs.HC 57%

Towards Consumer-Grade Cybersickness Prediction: Multi-Model Alignment for Real-Time Vision-Only Inference

Yitong Zhu, Zhuowen Liang, Yiming Wu, Tangyao Li, Yuyang Wang

机构 * The Hong Kong University of Science and Technology(Guangzhou)(香港科学与技术大学(广州)) Nanyang Technological University(南洋理工大学) Nanyang Technological University Singapore(南洋理工大学新加坡)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11419 2025-08-18 cs.CV 57%

Training-free Dimensionality Reduction via Feature Truncation: Enhancing Efficiency in Privacy-preserving Multi-Biometric Systems

Florian Bayer, Maximilian Russo, Christian Rathgeb

机构 * Hochschule Darmstadt(达姆斯塔特应用技术大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10704 2025-08-15 cs.CV 57%

Beyond conventional vision: RGB-event fusion for robust object detection in dynamic traffic scenarios

Zhanwen Liu, Yujing Sun, Yang Wang, Nan Yang, Shengbo Eben Li, Xiangmo Zhao

机构 * School of Information Engineering, Chang’an University(信息工程学院,长安大学) School of Vehicle Mobility & College of AI, Tsinghua University(车辆运动学院与人工智能学院,清华大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10232 2025-08-15 cs.CV 57%

CellSymphony: Deciphering the molecular and phenotypic orchestration of cells with single-cell pathomics

Paul H. Acosta, Pingjun Chen, Simon P. Castillo, Maria Esther Salvatierra, Yinyin Yuan, Xiaoxi Pan

机构 * Translational Molecular Pathology Department, Division of Pathology and Laboratory Medicine, The University of Texas MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心转化分子病理学部门) Institute for Data Science in Oncology, The University of Texas MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心肿瘤数据科学研究所)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11049 2025-08-15 cs.LG cs.AI 57%

15,500 Seconds: Lean UAV Classification Using EfficientNet and Lightweight Fine-Tuning

Andrew P. Berg, Qian Zhang, Mia Y. Wang

机构 * Department of Computer Science(计算机科学系) College of Charleston(查尔斯顿学院) Department of Engineering(工程系)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09802 2025-08-14 cs.CV 57%

MUJICA: Reforming SISR Models for PBR Material Super-Resolution via Cross-Map Attention

Xin Du, Maoyuan Xu, Zhi Ying

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14332 2025-08-14 cs.CV 57%

ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs

Yin Xie, Kaicheng Yang, Peirou Liang, Xiang An, Yongle Zhao, Yumeng Wang, Ziyong Feng, Roy Miles, Ismail Elezi, Jiankang Deng

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 10 pages, 6 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09227 2025-08-14 cs.LG cs.AI cs.CE 57%

GSMT: Graph Fusion and Spatiotemporal TaskCorrection for Multi-Bus Trajectory Prediction

Fan Ding, Hwa Hui Tew, Junn Yong Loo, Susilawati, LiTong Liu, Fang Yu Leong, Xuewen Luo, Kar Keong Chin, Jia Jun Gan

机构 * School of Information Technology, Monash University Malaysia(墨尔本大学马来西亚分校信息科技学院) School of Engineering, Monash University Malaysia(墨尔本大学马来西亚分校工程学院) Perunding Atur Trafik Sdn Bhd, Malaysia(马来西亚交通自动化有限公司)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments This paper has been accepted by ITSC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08917 2025-08-13 cs.CV 57%

A Pseudo Global Fusion Paradigm-Based Cross-View Network for LiDAR-Based Place Recognition

Jintao Cheng, Jiehao Luo, Xieyuanli Chen, Jin Wu, Rui Fan, Xiaoyu Tang, Wei Zhang

机构 * School of Electronics and Information Engineering, and Xingzhi College, South China Normal University(电子信息工程学院和星智学院,华南师范大学) School of Data Science and Engineering, and Xingzhi College, South China Normal University(数据科学与工程学院和星智学院,华南师范大学) College of Intelligence Science and Technology, National University of Defense Technology(智能科学与技术学院,国防科技大学) College of Electronics & Information Engineering, Shanghai Research Institute for Intelligent Autonomous Systems, the State Key Laboratory of Intelligent Autonomous Systems, and Frontiers Science Center for Intelligent Autonomous Systems, Tongji University(电子与信息工程学院,上海智能自主系统研究所,智能自主系统国家重点实验室,智能自主系统前沿科学中心,同济大学) Department of Electronic and Computer Engineering, Hong Kong University of Science and Technology(电子与计算机工程系,香港科技大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08697 2025-08-13 cs.CV 57%

ROD: RGB-Only Fast and Efficient Off-road Freespace Detection

Tong Sun, Hongliang Ye, Jilin Mei, Liang Chen, Fangzhou Zhao, Leiqiang Zong, Yu Hu

机构 * Research Center for Intelligent Computing Systems, Institute of Computing Technology, Chinese Academy of Sciences(智能计算系统研究室,计算技术研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Astronomical Computing Research Center, Zhejiang Lab(天文计算研究室,浙江实验室) Beijing Special Vehicle Academy(北京特种车辆学院) School of Computer Science and Technology, University of Chinese Academy of Sciences(计算机科学与技术学院,中国科学院大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Journal ref ICRA2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07804 2025-08-12 cs.CV 57%

Pose-RFT: Enhancing MLLMs for 3D Pose Generation via Hybrid Action Reinforcement Fine-Tuning

Bao Li, Xiaomei Zhang, Miao Xu, Zhaoxin Fan, Xiangyu Zhu, Zhen Lei

机构 * CASIA(中国科学院自动化研究所) UCAS(中国科学院大学) CAIR, HKISI, CAS(中国科学院自动化研究所) Beihang University(北京航空航天大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07480 2025-08-12 eess.SP cs.AI cs.LG 57%

EEG-Language Pretraining for Highly Label-Efficient Clinical Phenotyping

Sam Gijsen, Kerstin Ritter

机构 * Charité – Universitätsmedizin Berlin, Department of Psychiatry and Psychotherapy, Berlin, Germany(柏林查理医院医学大学精神病与心理治疗系) Hertie Institute for AI in Brain Health, University of Tübingen, Germany(图宾根大学健康人工智能研究所)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments Accepted to ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03060 2025-08-07 cs.CV 57%

CHARM: Collaborative Harmonization across Arbitrary Modalities for Modality-agnostic Semantic Segmentation

Lekang Wen, Jing Xiao, Liang Liao, Jiajun Chen, Mi Wang

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06133 2025-08-07 cs.CV 57%

BrainSegDMlF: A Dynamic Fusion-enhanced SAM for Brain Lesion Segmentation

Hongming Wang, Yifeng Wu, Huimin Huang, Hongtao Wu, Jia-Xuan Jiang, Xiaodong Zhang, Hao Zheng, Xian Wu, Yefeng Zheng, Jinping Xu, Jing Cheng

机构 * Southern University of Science and Technology(南方科技大学) Jarvis Research Center, Tencent YouTu Lab(腾讯优图实验室) Westlake University(西湖大学) Shenzhen University of Advanced Technology(深圳先进技术大学) Tencent YouTu Lab(腾讯优图实验室) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究所,中国科学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10691 2025-08-07 cs.AI 57%

Fine-Tuning and Deploying Large Language Models Over Edges: Issues and Approaches

Yanjie Dong, Haijun Zhang, Chengming Li, Song Guo, Victor C. M. Leung, Xiping Hu

机构 * Shenzhen MSU-BIT University(深圳MSU-BIT大学) University of Science and Technology Beijing(北京科技大学) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03277 2025-08-06 cs.CV 57%

Efficient Multi-Slide Visual-Language Feature Fusion for Placental Disease Classification

Hang Guo, Qing Zhang, Zixuan Gao, Siyuan Yang, Shulin Peng, Xiang Tao, Ting Yu, Yan Wang, Qingli Li

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted by ACMMM'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03257 2025-08-05 cs.CV 57%

LACONIC: A 3D Layout Adapter for Controllable Image Creation

Léopold Maillard, Tom Durand, Adrien Ramanana Rahary, Maks Ovsjanikov

机构 * LIX, École Polytechnique, IP Paris(巴黎高等理工学院LIX研究所) Dassault Systèmes(达索系统)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00248 2025-08-04 cs.CV 57%

Guided Depth Map Super-Resolution via Multi-Scale Fusion U-shaped Mamba Network

Chenggang Guo, Hao Xu, XianMing Wan

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00053 2025-08-04 cs.CV 57%

A Quality-Guided Mixture of Score-Fusion Experts Framework for Human Recognition

Jie Zhu, Yiyang Su, Minchul Kim, Anil Jain, Xiaoming Liu

机构 * Department of Computer Science and Engineering, Michigan State University(计算机科学与工程系,密歇根州立大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted to ICCV 2025. 11 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20216 2025-08-01 cs.CV 57%

Dual-Stream Global-Local Feature Collaborative Representation Network for Scene Classification of Mining Area

Shuqi Fan, Haoyi Wang, Xianju Li

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Accepted to IJCNN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21857 2025-07-30 cs.CV 57%

Unleashing the Power of Motion and Depth: A Selective Fusion Strategy for RGB-D Video Salient Object Detection

Jiahao He, Daerji Suolang, Keren Fu, Qijun Zhao

机构 * College of Computer Science, Sichuan University(四川大学计算机学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments submitted to TMM on 11-Jun-2024, ID: MM-020522, still in peer review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20842 2025-07-29 cs.CV 57%

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models

Yuchen Liu, Yaoming Wang, Bowen Shi, Xiaopeng Zhang, Wenrui Dai, Chenglin Li, Hongkai Xiong, Qi Tian

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20025 2025-07-29 cs.CV 57%

Region-based Cluster Discrimination for Visual Representation Learning

Yin Xie, Kaicheng Yang, Xiang An, Kun Wu, Yongle Zhao, Weimo Deng, Zimin Ran, Yumeng Wang, Ziyong Feng, Roy Miles, Ismail Elezi, Jiankang Deng

机构 * DeepGlint University of Technology Sydney(悉尼科技大学) Huawei London Research Center(华为伦敦研究中心) Imperial College London(伦敦帝国理工学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted as a highlight paper at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22373 2025-07-29 cs.LG cs.AI 57%

Analytic Continual Test-Time Adaptation for Multi-Modality Corruption

Yufei Zhang, Yicheng Xu, Hongxin Wei, Zhiping Lin, Xiaofeng Zou, Cen Chen, Huiping Zhuang

机构 * South China University of Technology(华南理工大学) Institute of Science Tokyo(东京科学研究所) Southern University of Science and Technology(南方科技大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18522 2025-07-25 cs.CV 57%

GaussianFusionOcc: A Seamless Sensor Fusion Approach for 3D Occupancy Prediction Using 3D Gaussians

Tomislav Pavković, Mohammad-Ali Nikouei Mahani, Johannes Niedermayer, Johannes Betz

机构 * Technical University of Munich(慕尼黑技术大学) BMW Group(宝马集团) Munich Institute of Robotics and Machine Intelligence(慕尼黑机器人与智能机械研究所)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17456 2025-07-24 cs.CV 57%

Dynamic Scoring with Enhanced Semantics for Training-Free Human-Object Interaction Detection

Francesco Tonini, Lorenzo Vaquero, Alessandro Conti, Cigdem Beyan, Elisa Ricci

机构 * University of Trento(特伦托大学) University of Verona(威尼斯大学) Department of Computer Science(计算机科学系)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted to ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07939 2025-07-23 cs.CL 57%

SAGE: A Visual Language Model for Anomaly Detection via Fact Enhancement and Entropy-aware Alignment

Guoxin Zang, Xue Li, Donglin Di, Lanshun Nie, Dechen Zhan, Yang Song, Lei Fan

机构 * Harbin Institute of Technology(哈尔滨工业大学) University of New South Wales(新南威尔士大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

Comments Accepted by ACMMM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00721 2025-07-22 cs.CV 57%

UPRE: Zero-Shot Domain Adaptation for Object Detection via Unified Prompt and Representation Enhancement

Xiao Zhang, Fei Wei, Yong Wang, Wenda Zhao, Feiyi Li, Xiangxiang Chu

机构 * Dalian University of Technology(大连理工大学) AMAP, Alibaba Group(阿里集团AMAP)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏