arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6903 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6903 篇

2504.09588 2025-08-22 cs.CV cs.AI 62%

TextSplat: Text-Guided Semantic Fusion for Generalizable Gaussian Splatting

Zhicong Wu, Hongbin Xu, Gang Xu, Ping Nie, Zhixin Yan, Jinkai Zheng, Liangqiong Qu, Ming Li, Liqiang Nie

机构 * Institute of Artificial Intelligence, Xiamen University(厦门大学人工智能研究所) Guangdong Laboratory of Artificial Intelligence(广东省人工智能与数字经济实验室) Peking University(北京大学) School of Future Technology, South China University of Technology(华南理工大学未来技术学院) Hangzhou Dianzi University(杭州电子科技大学) Central Laboratory of Lishui Hospital of Wenzhou Medical University, The First Affiliated Hospital of Lishui University, Lishui People's Hospital(丽水市人民医院中央实验室、丽水大学第一附属医院、丽水人民医院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.07796 2025-08-20 cs.CV cs.AI cs.LG 62%

Fusing Echocardiography Images and Medical Records for Continuous Patient Stratification

Nathan Painchaud, Jérémie Stym-Popper, Pierre-Yves Courand, Nicolas Thome, Pierre-Marc Jodoin, Nicolas Duchateau, Olivier Bernard

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 13 pages + 2 pages of supplementary material, accepted for publication in IEEE TUFFC

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12036 2025-08-19 cs.CV cs.AI 62%

Q-FSRU: Quantum-Augmented Frequency-Spectral Fusion for Medical Visual Question Answering

Rakesh Thakur, Yusra Tariq

机构 * Amity Centre for Artificial Intelligence, Amity University, Noida(阿米蒂人工智能中心,阿米蒂大学,诺伊达)

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV、cs.AI

Comments 8 pages, 4 figures Submitted to AAAI 26

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06163 2025-08-11 cs.CL cs.AI cs.LG 62%

One Size Does Not Fit All: A Distribution-Aware Sparsification for More Precise Model Merging

Yingfeng Luo, Dingyang Lin, Junxin Wang, Ziqiang Xu, Kaiyan Chang, Tong Zheng, Bei Li, Anxiang Ma, Tong Xiao, Zhengtao Yu, Jingbo Zhu

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03930 2025-08-08 cs.CL cs.AI 62%

GuARD: Effective Anomaly Detection through a Text-Rich and Graph-Informed Language Model

Yunhe Pang, Bo Chen, Fanjin Zhang, Yanghui Rao, Evgeny Kharlamov, Jie Tang

机构 * School of Computer Science and Engineering(计算机科学与工程学院) Sun Yat-Sen University(中山大学) Department of Computer Science and Technology(计算机科学与技术系) Tsinghua University(清华大学) Robert Bosch GmbH(博世集团)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Accepted at KDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01396 2025-08-07 cs.CV cs.AI 62%

Spatial-Frequency Aware for Object Detection in RAW Image

Zhuohua Ye, Liming Zhang, Hongru Han

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.02562 2025-08-06 cs.CV cs.AI cs.CY cs.LG 62%

Beyond Images: Adaptive Fusion of Visual and Textual Data for Food Classification

Prateek Mittal, Puneet Goyal, Joohi Chauhan

机构 * Motilal Nehru National Institute of Technology Allahabad, India(Motilal Nehru国立技术学院Allahabad分校) Indian Institute of Technology Ropar, India(印度理工学院Ropar分校)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00230 2025-08-04 cs.LG cs.CL cs.CV 62%

Towards Higher Effective Rank in Parameter-efficient Fine-tuning using Khatri--Rao Product

Paul Albert, Frederic Z. Zhang, Hemanth Saratchandran, Anton van den Hengel, Ehsan Abbasnejad

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL

Comments To appear in ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19119 2025-08-01 cs.CV cs.AI 62%

PatchTraj: Unified Time-Frequency Representation Learning via Dynamic Patches for Trajectory Prediction

Yanghong Liu, Xingping Dong, Ming Li, Weixing Zhang, Yidong Lou

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04151 2025-07-25 cs.CV cs.AI cs.LG 62%

Robust Multi-View Learning via Representation Fusion of Sample-Level Attention and Alignment of Simulated Perturbation

Jie Xu, Na Zhao, Gang Niu, Masashi Sugiyama, Xiaofeng Zhu

机构 * University of Electronic Science and Technology of China(电子科技大学) Hainan University(海南大学) Singapore University of Technology and Design(新加坡科技设计大学) Southeast University(东南大学) The University of Tokyo(东京大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06106 2025-07-25 cs.CV cs.AI 62%

Towards a Universal 3D Medical Multi-modality Generalization via Learning Personalized Invariant Representation

Zhaorui Tan, Xi Yang, Tan Pan, Tianyi Liu, Chen Jiang, Xin Guo, Qiufeng Wang, Anh Nguyen, Yuan Qi, Kaizhu Huang, Yuan Cheng

机构 * Shanghai Academy of Artificial Intelligence for Science(上海人工智能科学研究院) Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学) University of Liverpool(利物浦大学) AI 3 Fudan University(复旦大学) Duke Kunshan University(杜克-昆山大学) Zhongshan Hospital(中山医院)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV25

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09146 2025-07-23 cs.CV cs.AI 62%

Fusion-Mamba for Cross-modality Object Detection

Wenhao Dong, Haodong Zhu, Shaohui Lin, Xiaoyan Luo, Yunhang Shen, Xuhui Liu, Juan Zhang, Guodong Guo, Baochang Zhang

机构 * Beihang University(北航大学) East China Normal University(东华师范大学) Tencent Youtu Lab(腾讯优图实验室) Eastern Institute of Technology(东部技术研究所)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

Journal ref IEEE Transactions on Multimedia, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12774 2025-07-18 cs.LG cs.AI cs.CL 62%

A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models

Weijieying Ren, Jingxi Zhu, Zehao Liu, Tianxiang Zhao, Vasant Honavar

机构 * Information Sciences and Technology, The Pennsylvania State University(信息科学与技术系,宾夕法尼亚州立大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10015 2025-07-18 cs.CV cs.AI cs.LG 62%

(Almost) Free Modality Stitching of Foundation Models

Jaisidh Singh, Diganta Misra, Boris Knyazev, Antonio Orvieto

机构 * University of Tübingen(图宾根大学) Zuse School ELIZA(Zuse学校ELIZA) ELLIS Institute Tübingen(图宾根ELLIS研究所) MPI-IS Tübingen(图宾根MPI-IS研究所) SAIT AI Lab Montréal(蒙特利尔SAIT人工智能实验室) Tübingen AI Center(图宾根人工智能中心)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05221 2025-07-09 cs.CV cs.AI 62%

CTA: Cross-Task Alignment for Better Test Time Training

Samuel Barbeau, Pedram Fekri, David Osowiechi, Ali Bahri, Moslem Yazdanpanah, Masih Aminbeidokhti, Christian Desrosiers

机构 * ÉTS Montréal(蒙特利尔ÉTS) Concordia University(康科迪亚大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Preprint, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03893 2025-07-08 cs.CV cs.AI 62%

Hierarchical Semantic-Visual Fusion of Visible and Near-infrared Images for Long-range Haze Removal

Yi Li, Xiaoxiong Wang, Jiawei Wang, Yi Chang, Kai Cao, Luxin Yan

机构 * National Key Laboratory of Multispectral Information Intelligent Processing Technology, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(多光谱信息智能处理技术国家实验室,人工智能与自动化学院,华中科技大学) State Key Laboratory of Dynamic Optical Imaging and Measurement(动态光学成像与测量国家重点实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments This work has been accepted by IEEE Transactions on Multimedia for publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10967 2025-07-02 cs.CV cs.AI 62%

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Qizhe Zhang, Mengzhen Liu, Lichen Li, Ming Lu, Yuan Zhang, Junwen Pan, Qi She, Shanghang Zhang

机构 * National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(国家多媒体信息处理重点实验室,计算机学院,北京大学) ByteDance(字节跳动)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 22 pages, 5 figures, code: https://github.com/Theia-4869/CDPruner, project page: https://theia-4869.github.io/CDPruner

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12884 2025-07-01 cs.LG cs.AI cs.CV 62%

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks

Yuanze Hu, Zhaoxin Fan, Xinyu Wang, Gen Li, Ye Qiu, Zhichao Yang, Wenjun Wu, Kejian Wu, Yifan Sun, Xiaotie Deng, Jin Dong

机构 * Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing(北京未来区块链与隐私计算先进创新中心) Beihang University(北京航空航天大学) Hangzhou International Innovation Institute(杭州国际创新研究院) Xreal Renmin University(中国人民大学) Peking University(北京大学) Beijing Academy of Blockchain and Edge Computing (BABEC)(北京区块链与边缘计算研究院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21873 2025-06-30 cs.CV cs.AI 62%

Grounding-Aware Token Pruning: Recovering from Drastic Performance Drops in Visual Grounding Caused by Pruning

Tzu-Chun Chien, Chieh-Kai Lin, Shiang-Feng Tsai, Ruei-Chi Lai, Hung-Jen Chen, Min Sun

机构 * National Tsing Hua University (NTHU)(国立清华大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21294 2025-06-27 cs.CL cs.AI 62%

Detecting Referring Expressions in Visually Grounded Dialogue with Autoregressive Language Models

Bram Willemsen, Gabriel Skantze

机构 * Division of Speech, Music and Hearing(语音、音乐和听觉系) KTH Royal Institute of Technology(皇家理工学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Accepted for publication at XLLM @ ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.00863 2025-06-24 cs.CL cs.AI cs.LG 62%

UniMoT: Unified Molecule-Text Language Model with Discrete Token Representation

Shuhan Guo, Yatao Bian, Ruibing Wang, Nan Yin, Zhen Wang, Quanming Yao

机构 * Tsinghua University(清华大学) Tencent AI Lab(腾讯AI实验室) Northwestern Polytechnical University(西北工业大学) Hong Kong University of Science and Technology(香港科技大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments IJCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01287 2025-06-23 cs.CL cs.CV 62%

Deep Learning based Visually Rich Document Content Understanding: A Survey

Yihao Ding, Soyeon Caren Han, Jean Lee, Eduard Hovy

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Work in Progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14786 2025-06-19 cs.LG cs.AI cs.CV 62%

PIPE: Physics-Informed Position Encoding for Alignment of Satellite Images and Time Series

Haobo Li, Eunseo Jung, Zixin Chen, Zhaowei Wang, Yueya Wang, Huamin Qu, Alexis Kai Hon Lau

机构 * Department of Computer Science & Engineering(香港理工大学计算机科学与工程系) Division of Environment & Sustainability(香港理工大学环境与可持续发展学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19048 2025-06-12 cs.CV cs.AI 62%

BiCo-Fusion: Bidirectional Complementary LiDAR-Camera Fusion for Semantic- and Spatial-Aware 3D Object Detection

Yang Song, Lin Wang

机构 * AI Thrust, The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)人工智能方向) School of Electrical and Electronic Engineering (EEE), Nanyang Technological University (NTU)(南洋理工大学电子与电气工程学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by IEEE Robotics and Automation Letters (RA-L)

Journal ref IEEE Robotics and Automation Letters, Volume 10 Issue 2, 1457 - 1464, February 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04756 2025-06-06 cs.AI cs.CV eess.IV 62%

Ontology-based knowledge representation for bone disease diagnosis: a foundation for safe and sustainable medical artificial intelligence systems

Loan Dao, Ngoc Quoc Ly

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04706 2025-06-06 cs.CV cs.AI 62%

Line of Sight: On Linear Representations in VLLMs

Achyuta Rajaram, Sarah Schwettmann, Jacob Andreas, Arthur Conmy

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 8 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21556 2025-05-29 cs.CV cs.AI 62%

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts

Hee-Seon Kim, Minbeom Kim, Wonjun Lee, Kihyun Kim, Changick Kim

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments LVLM, Jailbreak

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17002 2025-05-29 cs.CV cs.AI 62%

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association

Abdul Hannan, Muhammad Arslan Manzoor, Shah Nawaz, Muhammad Irzam Liaqat, Markus Schedl, Mubashir Noman

机构 * University of TrentoItaly(特伦托大学) Mohamed bin Zayed University of Artificial IntelligenceU.A.E.(穆罕默德·本·扎耶德人工智能大学) Johannes Kepler UniversityAustria(约翰·凯撒大学) IMT LuccaItaly(利卡大学) Linz Institute of Technology, AI LabAustria(林茨技术研究所)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted at InterSpeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17161 2025-05-27 cs.AI cs.CV cs.LG 62%

SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Tianzhe Chu, Yuexiang Zhai, Jihan Yang, Shengbang Tong, Saining Xie, Dale Schuurmans, Quoc V. Le, Sergey Levine, Yi Ma

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Website at https://tianzhechu.com/SFTvsRL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19611 2025-05-27 cs.CV cs.AI 62%

Align and Surpass Human Camouflaged Perception: Visual Refocus Reinforcement Fine-Tuning

Ruolin Shen, Xiaozhong Ji, Kai WU, Jiangning Zhang, Yijun He, HaiHua Yang, Xiaobin Hu, Xiaoyu Sun

机构 * Technische Universität München(慕尼黑技术大学) ByteDance(字节跳动) Zhejiang University(浙江大学) Australian National University(澳大利亚国立大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Project Website: \url{https://github.com/HUuxiaobin/VRRF}

详情

展开后加载摘要…

URL PDF HTML 收藏