arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6918 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6918 篇

2505.09422 2025-05-15 cs.CV 57%

MoRAL: Motion-aware Multi-Frame 4D Radar and LiDAR Fusion for Robust 3D Object Detection

Xiangyuan Peng, Yu Wang, Miao Tang, Bierzynski Kay, Lorenzo Servadei, Robert Wille

机构 * Technical University of Munich(慕尼黑技术大学) Infineon Technologies AG(英飞凌科技AG) China University of Geosciences(中国地质大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07689 2025-05-13 cs.CV 57%

Anatomical Attention Alignment representation for Radiology Report Generation

Quang Vinh Nguyen, Minh Duc Nguyen, Thanh Hoang Son Vo, Hyung-Jeong Yang, Soo-Hyung Kim

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07398 2025-05-13 cs.CV 57%

DepthFusion: Depth-Aware Hybrid Feature Fusion for LiDAR-Camera 3D Object Detection

Mingqian Ji, Jian Yang, Shanshan Zhang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07172 2025-05-13 cs.CV 57%

Critique Before Thinking: Mitigating Hallucination through Rationale-Augmented Instruction Tuning

Zexian Yang, Dian Li, Dayan Wu, Gang Liu, Weiping Wang

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) Foundation Technology Center, Tencent PCG(腾讯PCG基础技术中心)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07110 2025-05-13 cs.HC cs.CV 57%

DeepSORT-Driven Visual Tracking Approach for Gesture Recognition in Interactive Systems

Tong Zhang, Fenghua Shao, Runsheng Zhang, Yifan Zhuang, Liuqingqing Yang

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06903 2025-05-13 cs.CV 57%

CheXLearner: Text-Guided Fine-Grained Representation Learning for Progression Detection

Yuanzhuo Wang, Junwen Duan, Xinyu Li, Jianxin Wang

机构 * School of Computer Science and Engineering, Central South University(计算机科学与工程学院,中南大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06151 2025-05-12 cs.CL 57%

Estimating Quality in Therapeutic Conversations: A Multi-Dimensional Natural Language Processing Framework

Alice Rueda, Argyrios Perivolaris, Niloy Roy, Dylan Weston, Sarmed Shaya, Zachary Cote, Martin Ivanov, Bazen G. Teferra, Yuqi Wu, Sirisha Rambhatla, Divya Sharma, Andrew Greenshaw, Rakesh Jetly, Yanbo Zhang, Bo Cao, Reza Samavi, Sridhar Krishnan, Venkat Bhat

机构 * St. Michael’s Hospital, Unity Health Toronto(圣米歇尔医院,统一健康多伦多) Toronto Metropolitan University(多伦多 Metropolitan 大学) University of Waterloo(滑铁卢大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

Comments 12 pages, 4 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04965 2025-05-09 cs.CV 57%

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding

Henry Zheng, Hao Shi, Qihang Peng, Yong Xien Chng, Rui Huang, Yepeng Weng, Zhongchao Shi, Gao Huang

机构 * Department of Automation, BNRist, Tsinghua University(自动化系、BNRist、清华大学) AI Lab, Lenovo Research(联想研究院人工智能实验室)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04758 2025-05-09 cs.CV 57%

Lightweight RGB-D Salient Object Detection from a Speed-Accuracy Tradeoff Perspective

Songsong Duan, Xi Yang, Nannan Wang, Xinbo Gao

机构 * State Key Laboratory of Integrated Services Networks, School of Telecommunications Engineering, Xidian University(西安电子科技大学信息与通信工程学院集成服务网络国家重点实验室)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Accepted by TIP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04616 2025-05-08 cs.CV 57%

Person Recognition at Altitude and Range: Fusion of Face, Body Shape and Gait

Feng Liu, Nicholas Chimitt, Lanqing Guo, Jitesh Jain, Aditya Kane, Minchul Kim, Wes Robbins, Yiyang Su, Dingqiang Ye, Xingguang Zhang, Jie Zhu, Siddharth Satyakam, Christopher Perry, Stanley H. Chan, Arun Ross, Humphrey Shi, Zhangyang Wang, Anil Jain, Xiaoming Liu

机构 * Department of Computer Science and Engineering, Michigan State University(计算机科学与工程系,密歇根州立大学) School of Electrical and Computer Engineering, Purdue University(电气与计算机工程学院,普渡大学) School of Interactive Computing, Georgia Tech(交互计算学院,佐治亚理工学院) Department of Electrical and Computer Engineering, University of Texas at Austin(电气与计算机工程系,德克萨斯大学奥斯汀分校)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 18 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04003 2025-05-08 eess.IV cs.CV 57%

Prototype-Based Information Compensation Network for Multi-Source Remote Sensing Data Classification

Feng Gao, Sheng Liu, Chuanzheng Gong, Xiaowei Zhou, Jiayi Wang, Junyu Dong, Qian Du

机构 * State Key Laboratory of Physical Oceanography, Ocean University of China(海洋大学物理海洋国家重点实验室) Department of Electrical and Computer Engineering, Mississippi State University(密苏里州立大学电气与计算机工程系)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments Accepted by IEEE TGRS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03798 2025-05-08 cs.LG cs.AI 57%

Position: Foundation Models Need Digital Twin Representations

Yiqing Shen, Hao Ding, Lalithkumar Seenivasan, Tianmin Shu, Mathias Unberath

机构 * Department of Computer Science(计算机科学系) Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02385 2025-05-06 eess.IV cs.CV 57%

An Arbitrary-Modal Fusion Network for Volumetric Cranial Nerves Tract Segmentation

Lei Xie, Huajun Zhou, Junxiong Huang, Jiahao Huang, Qingrun Zeng, Jianzhong He, Jiawei Zhang, Baohua Fan, Mingchu Li, Guoqiang Xie, Hao Chen, Yuanjing Feng

机构 * College of Information Engineering, Zhejiang University of Technology(浙江工业大学信息工程学院) Department of Computer Science and Engineering, Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系) Department of Computer Science and Engineering, Department of Chemical and Biological Engineering and Center for Aging Science , Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系) Department of Radiology, Xiangya Hospital, Central South University(中南大学湘雅医院放射科) Department of Neurosurgery, Nuclear Industry 215 Hospital of Shaanxi Province(陕西核工业215医院神经外科) Department of Neurosurgery, Taihe Hospital of Wannan Medical College(皖南医学院太和医院神经外科)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06222 2025-05-06 cs.CV 57%

Vision-based 3D Semantic Scene Completion via Capture Dynamic Representations

Meng Wang, Fan Wu, Yunchuan Qin, Ruihui Li, Zhuo Tang, Kenli Li

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00627 2025-05-02 cs.CV 57%

Brain Foundation Models with Hypergraph Dynamic Adapter for Brain Disease Analysis

Zhongying Deng, Haoyu Wang, Ziyan Huang, Lipei Zhang, Angelica I. Aviles-Rivero, Chaoyu Liu, Junjun He, Zoe Kourtzi, Carola-Bibiane Schönlieb

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 35 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20996 2025-04-30 cs.CV 57%

X-Fusion: Introducing New Modality to Frozen Large Language Models

Sicheng Mo, Thao Nguyen, Xun Huang, Siddharth Srinivasan Iyer, Yijun Li, Yuchen Liu, Abhishek Tandon, Eli Shechtman, Krishna Kumar Singh, Yong Jae Lee, Bolei Zhou, Yuheng Li

机构 * University of California, Los Angeles(加州大学洛杉矶分校) University of Wisconsin–Madison(威斯康星大学麦迪逊分校) Adobe Research(Adobe研究院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Project Page: https://sichengmo.github.io/XFusion/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18631 2025-04-29 cs.AI cs.LG 57%

Research on Personalized Medical Intervention Strategy Generation System based on Group Relative Policy Optimization and Time-Series Data Fusion

Dingxin Lu, Shurui Wu, Xinyi Huang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18096 2025-04-28 cs.AI cs.LG 57%

Combating the Bucket Effect:Multi-Knowledge Alignment for Medication Recommendation

Xiang Li, Haixu Ma, Guanyong Wu, Shi Mu, Chen Li, Shunpan Liang

机构 * School of Information Science and Engineering, Yanshan University(燕山大学信息科学与工程学院) Xinjiang College Of Science & Technology(新疆科学与技术学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.AI

Comments 18 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06018 2025-04-28 cs.LG cs.AI 57%

A Picture is Worth A Thousand Numbers: Enabling LLMs Reason about Time Series via Visualization

Haoxin Liu, Chenghao Liu, B. Aditya Prakash

机构 * Georgia Institute of Technology(佐治亚理工学院) Salesforce Research Asia(Salesforce亚洲研究院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18025 2025-04-28 cs.CV 57%

ShapeSpeak: Body Shape-Aware Textual Alignment for Visible-Infrared Person Re-Identification

Shuanglin Yan, Neng Dong, Shuang Li, Rui Yan, Hao Tang, Jing Qin

机构 * Nanjing University of Science and Technology(南京理工大学) Chongqing University of Posts and Telecommunications(重庆邮电大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01163 2025-04-25 cs.CV 57%

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer

Jiajun Deng, Tianyu He, Li Jiang, Tianyu Wang, Feras Dayoub, Ian Reid

机构 * Australian Institute for Machine Learning, The University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学) Microsoft Research(微软研究院) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Mohamed bin Zayed University of AI(穆罕默德·本·扎耶德人工智能大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.15473 2025-04-24 cs.CL cs.LG 57%

Beyond Self Attention: A Subquadratic Fourier Wavelet Transformer with Multi Modal Fusion

Andrew Kiruluta, Andreas Lemos, Eric Lundy

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CL

Comments 7 pages, 4 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04797 2025-04-23 cs.CL cs.LG 57%

Parallel Corpora for Machine Translation in Low-resource Indic Languages: A Comprehensive Review

Rahul Raja, Arpita Vats

机构 * Carnegie Mellon University(卡内基梅隆大学) Stanford University(斯坦福大学) Boston University(波士顿大学) Santa Clara University(圣克拉拉大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

Comments Accepted in NACCL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14231 2025-04-22 cs.CV 57%

Exploring Modality Guidance to Enhance VFM-based Feature Fusion for UDA in 3D Semantic Segmentation

Johannes Spoecklberger, Wei Lin, Pedro Hermosilla, Sivan Doveh, Horst Possegger, M. Jehanzeb Mirza

机构 * Institute of Visual Computing, Graz University of Technology(视觉计算研究所,格拉茨技术大学) JKU LINZ(JKU林茨) TU Wien(维也纳技术大学) IBM Research(IBM研究院) MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12966 2025-04-18 cs.CV cs.LG 57%

Vision and Language Integration for Domain Generalization

Yanmei Wang, Xiyao Liu, Fupeng Chu, Zhi Han

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00246 2025-04-15 cs.CV cs.LG 57%

ResiDual Transformer Alignment with Spectral Decomposition

Lorenzo Basile, Valentino Maiorca, Luca Bortolussi, Emanuele Rodolà, Francesco Locatello

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Published in Transactions on Machine Learning Research (TMLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08912 2025-04-15 cs.LG cs.AI 57%

HyperCore: The Core Framework for Building Hyperbolic Foundation Models with Comprehensive Modules

Neil He, Menglin Yang, Rex Ying

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08676 2025-04-15 cs.CV 57%

Language-Depth Navigated Thermal and Visible Image Fusion

Jinchang Zhang, Zijun Li, Guoyu Lu

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07441 2025-04-11 cs.CV 57%

WS-DETR: Robust Water Surface Object Detection through Vision-Radar Fusion with Detection Transformer

Huilin Yin, Pengyu Wang, Senmao Li, Jun Yan, Daniel Watzenig

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07061 2025-04-10 cs.CV 57%

Teaching pathology foundation models to accurately predict gene expression with parameter efficient knowledge transfer

Shi Pan, Jianan Chen, Maria Secrier

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏