arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6918 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6918 篇

2411.19187 2025-02-20 cs.CL 57%

Beyond Logit Lens: Contextual Embeddings for Robust Hallucination Detection & Grounding in VLMs

Anirudh Phukan, Divyansh, Harshit Kumar Morj, Vaishnavi, Apoorv Saxena, Koustava Goswami

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

Comments Accepted to NAACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13071 2025-02-19 cs.CV 57%

RobuRCDet: Enhancing Robustness of Radar-Camera Fusion in Bird's Eye View for 3D Object Detection

Jingtong Yue, Zhiwei Lin, Xin Lin, Xiaoyu Zhou, Xiangtai Li, Lu Qi, Yongtao Wang, Ming-Hsuan Yang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ICLR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11724 2025-02-18 cs.CV 57%

Incomplete Modality Disentangled Representation for Ophthalmic Disease Grading and Diagnosis

Chengzhi Liu, Zile Huang, Zhe Chen, Feilong Tang, Yu Tian, Zhongxing Xu, Zihong Luo, Yalin Zheng, Yanda Meng

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 7 Pages, 6 figures

Journal ref AAAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19684 2025-02-18 cs.AI 57%

Boosting Private Domain Understanding of Efficient MLLMs: A Tuning-free, Adaptive, Universal Prompt Optimization Framework

Jiang Liu, Bolin Li, Haoyuan Li, Tianwei Lin, Wenqiao Zhang, Tao Zhong, Zhelun Yu, Jinghao Wei, Hao Cheng, Wanggui He, Fangxun Shu, Hao Jiang, Zheqi Lv, Juncheng Li, Siliang Tang, Yueting Zhuang

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.03677 2025-02-18 cs.CV 57%

L4DR: LiDAR-4DRadar Fusion for Weather-Robust 3D Object Detection

Xun Huang, Ziyu Xu, Hai Wu, Jinlong Wang, Qiming Xia, Yan Xia, Jonathan Li, Kyle Gao, Chenglu Wen, Cheng Wang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Accepted by AAAI2025(Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04815 2025-02-18 q-bio.PE cs.AI 57%

A Review of BioTree Construction in the Context of Information Fusion: Priors, Methods, Applications and Trends

Zelin Zang, Yongjie Xu, Chenrui Duan, Yue Yuan, Jinlin Wu, Zhen Lei, Stan Z. Li

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments 115 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.00129 2025-02-18 cs.CV cs.RO 57%

CMRNext: Camera to LiDAR Matching in the Wild for Localization and Extrinsic Calibration

Daniele Cattaneo, Abhinav Valada

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments Accepted to IEEE Transactions on Robotics (T-RO), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09093 2025-02-14 cs.CV 57%

From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs

Mingxiao Li, Fang Qu, Zhanpeng Chen, Na Su, Zhizhou Zhong, Ziyang Chen, Nan Du, Xiaolong Li

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08486 2025-02-13 cs.CV 57%

Referring Remote Sensing Image Segmentation via Bidirectional Alignment Guided Joint Prediction

Tianxiang Zhang, Zhaokun Wen, Bo Kong, Kecheng Liu, Yisi Zhang, Peixian Zhuang, Jiangyun Li

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12197 2025-02-10 cs.RO cs.AI 57%

Towards Interpretable Visuo-Tactile Predictive Models for Soft Robot Interactions

Enrico Donato, Thomas George Thuruthel, Egidio Falotico

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

Comments IEEE RAS EMBS 10th International Conference on Biomedical Robotics and Biomechatronics (BioRob 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14739 2025-02-07 cs.CV 57%

Maintaining Structural Integrity in Parameter Spaces for Parameter Efficient Fine-tuning

Chongjie Si, Xuehui Wang, Xue Yang, Zhengqin Xu, Qingyun Li, Jifeng Dai, Yu Qiao, Xiaokang Yang, Wei Shen

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 2025 ICLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01959 2025-02-05 cs.CV 57%

MATCNN: Infrared and Visible Image Fusion Method Based on Multi-scale CNN with Attention Transformer

Jingjing Liu, Li Zhang, Xiaoyang Zeng, Wanquan Liu, Jianhua Zhang

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01377 2025-02-04 cs.CE cs.AI 57%

Data-Efficient Model for Psychological Resilience Prediction based on Neurological Data

Zhi Zhang, Yan Liu, Mengxia Gao, Yu Yang, Jiannong Cao, Wai Kai Hou, Shirley Li, Sonata Yau, Yun Kwok Wing, Tatia M. C. Lee

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14568 2025-02-04 eess.IV cs.CV 57%

Policy Gradient-Driven Noise Mask

Mehmet Can Yavuz, Yang Yang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments International Conference on Pattern Recognition (2024) Accepted Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11276 2025-02-03 cs.LG cs.AI eess.SP 57%

Wearable Accelerometer Foundation Models for Health via Knowledge Distillation

Salar Abbaspourazad, Anshuman Mishra, Joseph Futoma, Andrew C. Miller, Ian Shapiro

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.AI

Comments updated format

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08703 2025-01-28 cs.CV 57%

TsCA: On the Semantic Consistency Alignment via Conditional Transport for Compositional Zero-Shot Learning

Miaoge Li, Jingcai Guo, Richard Yi Da Xu, Dongsheng Wang, Xiaofeng Cao, Zhijie Rao, Song Guo

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14859 2025-01-28 cs.CL cs.LG 57%

Dynamic Adaptation of LoRA Fine-Tuning for Efficient and Task-Specific Optimization of Large Language Models

Xiaoxuan Liao, Chihang Wang, Shicheng Zhou, Jiacheng Hu, Hongye Zheng, Jia Gao

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14199 2025-01-27 cs.LG cs.AI cs.ET 57%

Coordinating Ride-Pooling with Public Transit using Reward-Guided Conservative Q-Learning: An Offline Training and Online Fine-Tuning Reinforcement Learning Framework

Yulong Hu, Tingting Dong, Sen Li

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.09466 2025-01-27 eess.IV cs.CV cs.LG 57%

MAMAF-Net: Motion-Aware and Multi-Attention Fusion Network for Stroke Diagnosis

Aysen Degerli, Pekka Jakala, Juha Pajula, Milla Immonen, Miguel Bordallo Lopez

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13479 2025-01-24 cs.LG cs.AI 57%

Adaptive Few-Shot Learning (AFSL): Tackling Data Scarcity with Stability, Robustness, and Versatility

Rishabh Agrawal

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.12481 2025-01-24 cs.LG cs.CV 57%

TT-BLIP: Enhancing Fake News Detection Using BLIP and Tri-Transformer

Eunjee Choi, Jong-Kook Kim

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 8 pages, Accepted 27th International Conference on Information Fusion, FUSION 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09707 2025-01-17 cs.AI 57%

The Goofus & Gallant Story Corpus for Practical Value Alignment

Md Sultan Al Nahian, Tasmia Tasrin, Spencer Frazier, Mark Riedl, Brent Harrison

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

Comments Accepted by International Conference on Machine Learning and Applications (ICMLA) 2024. Main Conference, Long Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07901 2025-01-15 cs.CV eess.IV 57%

Cloud Removal With PolSAR-Optical Data Fusion Using A Two-Flow Residual Network

Yuxi Wang, Wenjuan Zhang, Bing Zhang

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07556 2025-01-14 cs.CV 57%

MatchAnything: Universal Cross-Modality Image Matching with Large-Scale Pre-Training

Xingyi He, Hao Yu, Sida Peng, Dongli Tan, Zehong Shen, Hujun Bao, Xiaowei Zhou

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments Project page: https://zju3dv.github.io/MatchAnything/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15770 2025-01-14 cs.CV 57%

Text-Guided Coarse-to-Fine Fusion Network for Robust Remote Sensing Visual Question Answering

Zhicheng Zhao, Changfu Zhou, Yu Zhang, Chenglong Li, Xiaoliang Ma, Jin Tang

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07048 2025-01-14 cs.AI 57%

Unveiling the Potential of Text in High-Dimensional Time Series Forecasting

Xin Zhou, Weiqing Wang, Shilin Qu, Zhiqiang Zhang, Christoph Bergmeir

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments Accepted by NeurIPS24 TSALM Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12460 2025-01-14 cs.CV 57%

PromptDet: A Lightweight 3D Object Detection Framework with LiDAR Prompts

Kun Guo, Qiang Ling

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Accepted by AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05631 2025-01-13 cs.CV 57%

HFMF: Hierarchical Fusion Meets Multi-Stream Models for Deepfake Detection

Anant Mehta, Bryant McArthur, Nagarjuna Kolloju, Zhengzhong Tu

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments This work is accepted to WACV 2025 Workshop on AI for Multimedia Forensics & Disinformation Detection. Code is available at: https://github.com/taco-group/HFMF

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05075 2025-01-10 cs.AI cs.LG 57%

A Text-Based Knowledge-Embedded Soft Sensing Modeling Approach for General Industrial Process Tasks Based on Large Language Model

Shuo Tong, Han Liu, Runyuan Guo, Xueqiong Tian, Wenqing Wang, Ding Liu, Youmin Zhang

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04815 2025-01-10 cs.CV 57%

Towards Generalizable Trajectory Prediction Using Dual-Level Representation Learning And Adaptive Prompting

Kaouther Messaoud, Matthieu Cord, Alexandre Alahi

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏