arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-03 至 2025-09-03 共收录 18 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 18 篇

2509.00622 2025-09-03 cs.AI cs.IR 83%

BALM-TSF: Balanced Multimodal Alignment for LLM-Based Time Series Forecasting

Shiqiao Zhou, Holger Schöner, Huanbo Lyu, Edouard Fouché, Shuo Wang

机构 * University of Birmingham(伯明翰大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00664 2025-09-03 cs.CV cs.AI 81%

Fusion to Enhance: Fusion Visual Encoder to Enhance Multimodal Language Model

Yifei She, Huangxuan Wu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06594 2025-09-03 q-fin.CP cs.AI cs.LG 79%

Stock Movement Prediction with Multimodal Stable Fusion via Gated Cross-Attention Mechanism

Chang Zong, Hang Zhou

机构 * School of Information and Electronic Engineering, Zhejiang University of Science and Technology(浙江理工大学信息与电子工程学院) Department of Finance Accounting and Economics, Business School of Nottingham University(诺丁汉大学商学院金融会计与经济学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 14 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18057 2025-09-03 cs.SD cs.AI 79%

Dynamic Fusion Multimodal Network for SpeechWellness Detection

Wenqiang Sun, Han Yin, Jisheng Bai, Jianfeng Chen

机构 * Northwestern Polytechnical University(西北工业大学) School of Electrical Engineering, KAIST(韩国成均馆大学电气工程学院) LianFeng Acoustic Technologies Co., Ltd.(联丰声学科技有限公司) Xi’an University of Posts & Telecommunications(西安邮电大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 6 pages, 5figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00677 2025-09-03 cs.CV 79%

CSFMamba: Cross State Fusion Mamba Operator for Multimodal Remote Sensing Image Classification

Qingyu Wang, Xue Jiang, Guozheng Xu

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 5 pages, 2 figures, accpeted by 2025 IEEE International Geoscience and Remote Sensing Symposium(IGARSS 2025),not published yet

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02275 2025-09-03 cs.RO 78%

Human-Inspired Soft Anthropomorphic Hand System for Neuromorphic Object and Pose Recognition Using Multimodal Signals

Fengyi Wang, Xiangyu Fu, Nitish Thakor, Gordon Cheng

机构 * Institute for Cognitive Systems, Technical University of Munich(认知系统研究所,慕尼黑技术大学) Department of Biomedical Engineering, Johns Hopkins University(生物医学工程系,约翰霍普金斯大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00419 2025-09-03 cs.CV 74%

LightVLM: Acceleraing Large Multimodal Models with Pyramid Token Merging and KV Cache Compression

Lianyu Hu, Fanhua Shang, Wei Feng, Liang Wan

机构 * College of Intelligence and Computing, Tianjin University(智能与计算学院,天津大学)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

Comments EMNLP2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00752 2025-09-03 cs.CV 70%

Multi-Level CLS Token Fusion for Contrastive Learning in Endoscopy Image Classification

Y Hop Nguyen, Doan Anh Phan Huu, Trung Thai Tran, Nhat Nam Mai, Van Toi Giap, Thao Thi Phuong Dao, Trung-Nghia Le

机构 * University of Science, VNU-HCM(越南胡志明市科学大学) Thong Nhat Hospital(通纳特医院)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00371 2025-09-03 cs.CV 70%

Two Causes, Not One: Rethinking Omission and Fabrication Hallucinations in MLLMs

Guangzong Si, Hao Yin, Xianfei Li, Qing Ding, Wenlong Liao, Tao He, Pai Peng

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Preprint,Underreview

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01214 2025-09-03 cs.CV cs.MM 62%

PRINTER:Deformation-Aware Adversarial Learning for Virtual IHC Staining with In Situ Fidelity

Yizhe Yuan, Bingsen Xue, Bangzheng Pu, Chengxiang Wang, Cheng Jin

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.MM

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00936 2025-09-03 cs.AI 57%

UrbanInsight: A Distributed Edge Computing Framework with LLM-Powered Data Filtering for Smart City Digital Twins

Kishor Datta Gupta, Md Manjurul Ahsan, Mohd Ariful Haque, Roy George, Azmine Toushik Wasi

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00346 2025-09-03 cs.CV 57%

LUT-Fuse: Towards Extremely Fast Infrared and Visible Image Fusion via Distillation to Learnable Look-Up Tables

Xunpeng Yi, Yibing Zhang, Xinyu Xiang, Qinglong Yan, Han Xu, Jiayi Ma

机构 * Electronic Information School, Wuhan University(武汉大学电子信息学院) School of Automation, Southeast University(东南大学自动化学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15085 2025-09-03 cs.CV 57%

Cognitive-Inspired Hierarchical Attention Fusion With Visual and Textual for Cross-Domain Sequential Recommendation

Wangyu Wu, Zhenhong Chen, Siqi Song, Xianglin Qiu, Xiaowei Huang, Fei Ma, Jimin Xiao

机构 * Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学) The University of Liverpool(利物浦大学) Microsoft(微软公司)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted at CogSCI 2025. arXiv admin note: text overlap with arXiv:2502.15694

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02011 2025-09-03 cs.RO 50%

Generalizing Unsupervised Lidar Odometry Model from Normal to Snowy Weather Conditions

Beibei Zhou, Zhiyuan Zhang, Zhenbo Song, Jianhui Guo, Hui Kong

机构 * Shanghai Polytechnic University(上海理工大学) Singapore Management University(新加坡国立大学) Nanjing University of Science and Technology(南京理工大学) University of Macau(澳门大学)

专题命中 多模态训练与对齐 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01611 2025-09-03 cs.RO 50%

A Hybrid Input based Deep Reinforcement Learning for Lane Change Decision-Making of Autonomous Vehicle

Ziteng Gao, Jiaqi Qu, Chaoyu Chen

专题命中 多模态训练与对齐 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06156 2025-09-03 cs.RO 50%

ViTaMIn: Learning Contact-Rich Tasks Through Robot-Free Visuo-Tactile Manipulation Interface

Fangchen Liu, Chuanyu Li, Yihua Qin, Jing Xu, Pieter Abbeel, Rui Chen

机构 * Tsinghua University(清华大学) University of California, Berkeley(加州大学伯克利分校)

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17761 2025-09-03 cs.LG 50%

Towards a Unified Textual Graph Framework for Spectral Reasoning via Physical and Chemical Information Fusion

Jiheng Liang, Ziru Yu, Zujie Xie, Yuchen Guo, Yulan Guo, Xiangyang Yu

专题命中 多模态训练与对齐 :multi-modal(abstract)

Comments We need to further modify and supplement the experiment

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19338 2025-09-03 cs.LG cs.CR 50%

Membership Inference Attacks on Large-Scale Models: A Survey

Hengyu Wu, Yang Cao

机构 * Institute of Science Tokyo(东京科学研究所)

专题命中 多模态训练与对齐 :multimodal(abstract)

Comments Preprint. Submitted for peer review. The final version may differ

详情

展开后加载摘要…

URL PDF HTML 收藏