arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-03 至 2025-09-03 共收录 120 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 10 篇

2504.06661 2025-09-03 cs.RO 50%

Domain-Conditioned Scene Graphs for State-Grounded Task Planning

Jonas Herzog, Jiangpin Liu, Yue Wang

机构 * Zhejiang University(浙江大学)

专题命中 多模态Agent :multimodal(abstract)

Comments Accepted for IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态训练与对齐 18 篇

2509.00622 2025-09-03 cs.AI cs.IR 83%

BALM-TSF: Balanced Multimodal Alignment for LLM-Based Time Series Forecasting

Shiqiao Zhou, Holger Schöner, Huanbo Lyu, Edouard Fouché, Shuo Wang

机构 * University of Birmingham(伯明翰大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00664 2025-09-03 cs.CV cs.AI 81%

Fusion to Enhance: Fusion Visual Encoder to Enhance Multimodal Language Model

Yifei She, Huangxuan Wu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06594 2025-09-03 q-fin.CP cs.AI cs.LG 79%

Stock Movement Prediction with Multimodal Stable Fusion via Gated Cross-Attention Mechanism

Chang Zong, Hang Zhou

机构 * School of Information and Electronic Engineering, Zhejiang University of Science and Technology(浙江理工大学信息与电子工程学院) Department of Finance Accounting and Economics, Business School of Nottingham University(诺丁汉大学商学院金融会计与经济学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 14 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18057 2025-09-03 cs.SD cs.AI 79%

Dynamic Fusion Multimodal Network for SpeechWellness Detection

Wenqiang Sun, Han Yin, Jisheng Bai, Jianfeng Chen

机构 * Northwestern Polytechnical University(西北工业大学) School of Electrical Engineering, KAIST(韩国成均馆大学电气工程学院) LianFeng Acoustic Technologies Co., Ltd.(联丰声学科技有限公司) Xi’an University of Posts & Telecommunications(西安邮电大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 6 pages, 5figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00677 2025-09-03 cs.CV 79%

CSFMamba: Cross State Fusion Mamba Operator for Multimodal Remote Sensing Image Classification

Qingyu Wang, Xue Jiang, Guozheng Xu

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 5 pages, 2 figures, accpeted by 2025 IEEE International Geoscience and Remote Sensing Symposium(IGARSS 2025),not published yet

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02275 2025-09-03 cs.RO 78%

Human-Inspired Soft Anthropomorphic Hand System for Neuromorphic Object and Pose Recognition Using Multimodal Signals

Fengyi Wang, Xiangyu Fu, Nitish Thakor, Gordon Cheng

机构 * Institute for Cognitive Systems, Technical University of Munich(认知系统研究所,慕尼黑技术大学) Department of Biomedical Engineering, Johns Hopkins University(生物医学工程系,约翰霍普金斯大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00419 2025-09-03 cs.CV 74%

LightVLM: Acceleraing Large Multimodal Models with Pyramid Token Merging and KV Cache Compression

Lianyu Hu, Fanhua Shang, Wei Feng, Liang Wan

机构 * College of Intelligence and Computing, Tianjin University(智能与计算学院,天津大学)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

Comments EMNLP2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00752 2025-09-03 cs.CV 70%

Multi-Level CLS Token Fusion for Contrastive Learning in Endoscopy Image Classification

Y Hop Nguyen, Doan Anh Phan Huu, Trung Thai Tran, Nhat Nam Mai, Van Toi Giap, Thao Thi Phuong Dao, Trung-Nghia Le

机构 * University of Science, VNU-HCM(越南胡志明市科学大学) Thong Nhat Hospital(通纳特医院)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00371 2025-09-03 cs.CV 70%

Two Causes, Not One: Rethinking Omission and Fabrication Hallucinations in MLLMs

Guangzong Si, Hao Yin, Xianfei Li, Qing Ding, Wenlong Liao, Tao He, Pai Peng

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Preprint,Underreview

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01214 2025-09-03 cs.CV cs.MM 62%

PRINTER:Deformation-Aware Adversarial Learning for Virtual IHC Staining with In Situ Fidelity

Yizhe Yuan, Bingsen Xue, Bangzheng Pu, Chengxiang Wang, Cheng Jin

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.MM

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00936 2025-09-03 cs.AI 57%

UrbanInsight: A Distributed Edge Computing Framework with LLM-Powered Data Filtering for Smart City Digital Twins

Kishor Datta Gupta, Md Manjurul Ahsan, Mohd Ariful Haque, Roy George, Azmine Toushik Wasi

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00346 2025-09-03 cs.CV 57%

LUT-Fuse: Towards Extremely Fast Infrared and Visible Image Fusion via Distillation to Learnable Look-Up Tables

Xunpeng Yi, Yibing Zhang, Xinyu Xiang, Qinglong Yan, Han Xu, Jiayi Ma

机构 * Electronic Information School, Wuhan University(武汉大学电子信息学院) School of Automation, Southeast University(东南大学自动化学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15085 2025-09-03 cs.CV 57%

Cognitive-Inspired Hierarchical Attention Fusion With Visual and Textual for Cross-Domain Sequential Recommendation

Wangyu Wu, Zhenhong Chen, Siqi Song, Xianglin Qiu, Xiaowei Huang, Fei Ma, Jimin Xiao

机构 * Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学) The University of Liverpool(利物浦大学) Microsoft(微软公司)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted at CogSCI 2025. arXiv admin note: text overlap with arXiv:2502.15694

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02011 2025-09-03 cs.RO 50%

Generalizing Unsupervised Lidar Odometry Model from Normal to Snowy Weather Conditions

Beibei Zhou, Zhiyuan Zhang, Zhenbo Song, Jianhui Guo, Hui Kong

机构 * Shanghai Polytechnic University(上海理工大学) Singapore Management University(新加坡国立大学) Nanjing University of Science and Technology(南京理工大学) University of Macau(澳门大学)

专题命中 多模态训练与对齐 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01611 2025-09-03 cs.RO 50%

A Hybrid Input based Deep Reinforcement Learning for Lane Change Decision-Making of Autonomous Vehicle

Ziteng Gao, Jiaqi Qu, Chaoyu Chen

专题命中 多模态训练与对齐 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06156 2025-09-03 cs.RO 50%

ViTaMIn: Learning Contact-Rich Tasks Through Robot-Free Visuo-Tactile Manipulation Interface

Fangchen Liu, Chuanyu Li, Yihua Qin, Jing Xu, Pieter Abbeel, Rui Chen

机构 * Tsinghua University(清华大学) University of California, Berkeley(加州大学伯克利分校)

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17761 2025-09-03 cs.LG 50%

Towards a Unified Textual Graph Framework for Spectral Reasoning via Physical and Chemical Information Fusion

Jiheng Liang, Ziru Yu, Zujie Xie, Yuchen Guo, Yulan Guo, Xiangyang Yu

专题命中 多模态训练与对齐 :multi-modal(abstract)

Comments We need to further modify and supplement the experiment

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19338 2025-09-03 cs.LG cs.CR 50%

Membership Inference Attacks on Large-Scale Models: A Survey

Hengyu Wu, Yang Cao

机构 * Institute of Science Tokyo(东京科学研究所)

专题命中 多模态训练与对齐 :multimodal(abstract)

Comments Preprint. Submitted for peer review. The final version may differ

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他多模态 11 篇

2509.01337 2025-09-03 cs.MM cs.AI cs.CL 82%

LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition

Qianrui Zhou, Hua Xu, Yifan Wang, Xinzhi Dong, Hanlei Zhang

机构 * Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) School of Information Science and Engineering, Hebei University of Science and Technology(河北科技大学信息科学与工程学院)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

Comments Accepted by EMNLP 2025 (Main Track, Long Paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01996 2025-09-03 cs.RO cs.HC 78%

MIRAGE: Multimodal Intention Recognition and Admittance-Guided Enhancement in VR-based Multi-object Teleoperation

Chi Sun, Xian Wang, Abhishek Kumar, Chengbin Cui, Lik-Hang Lee

机构 * The Hong Kong Polytechnic University(香港理工大学) University of Jyväskylä(于韦斯屈莱大学)

专题命中 其他多模态 :multimodal(title,abstract)

Comments Accepted by ISMAR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01164 2025-09-03 cs.LG eess.IV 78%

A Multimodal Deep Learning Framework for Early Diagnosis of Liver Cancer via Optimized BiLSTM-AM-VMD Architecture

Cheng Cheng, Zeping Chen, Xavier Wang

机构 * Department of Encephalopathy, Chengdu Pidu District Hospital of Traditional Chinese Medicine, Chengdu 611730, China(成都Pidu区中医医院神经病科) Department of Tuina, Chengdu Pidu District Hospital of Traditional Chinese Medicine, Chengdu 611730, China(成都Pidu区中医医院推拿科) Department of Electrical and Computer Engineering, Carnegie Mellon University, Pittsburgh, PA 15213, USA(卡内基梅隆大学电气与计算机工程系)

专题命中 其他多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01123 2025-09-03 eess.SY cs.SI cs.SY 78%

Using Gaussian Mixtures to Model Evolving Multi-Modal Beliefs Across Social Media

Yijun Chen, Farhad Farokhi, Yutong Bu, Nicholas Kah Yean Low, Jarra Horstman, Julian Greentree, Robin Evans, Andrew Melatos

专题命中 其他多模态 :multi-modal(title,abstract)

Comments 8 pages, 5 figures, IEEE Conference on Decision and Control

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00775 2025-09-03 physics.optics 78%

Multimodal near-infrared spectroscopy for biosensing applications

Krisztian Neutsch, Jana Nikolic, Asra Mafakheri, Jan Stegemann, Sebastian Kruss

专题命中 其他多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02540 2025-09-03 eess.SP 50%

LLM-Enhanced Space-Air-Ground-Sea Integrated Networks

Halvin Yang, Sangarapillai Lambotharan, Mahsa Derakhshani, Lajos Hanzo

专题命中 其他多模态 :multimodal(abstract)

Comments 6 figures, 7 pages, magazine

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02003 2025-09-03 cs.LG stat.ML 50%

Bouncy particle sampler with infinite exchanging parallel tempering

Yohei Saito, Shun Kimura, Koujin Takeda

机构 * Center for Mathematics and Data Science, Gunma University(数学与数据科学中心,gunma大学) Graduate School of Science and Engineering, Ibaraki University(科学与工程研究生院,ibaraki大学)

专题命中 其他多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01609 2025-09-03 cs.HC 50%

Quantifying the Effect of Thermal Illusions in Virtual Reality

Yannick Weiss, Marlene Eder, Oguzhan Cesur, Steeven Villa

专题命中 其他多模态 :cross-modal(abstract)

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22903 2025-09-03 eess.SP cs.IT math.IT 50%

Limited Feedback in RIS-Assisted Wireless Communications: Use Cases, Challenges, and Future Directions

Weicong Chen, Jiajia Guo, Yiming Cui, Xiao Li, Shi Jin

专题命中 其他多模态 :multi-modal(abstract)

Comments This work has been submitted for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00863 2025-09-03 cs.LG 50%

Predicting Multi-Type Talented Students in Secondary School Using Semi-Supervised Machine Learning

Xinzhe Zheng, Zhen-Qun Yang, Jiannong Cao, Jiabei Cheng

机构 * The Hong Kong Polytechnic University(香港理工大学)

专题命中 其他多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18464 2025-09-03 hep-ph stat.ML 50%

A comparison of Bayesian sampling algorithms for high-dimensional particle physics and cosmology applications

Joshua Albert, Csaba Balazs, Andrew Fowlie, Will Handley, Nicholas Hunt-Smith, Roberto Ruiz de Austri, Martin White

专题命中 其他多模态 :multimodal(abstract)

Comments 45 pages, 8 figures, 20 tables

详情

展开后加载摘要…

URL PDF HTML 收藏