arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-29 至 2025-09-29 共收录 79 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 20 篇

2503.10434 2025-09-29 cs.RO cs.CV cs.LG 57%

Learning Personalized Driving Styles via Reinforcement Learning from Human Feedback

Derun Li, Changye Li, Yue Wang, Jianwei Ren, Xin Wen, Pengxiang Li, Leimeng Xu, Kun Zhan, Peng Jia, Xianpeng Lang, Ningyi Xu, Hang Zhao

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Qi Zhi Institute(上海启智研究院) Peking University(北京大学) LiAuto Tsinghua University(清华大学)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

Comments 20 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18788 2025-09-29 cs.RO cs.CV 57%

Excavating in the Wild: The GOOSE-Ex Dataset for Semantic Segmentation

Raphael Hagmanns, Peter Mortimer, Miguel Granero, Thorsten Luettel, Janko Petereit

机构 * Fraunhofer Institute of Optronics, System Technologies and Image Exploitation(弗劳恩霍夫光学与系统技术研究所) Institute for Autonomous Systems Technology, University of the Bundeswehr Munich(联邦国防军慕尼黑大学自主系统技术研究所) Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments Accepted for publication at 2025 IEEE International Conference on Robotics and Automation (ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21341 2025-09-29 cs.NE cs.AI cs.LG 57%

From Embeddings to Equations: Genetic-Programming Surrogates for Interpretable Transformer Classification

Mohammad Sadegh Khorshidi, Navid Yazdanjue, Hassan Gharoun, Mohammad Reza Nikoo, Fang Chen, Amir H. Gandomi

机构 * Faculty of Engineering & Information Technology, University of Technology Sydney(工程与信息技术学院,技术悉尼大学) Department of Civil and Architectural Engineering, Sultan Qaboos University(土木与建筑学院,苏丹·卡布斯大学) University Research and Innovation Center (EKIK), Obuda University(研究与创新中心(EKIK),布达佩斯大学)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.AI

Comments 20 pages, 8 tables, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22334 2025-09-29 cs.CY 50%

MakOne: Behavioural Data of University Students' Smart Devices in Uganda

Michael Kizito, Ivan Kayongo, Hawa Nyende, Halimu Chongomweru, Lillian Muyama, Roy Alia Asiku, Alice Mugisha

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04278 2025-09-29 cs.HC 50%

EmoPrefer: Can Large Language Models Understand Human Emotion Preferences?

Zheng Lian, Licai Sun, Lan Chen, Haoyu Chen, Zebang Cheng, Fan Zhang, Ziyu Jia, Ziyang Ma, Fei Ma, Xiaojiang Peng, Jianhua Tao

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 4 篇

2509.21662 2025-09-29 cs.LG 82%

MMPlanner: Zero-Shot Multimodal Procedural Planning with Chain-of-Thought Object State Reasoning

Afrina Tabassum, Bin Guo, Xiyao Ma, Hoda Eldardiry, Ismini Lourentzou

机构 * Amazon(亚马逊公司) Alexa, Amazon(亚马逊Alexa部门) Virginia Tech(弗吉尼亚理工大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract)

Comments 17 pages, 9 figures, 14 tables, Findings of the Association for Computational Linguistics: EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08524 2025-09-29 cs.HC cs.AI 79%

StreetReaderAI: Making Street View Accessible Using Context-Aware Multimodal AI

Jon E. Froehlich, Alexander Fiannaca, Nimer Jaber, Victor Tsaran, Shaun Kane

机构 * Google Research(谷歌研究) Google DeepMind(谷歌DeepMind) Google(谷歌)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Comments Accepted to UIST'25; v2. Fixed a missing word in the PDF; v3. Fixed a typo in an author's name; v4. Changed system name and title

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22441 2025-09-29 cs.RO 67%

UnderwaterVLA: Dual-brain Vision-Language-Action architecture for Autonomous Underwater Navigation

Zhangyuan Wang, Yunpeng Zhu, Yuqi Yan, Xiaoyuan Tian, Xinhao Shao, Meixuan Li, Weikun Li, Guangsheng Su, Weicheng Cui, Dixia Fan

机构 * School of Engineering, Westlake University(西lake大学工程学院) College of Information Science & Electronic Engineering, Zhejiang University(浙江大学信息科学与电子工程学院) College of Environmental and Resource Sciences, Zhejiang University(浙江大学环境与资源科学学院) Australian National University(澳大利亚国立大学)

专题命中 多模态Agent :multimodal(abstract);multimodal foundation model(abstract)

Comments This paper introduces the first VLA framework for AUVs, featuring a dual-brain architecture and zero-data MPC for real-world underwater navigation

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22631 2025-09-29 cs.CV cs.CL 62%

LABELING COPILOT: A Deep Research Agent for Automated Data Curation in Computer Vision

Debargha Ganguly, Sumit Kumar, Ishwar Balappanawar, Weicong Chen, Shashank Kambhatla, Srinivasan Iyengar, Shivkumar Kalyanaraman, Ponnurangam Kumaraguru, Vipin Chaudhary

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 8 篇

2509.21358 2025-09-29 cs.CV cs.AI 90%

MDF-MLLM: Deep Fusion Through Cross-Modal Feature Alignment for Contextually Aware Fundoscopic Image Classification

Jason Jordan, Mohammadreza Akbari Lor, Peter Koulen, Mei-Ling Shyu, Shu-Ching Chen

专题命中 多模态训练与对齐 :MLLM(title,abstract);cross-modal(title);multimodal(abstract);image-text(abstract)

Comments Word count: 5157, Table count: 2, Figure count: 5

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10143 2025-09-29 cs.LG cs.CV 89%

On the Value of Cross-Modal Misalignment in Multimodal Representation Learning

Yichao Cai, Yuhang Liu, Erdun Gao, Tianjiao Jiang, Zhen Zhang, Anton van den Hengel, Javen Qinfeng Shi

机构 * Australian Institute for Machine Learning(澳大利亚机器学习研究所) The University of Adelaide(阿德莱德大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments NeurIPS 2025 camera-ready version (with checklist removed for presentation clarity)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22261 2025-09-29 cs.AI cs.CL 81%

InfiMed-Foundation: Pioneering Advanced Multimodal Medical Models with Compute-Efficient Pre-Training and Multi-Stage Fine-Tuning

Guanghao Zhu, Zhitian Hou, Zeyu Liu, Zhijie Sang, Congkai Xie, Hongxia Yang

机构 * The Hong Kong Polytechnic University(香港理工大学) Sun Yat-sen University(中山大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21854 2025-09-29 cs.MM cs.CV 81%

Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization

Songjun Tu, Qichao Zhang, Jingbo Sun, Yuqian Fu, Linjing Li, Xiangyuan Lan, Dongmei Jiang, Yaowei Wang, Dongbin Zhao

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) Pengcheng Laboratory(鹏城实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments 12pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12716 2025-09-29 cs.CL cs.AI 62%

Shadow-FT: Tuning Instruct Model via Training on Paired Base Model

Taiqiang Wu, Runming Yang, Jiayi Li, Pengfei Hu, Yik-Chung Wu, Ngai Wong, Yujiu Yang

机构 * The University of Hong Kong(香港大学) Tsinghua University(清华大学) Tencent(腾讯)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 24 pages, 12 tables, 8 figures. Previous name: Shadow-FT: Tuning Instruct via Base

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.05798 2025-09-29 cs.CV cs.AI cs.LG 62%

Multi-View Hypercomplex Learning for Breast Cancer Screening

Eleonora Lopez, Eleonora Grassucci, Danilo Comminiello

机构 * Department of Information Engineering, Electronics and Telecommunications (DIET), Sapienza University of Rome(信息工程、电子与电信系(DIET),罗马萨皮恩扎大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments This paper has been submitted to Expert Systems with Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14364 2025-09-29 cs.CL 57%

Position IDs Matter: An Enhanced Position Layout for Efficient Context Compression in Large Language Models

Runsong Zhao, Xin Liu, Xinyu Liu, Pengcheng Huang, Chunyang Xiao, Tong Xiao, Jingbo Zhu

机构 * NLP Lab, School of Computer Science and Engineering, Northeastern University, Shenyang, China(东北大学计算机科学与工程学院自然语言处理实验室) NiuTrans Research, Shenyang, China(NiuTrans研究院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22195 2025-09-29 cs.RO 50%

Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic Forgetting

Asher J. Hancock, Xindi Wu, Lihan Zha, Olga Russakovsky, Anirudha Majumdar

机构 * Princeton University(普林斯顿大学)

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 其他多模态 2 篇

2509.22079 2025-09-29 physics.optics 78%

Multimodal Speckle-polarization Fiber-optic Sensing for Localized and High-bandwidth Vibration Monitoring

Catarina S. Monteiro, Tomás Lopes, Joana Teixeira, Tiago Ferreira, Pedro A. S. Jorge, Nuno A. Silva

专题命中 其他多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13920 2025-09-29 cs.AI 57%

Integrating Knowledge Graphs and Bayesian Networks: A Hybrid Approach for Explainable Disease Risk Prediction

Mbithe Nzomo, Deshendran Moodley

机构 * Centre for Artificial Intelligence Research (CAIR)(人工智能研究中心) Department of Computer Science(计算机科学系) University of Cape Town(开普敦大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments This work has been accepted for presentation at the 49th IEEE International Conference on Computers, Software, and Applications (COMPSAC 2025). The final published version will be available via IEEE Xplore

Journal ref Proceedings of the 2025 IEEE 49th Annual Computers, Software, and Applications Conference (COMPSAC), Toronto, ON, Canada, 2025, pp. 834-844

详情

展开后加载摘要…

URL PDF HTML 收藏