arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-17 至 2025-10-17 共收录 7 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 7 篇

2403.15356 2025-10-17 cs.CV 86%

Neural Plasticity-Inspired Multimodal Foundation Model for Earth Observation

Zhitong Xiong, Yi Wang, Fahong Zhang, Adam J. Stewart, Joëlle Hanna, Damian Borth, Ioannis Papoutsis, Bertrand Le Saux, Gustau Camps-Valls, Xiao Xiang Zhu

机构 * Chair of Data Science in Earth Observation, Technical University of Munich (TUM)(地球观测数据科学教授职位,慕尼黑技术大学) Munich Center for Machine Learning(慕尼黑机器学习中心) AIML Lab, School of Computer Science, University of St. Gallen(人工智能实验室,圣加尔登大学计算机科学学院) School of Rural, Surveying and Geoinformatics Engineering, National Technical University of Athens(农村、测绘与地理信息工程学院,国家技术大学雅典) Image Processing Laboratory (IPL), Universitat de València(图像处理实验室(IPL),瓦伦西亚大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(title);分类 cs.CV

Comments 18 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14344 2025-10-17 cs.CR cs.AI 79%

BinCtx: Multi-Modal Representation Learning for Robust Android App Behavior Detection

Zichen Liu, Shao Yang, Xusheng Xiao

机构 * Arizona State University(亚利桑那州立大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20401 2025-10-17 cs.CV cs.RO 79%

SGAligner++: Cross-Modal Language-Aided 3D Scene Graph Alignment

Binod Singh, Sayan Deb Sarkar, Iro Armeni

机构 * Technical University of Munich(慕尼黑技术大学) Stanford University(斯坦福大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments Project Page: https://singhbino3d.github.io/sgpp/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14411 2025-10-17 cs.LG cs.MM cs.SD eess.AS 74%

Revisit Modality Imbalance at the Decision Layer

Xiaoyu Ma, Hao Chen

机构 * School of Computer Science and Engineering, Southeast University, Nanjing, China(计算机科学与工程学院,东南大学,南京,中国) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及其交叉应用重点实验室(东南大学),教育部,中国)

专题命中 多模态训练与对齐 :multimodal(abstract,comments);audio-visual(abstract);分类 cs.MM、eess.AS

Comments Some Insights in Balanced Multimodal Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14387 2025-10-17 cs.AI 70%

Can MLLMs Absorb Math Reasoning Abilities from LLMs as Free Lunch?

Yijie Hu, Zihao Zhou, Kaizhu Huang, Xiaowei Huang, Qiufeng Wang

机构 * Xi’an-Jiaotong Liverpool University(西交利物浦大学) University of Liverpool(利物浦大学) Duke Kunshan University(杜克昆山大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14374 2025-10-17 cs.CV 70%

Spatial Preference Rewarding for MLLMs Spatial Understanding

Han Qiu, Peng Gao, Lewei Lu, Xiaoqin Zhang, Ling Shao, Shijian Lu

机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室) Shanghai AI Laboratory(上海人工智能实验室) Sensetime Research(商汤科技研究院) Zhejiang University of Technology(浙江工业大学) UCAS-Terminus AI Lab,University of Chinese Academy of Sciences(中国科学院大学Terminus AI实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01411 2025-10-17 cs.CV cs.AI 62%

ViTA-PAR: Visual and Textual Attribute Alignment with Attribute Prompting for Pedestrian Attribute Recognition

Minjeong Park, Hongbeen Park, Jinkyu Kim

机构 * Department of Computer Science and Engineering, Korea University, Seoul 02841, Korea(计算机科学与工程系,韩国大学,首尔02841,韩国)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted to IEEE ICIP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏