arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-23 至 2025-10-23 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 9 篇

2412.12718 2025-10-23 cs.CV cs.MM 88%

ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding

Zhenxing Zhang, Yaxiong Wang, Lechao Cheng, Zhun Zhong, Dan Guo, Meng Wang

机构 * School of Computer Science and Information Engineering, Hefei University of Technology, China(计算机科学与信息工程学院,合肥工业大学,中国) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center, China(人工智能研究所,合肥综合性国家科学中心,中国)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract);image-text(abstract)

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16895 2025-10-23 cs.CV cs.AI cs.LG 84%

With Limited Data for Multimodal Alignment, Let the STRUCTURE Guide You

Fabian Gröger, Shuo Wen, Huyen Le, Maria Brbić

机构 * EPFL(瑞士联邦理工学院) University of Basel(巴塞尔大学) HSLU(苏黎世联邦理工学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2025 camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19336 2025-10-23 cs.CV 83%

DaMo: Data Mixing Optimizer in Fine-tuning Multimodal LLMs for Mobile Phone Agents

Kai Shi, Jun Yang, Ni Yang, Binqiang Pan, Qingsong Xie, Chao Zhang, Zhenyu Yang, Tianhuang Su, Haonan Lu

机构 * OPPO AI Center(OPPO人工智能中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18279 2025-10-23 cs.CL cs.AI 81%

Text or Pixels? It Takes Half: On the Token Efficiency of Visual Text Inputs in Multimodal LLMs

Yanhong Li, Zixuan Lan, Jiawei Zhou

机构 * Allen Institute for AI(艾伦人工智能研究所) University of Chicago(芝加哥大学) Stony Brook University(石溪大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025 Findings ("Text or Pixels? Evaluating Efficiency and Understanding of LLMs with Visual Text Inputs")

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19215 2025-10-23 cs.CV 70%

SFGFusion: Surface Fitting Guided 3D Object Detection with 4D Radar and Camera Fusion

Xiaozhi Li, Huijun Di, Jian Li, Feng Liu, Wei Liang

机构 * Radar Technology Research Institute, School of Information and Electronics, Beijing Institute of Technology(雷达技术研究院,信息电子学院,北京理工大学) School of Computer Science and Technology, Beijing Institute of Technology(计算机科学与技术学院,北京理工大学) Innovative Equipment Research Institute, Beijing Institute of Technology(创新装备研究院,北京理工大学) Key Laboratory of Electronic and Information Technology in Satellite Navigation (Beijing Institute of Technology), Ministry of Education(卫星导航电子信息技术重点实验室(北京理工大学),教育部) Beijing Racobit Electronic Information Technology Co., Ltd.(北京瑞科比特电子信息技术有限公司)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments Submitted to Pattern Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19384 2025-10-23 cs.LG 67%

Learning Noise-Resilient and Transferable Graph-Text Alignment via Dynamic Quality Assessment

Yuhang Liu, Minglai Shao, Zengyi Wo, Yunlong Chu, Bing Hao, Shengzhong Liu, Ruijie Wang, Jianxin Li

机构 * School of New Media and Communication, Tianjin University(新媒体与传播学院,天津大学) Baidu(百度) Shanghai Jiao Tong University(上海交通大学) School of Computer Science and Engineering, Beihang University(计算机科学与工程学院,北航)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19520 2025-10-23 cs.MM 57%

CDI-DTI: A Strong Cross-domain Interpretable Drug-Target Interaction Prediction Framework Based on Multi-Strategy Fusion

Xiangyu Li, Haojie Yang, Kaimiao Hu, Runzhi Wu, Liangliang Liu, Ran Su

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19484 2025-10-23 q-bio.BM cs.AI cs.LG 57%

KnowMol: Advancing Molecular Large Language Models with Multi-Level Chemical Knowledge

Zaifei Yang, Hong Chang, Ruibing Hou, Shiguang Shan, Xilin Chen

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, CAS, China(人工智能安全国家重点实验室,计算技术研究所,中国科学院,中国) University of Chinese Academy of Sciences (CAS), China(中国科学院大学(中国科学院))

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19078 2025-10-23 cs.CV 57%

UniHPR: Unified Human Pose Representation via Singular Value Contrastive Learning

Zhongyu Jiang, Wenhao Chai, Lei Li, Zhuoran Zhou, Cheng-Yen Yang, Jenq-Neng Hwang

机构 * University of Washington(华盛顿大学) University of Copenhagen(哥本哈根大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏