arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-23 至 2025-10-23 共收录 50 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 11 篇

2509.18582 2025-10-23 cs.CV 79%

The Photographer Eye: Teaching Multimodal Large Language Models to Understand Image Aesthetics like Photographers

Daiqing Qi, Handong Zhao, Jing Shi, Simon Jenni, Yifei Fan, Franck Dernoncourt, Scott Cohen, Sheng Li

机构 * University of Virginia(弗吉尼亚大学) Adobe(Adobe公司)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Journal ref CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.19339 2025-10-23 cs.CV cs.CL 73%

PixelWorld: How Far Are We from Perceiving Everything as Pixels?

Zhiheng Lyu, Xueguang Ma, Wenhu Chen

机构 * University of Waterloo(滑铁卢大学) Vector Institute, Toronto(多伦多向量研究所)

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19599 2025-10-23 cs.CV cs.AI 62%

XBench: A Comprehensive Benchmark for Visual-Language Explanations in Chest Radiography

Haozhe Luo, Shelley Zixin Shu, Ziyu Zhou, Sebastian Otalora, Mauricio Reyes

机构 * ARTORG Center for Biomedical Engineering Research, University of Bern, Switzerland(ARTORG生物医学工程研究中心,伯尔尼大学,瑞士) Shanghai Jiao Tong University, China(上海交通大学,中国) Kaiko.AI, Switzerland(Kaiko.AI,瑞士) Dept. of Radiation Oncology, Inselspital, Bern University Hospital(放射肿瘤科,因斯普尔茨医院,伯尔尼大学医院)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05571 2025-10-23 cs.CL 57%

Presenting a Paper is an Art: Self-Improvement Aesthetic Agents for Academic Presentations

Chengzhi Liu, Yuzhe Yang, Kaiwen Zhou, Zhen Zhang, Yue Fan, Yanan Xie, Peng Qi, Xin Eric Wang

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09024 2025-10-23 q-bio.NC cs.CV 57%

CNeuroMod-THINGS, a densely-sampled fMRI dataset for visual neuroscience

Marie St-Laurent, Basile Pinsard, Oliver Contier, Elizabeth DuPre, Katja Seeliger, Valentina Borghesani, Julie A. Boyle, Lune Bellec, Martin N. Hebart

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments 16 pages manuscript, 5 figures, 9 pages supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10043 2025-10-23 cs.SE 50%

TrioXpert: An Automated Incident Management Framework for Microservice System

Yongqian Sun, Yu Luo, Xidao Wen, Yuan Yuan, Xiaohui Nie, Shenglin Zhang, Tong Liu, Xi Luo

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 3 篇

2510.19497 2025-10-23 cs.MA cs.AI 79%

Modeling realistic human behavior using generative agents in a multimodal transport system: Software architecture and Application to Toulouse

Trung-Dung Vu, Benoit Gaudou, Kamaldeep Singh Oberoi

机构 * UMR IRIT University Toulouse Capitole(IRIT大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19245 2025-10-23 cs.CY cs.AI cs.HC cs.LG cs.MM 62%

See, Think, Act: Online Shopper Behavior Simulation with VLM Agents

Yimeng Zhang, Jiri Gesi, Ran Xue, Tian Wang, Ziyi Wang, Yuxuan Lu, Sinong Zhan, Huimin Zeng, Qingjun Cui, Yufan Guo, Jing Huang, Mubarak Shah, Dakuo Wang

机构 * Michigan State University(密歇根州立大学) Amazon(亚马逊) Northeastern University(东北大学) Northwestern University(西北大学) University of Central Florida(佛罗里达中央大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19747 2025-10-23 cs.SE 50%

Review of Tools for Zero-Code LLM Based Application Development

Priyaranjan Pattnayak, Hussain Bohra

专题命中 多模态Agent :multimodal(abstract)

Comments Accepted in 6th World Conference on Artificial Intelligence: Advances and Applications (WCAIAA 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 9 篇

2412.12718 2025-10-23 cs.CV cs.MM 88%

ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding

Zhenxing Zhang, Yaxiong Wang, Lechao Cheng, Zhun Zhong, Dan Guo, Meng Wang

机构 * School of Computer Science and Information Engineering, Hefei University of Technology, China(计算机科学与信息工程学院,合肥工业大学,中国) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center, China(人工智能研究所,合肥综合性国家科学中心,中国)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract);image-text(abstract)

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16895 2025-10-23 cs.CV cs.AI cs.LG 84%

With Limited Data for Multimodal Alignment, Let the STRUCTURE Guide You

Fabian Gröger, Shuo Wen, Huyen Le, Maria Brbić

机构 * EPFL(瑞士联邦理工学院) University of Basel(巴塞尔大学) HSLU(苏黎世联邦理工学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2025 camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19336 2025-10-23 cs.CV 83%

DaMo: Data Mixing Optimizer in Fine-tuning Multimodal LLMs for Mobile Phone Agents

Kai Shi, Jun Yang, Ni Yang, Binqiang Pan, Qingsong Xie, Chao Zhang, Zhenyu Yang, Tianhuang Su, Haonan Lu

机构 * OPPO AI Center(OPPO人工智能中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18279 2025-10-23 cs.CL cs.AI 81%

Text or Pixels? It Takes Half: On the Token Efficiency of Visual Text Inputs in Multimodal LLMs

Yanhong Li, Zixuan Lan, Jiawei Zhou

机构 * Allen Institute for AI(艾伦人工智能研究所) University of Chicago(芝加哥大学) Stony Brook University(石溪大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025 Findings ("Text or Pixels? Evaluating Efficiency and Understanding of LLMs with Visual Text Inputs")

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19215 2025-10-23 cs.CV 70%

SFGFusion: Surface Fitting Guided 3D Object Detection with 4D Radar and Camera Fusion

Xiaozhi Li, Huijun Di, Jian Li, Feng Liu, Wei Liang

机构 * Radar Technology Research Institute, School of Information and Electronics, Beijing Institute of Technology(雷达技术研究院,信息电子学院,北京理工大学) School of Computer Science and Technology, Beijing Institute of Technology(计算机科学与技术学院,北京理工大学) Innovative Equipment Research Institute, Beijing Institute of Technology(创新装备研究院,北京理工大学) Key Laboratory of Electronic and Information Technology in Satellite Navigation (Beijing Institute of Technology), Ministry of Education(卫星导航电子信息技术重点实验室(北京理工大学),教育部) Beijing Racobit Electronic Information Technology Co., Ltd.(北京瑞科比特电子信息技术有限公司)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments Submitted to Pattern Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19384 2025-10-23 cs.LG 67%

Learning Noise-Resilient and Transferable Graph-Text Alignment via Dynamic Quality Assessment

Yuhang Liu, Minglai Shao, Zengyi Wo, Yunlong Chu, Bing Hao, Shengzhong Liu, Ruijie Wang, Jianxin Li

机构 * School of New Media and Communication, Tianjin University(新媒体与传播学院,天津大学) Baidu(百度) Shanghai Jiao Tong University(上海交通大学) School of Computer Science and Engineering, Beihang University(计算机科学与工程学院,北航)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19520 2025-10-23 cs.MM 57%

CDI-DTI: A Strong Cross-domain Interpretable Drug-Target Interaction Prediction Framework Based on Multi-Strategy Fusion

Xiangyu Li, Haojie Yang, Kaimiao Hu, Runzhi Wu, Liangliang Liu, Ran Su

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19484 2025-10-23 q-bio.BM cs.AI cs.LG 57%

KnowMol: Advancing Molecular Large Language Models with Multi-Level Chemical Knowledge

Zaifei Yang, Hong Chang, Ruibing Hou, Shiguang Shan, Xilin Chen

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, CAS, China(人工智能安全国家重点实验室,计算技术研究所,中国科学院,中国) University of Chinese Academy of Sciences (CAS), China(中国科学院大学(中国科学院))

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19078 2025-10-23 cs.CV 57%

UniHPR: Unified Human Pose Representation via Singular Value Contrastive Learning

Zhongyu Jiang, Wenhao Chai, Lei Li, Zhuoran Zhou, Cheng-Yen Yang, Jenq-Neng Hwang

机构 * University of Washington(华盛顿大学) University of Copenhagen(哥本哈根大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 其他多模态 2 篇

2503.07663 2025-10-23 cs.LG cs.AI 83%

Merge then Realign: Simple and Effective Modality-Incremental Continual Learning for Multimodal LLMs

Dingkun Zhang, Shuhan Qi, Xinyu Xiao, Kehai Chen, Xuan Wang

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Guangdong Provincial Key Laboratory of Novel Security Intelligence Technologies(广东省新型安全智能技术重点实验室)

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19684 2025-10-23 quant-ph physics.app-ph 50%

Addressing spins at the clock transitions with a frequency- and bandwidth-tunable superconducting resonator

Yutian Wen, V. Ranjan, T. Lorriaux, D. Vion, B. Huard, A. Bienfait, E. Flurin, P. Bertet

专题命中 其他多模态 :multimodal(abstract)

Comments 11 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏