arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-08 至 2025-10-08 共收录 46 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4 篇

2508.07216 2025-10-08 cs.CV 83%

Bridging Semantic Logic Gaps: A Cognition Inspired Multimodal Boundary Preserving Network for Image Manipulation Localization

Songlin Li, Zhiqing Guo, Yuanman Li, Zeyu Li, Yunfeng Diao, Gaobo Yang, Liejun Wang

机构 * School of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院) School of Electronic and Information Engineering, Shenzhen University(深圳大学电子与信息工程学院) College of Computer Science and Electronic Engineering, Hunan University(湖南大学计算机科学与电子工程学院) School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23250 2025-10-08 cs.AI cs.CV 81%

Training Vision-Language Process Reward Models for Test-Time Scaling in Multimodal Reasoning: Key Insights and Lessons Learned

Brandon Ong, Tej Deep Pala, Vernon Toh, William Chandra Tjhi, Soujanya Poria

机构 * AI Singapore(AI新加坡) Nanyang Technological University(南洋理工大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22688 2025-10-08 cs.CV 57%

Robust Object Detection for Autonomous Driving via Curriculum-Guided Group Relative Policy Optimization

Xu Jia

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17132 2025-10-08 cs.AI 57%

Applications of Large Models in Medicine

YunHe Su, Zhengyang Lu, Junhui Liu, Ke Pang, Haoran Dai, Sa Liu, Yuxin Jia, Lujia Ge, Jing-min Yang

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 3 篇

2510.06060 2025-10-08 cs.MM cs.AI cs.CV 82%

Controllable Audio-Visual Viewpoint Generation from 360° Spatial Information

Christian Marinoni, Riccardo Fosco Gramaccioni, Eleonora Grassucci, Danilo Comminiello

机构 * Sapienza University of Rome, Italy(罗马大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05417 2025-10-08 cs.HC cs.AI 79%

Exploring Student Choice and the Use of Multimodal Generative AI in Programming Learning

Xinying Hou, Ruiwei Xiao, Runlong Ye, Michael Liut, John Stamper

机构 * University of Michigan(密歇根大学) Carnegie Mellon University(卡内基梅隆大学) University of Toronto(多伦多大学) University of Toronto Mississauga(多伦多大学滑铁卢分校)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments 7 pages, accepted to SIGCSE2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05249 2025-10-08 cs.HC 50%

CLAd-VR: Cognitive Load-based Adaptive Training for Machining Tasks in Virtual Reality

Bhavya Matam, Adamay Mann, Kachina Studer, Christian Gabbianelli, Sonia Castelo, John Liu, Claudio Silva, Dishita Turakhia

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 7 篇

2510.05829 2025-10-08 cs.SD cs.CV cs.LG cs.MM eess.AS 82%

FoleyGRAM: Video-to-Audio Generation with GRAM-Aligned Multimodal Encoders

Riccardo Fosco Gramaccioni, Christian Marinoni, Eleonora Grassucci, Giordano Cicchetti, Aurelio Uncini, Danilo Comminiello

机构 * Dept. Information Engineering, Electronics and Telecommunications (DIET), Sapienza University of Rome(信息工程、电子与电信系(DIET),罗马萨皮恩扎大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM、eess.AS

Comments Acepted at IJCNN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15192 2025-10-08 cs.CV 79%

Leveraging Foundation Models for Multimodal Graph-Based Action Recognition

Fatemeh Ziaeetabar, Florentin Wörgötter

机构 * School of Mathematics, Statistics and Computer Science, College of Science, University of Tehran(数学、统计与计算机科学学院,科学学院,塔里斯坦大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05836 2025-10-08 cs.CV 70%

Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow

Ruyang Liu, Shangkun Sun, Haoran Tang, Ge Li, Wei Gao

机构 * School of Electronic and Computer Engineering, Shenzhen Graduate School, 2 Peng Cheng LaboratoryPeking University(1 电子与计算机工程学院,深圳研究生院,2 深圳鹏城实验室,北京大学)

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted to ICCV' 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06077 2025-10-08 cs.CV cs.AI 62%

When Thinking Drifts: Evidential Grounding for Robust Video Reasoning

Mi Luo, Zihui Xue, Alex Dimakis, Kristen Grauman

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) UC Berkeley(伯克利大学) Bespoke Labs(Bespoke实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted by NeurIPS 2025, Project page: https://vision.cs.utexas.edu/projects/video-ver/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06040 2025-10-08 cs.CV cs.AI 62%

VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-based Group Relative Policy Optimization

Xinye Cao, Hongcan Guo, Jiawen Qian, Guoshun Nan, Chao Wang, Yuqi Pan, Tianhao Hou, Xiaojuan Wang, Yutong Gao

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Minzu University of China(民族大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03501 2025-10-08 cs.CV 57%

LV-MAE: Learning Long Video Representations through Masked-Embedding Autoencoders

Ilan Naiman, Emanuel Ben-Baruch, Oron Anschel, Alon Shoshan, Igor Kviatkovsky, Manoj Aggarwal, Gerard Medioni

机构 * Amazon(亚马逊)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to the International Conference on Computer Vision, ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05533 2025-10-08 q-fin.PM 50%

The New Quant: A Survey of Large Language Models in Financial Prediction and Trading

Weilong Fu

专题命中 视频多模态 :multimodal(abstract)

Comments 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 2 篇

2510.05949 2025-10-08 cs.LG cs.AI cs.CV stat.ML 62%

Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density

Randall Balestriero, Nicolas Ballas, Mike Rabbat, Yann LeCun

机构 * Meta-FAIR Brown University(布朗大学) NYU(纽约大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05016 2025-10-08 astro-ph.IM cs.AI cs.CL 62%

Large Language Models Achieve Gold Medal Performance at the International Olympiad on Astronomy & Astrophysics (IOAA)

Lucas Carrit Delgado Pinheiro, Ziru Chen, Bruno Caixeta Piazza, Ness Shroff, Yingbin Liang, Yuan-Sen Ting, Huan Sun

机构 * The Ohio State University(俄亥俄州立大学) Universidade de São Paulo(圣保罗大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 18 pages, 6 figures, to be submitted, comments are welcome. Reproducibility details can be found at: https://github.com/OSU-NLP-Group/LLM-IOAA

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 7 篇

2510.06131 2025-10-08 cs.CV cs.AI 86%

Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation

Jiawei Mao, Yuhan Wang, Lifeng Chen, Can Zhao, Yucheng Tang, Dong Yang, Liangqiong Qu, Daguang Xu, Yuyin Zhou

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments 16 pages,6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05661 2025-10-08 cs.CV cs.MM 81%

When and How to Cut Classical Concerts? A Multimodal Automated Video Editing Approach

Daniel Gonzálbez-Biosca, Josep Cabacas-Maso, Carles Ventura, Ismael Benito-Altamirano

机构 * eHealth Center, Faculty of Computer Science, Multimedia and Telecommunications, Universitat Oberta de Catalunya(eHealth中心,计算机科学、多媒体与电信学院,开放大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21653 2025-10-08 cs.CV 70%

Think Before You Diffuse: Infusing Physical Rules into Video Diffusion

Ke Zhang, Cihan Xiao, Jiacong Xu, Yiqun Mei, Vishal M. Patel

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments 19 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05593 2025-10-08 cs.CV cs.AI cs.CL 67%

Improving Chain-of-Thought Efficiency for Autoregressive Image Generation

Zeqi Gu, Markos Georgopoulos, Xiaoliang Dai, Marjan Ghazvininejad, Chu Wang, Felix Juefei-Xu, Kunpeng Li, Yujun Shi, Zecheng He, Zijian He, Jiawei Zhou, Abe Davis, Jialiang Wang

机构 * Meta Superintelligence Labs(Meta超智能实验室) Meta FAIR Cornell University(康奈尔大学) Stony Brook University(石溪大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05976 2025-10-08 cs.CV cs.AI cs.LG 62%

Diffusion Models for Low-Light Image Enhancement: A Multi-Perspective Taxonomy and Performance Analysis

Eashan Adhikarla, Yixin Liu, Brian D. Davison

机构 * Lehigh University(莱维大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15046 2025-10-08 cs.CL cs.AI 62%

ChartCards: A Chart-Metadata Generation Framework for Multi-Task Chart Understanding

Yifan Wu, Lutao Yan, Leixian Shen, Yinan Mei, Jiannan Wang, Yuyu Luo

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Need to be revised

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03498 2025-10-08 cs.CV 57%

OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation

Han Li, Xinyu Peng, Yaoming Wang, Zelin Peng, Xin Chen, Rongxiang Weng, Jingang Wang, Xunliang Cai, Wenrui Dai, Hongkai Xiong

机构 * Meituan Inc(美团公司) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments technical report, project url:https://onecat-ai.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 12 篇

2410.19378 2025-10-08 cs.CV cs.LG eess.IV 83%

Unified Cross-Modal Medical Image Synthesis with Hierarchical Mixture of Product-of-Experts

Reuben Dorent, Nazim Haouchine, Alexandra Golby, Sarah Frisken, Tina Kapur, William Wells

机构 * Harvard Medical School and the Brigham and Women’s Hospital(哈佛医学院和布里洛妇产科医院) MIND Team, Inria Saclay, Université Paris-Saclay, Palaiseau, France and Sorbonne Université, Institut du Cerveau - Paris Brain Institute - ICM, CNRS, Inria, Inserm, AP-HP, Hôpital de la Pitié Salpêtrière(MIND团队、Inria Saclay、巴黎萨克雷大学、Palaiseau,法国以及索邦大学、巴黎脑研究所-ICM、CNRS、Inria、Inserm、AP-HP、皮蒂埃萨勒普蒂耶医院) Massachusetts Institute of Technology(麻省理工学院)

专题命中 多模态评测 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Accepted in IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07966 2025-10-08 cs.CV 83%

SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence

Ziyang Gong, Wenhao Li, Oliver Ma, Songyuan Li, Zhaokai Wang, Songyuan Li, Jiayi Ji, Xue Yang, Gen Luo, Junchi Yan, Rongrong Ji

机构 * SJTU(上海交通大学) XMU(厦门大学) Shanghai AI Lab(上海人工智能实验室) SYSU(南方科技大学) FDU(福建大学) NUS(新加坡国立大学)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22385 2025-10-08 cs.CV cs.AI cs.CL 82%

Can Video Large Multimodal Models Think Like Doubters-or Double-Down: A Study on Defeasible Video Entailment

Yue Zhang, Jilei Sun, Yunhui Guo, Vibhav Gogate

机构 * Department of Computer Science(计算机科学系) The University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05551 2025-10-08 cs.CV 79%

When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding

Yan Shu, Hangui Lin, Yexin Liu, Yan Zhang, Gangyan Zeng, Yan Li, Yu Zhou, Ser-Nam Lim, Harry Yang, Nicu Sebe

机构 * University of Trento(特伦托大学) Hong Kong University of Science and Technology(香港科学与技术大学) University of International Relations(国际关系大学) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) Nanjing University of Science and Technology(南京理工大学) VCIP & TMCC & DISSec, College of Computer Science, Nankai University(南开大学计算机学院) University of Central Florida(佛罗里达中央大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17978 2025-10-08 cs.CL 79%

AVerImaTeC: A Dataset for Automatic Verification of Image-Text Claims with Evidence from the Web

Rui Cao, Zifeng Ding, Zhijiang Guo, Michael Schlichtkrull, Andreas Vlachos

机构 * University of Cambridge(剑桥大学) Queen Mary University of London(伦敦大学玛丽女王学院)

专题命中 多模态评测 :image-text(title,abstract);分类 cs.CL

Comments accepted at NeurIPS 2025 Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00192 2025-10-08 cs.CV 70%

Safe-LLaVA: A Privacy-Preserving Vision-Language Dataset and Benchmark for Biometric Safety

Younggun Kim, Sirnam Swetha, Fazil Kagdi, Mubarak Shah

机构 * Center For Research in Computer Vision, University of Central Florida, USA(计算机视觉研究中心,中央佛罗里达大学) Department of Civil Environmental and Construction Engineering, University of Central Florida, USA(土木环境与建设工程系,中央佛罗里达大学) Department of Computer Science, University of Central Florida, USA(计算机科学系,中央佛罗里达大学)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05644 2025-10-08 cs.CL cs.AI 62%

The African Languages Lab: A Collaborative Approach to Advancing Low-Resource African NLP

Sheriff Issaka, Keyi Wang, Yinka Ajibola, Oluwatumininu Samuel-Ipaye, Zhaoyi Zhang, Nicte Aguillon Jimenez, Evans Kofi Agyei, Abraham Lin, Rohan Ramachandran, Sadick Abdul Mumin, Faith Nchifor, Mohammed Shuraim, Lieqi Liu, Erick Rosas Gonzalez, Sylvester Kpei, Jemimah Osei, Carlene Ajeneza, Persis Boateng, Prisca Adwoa Dufie Yeboah, Saadia Gabriel

机构 * University of California, Los Angeles(加州大学洛杉矶分校) Georgia Institute of Technology(佐治亚理工学院) University of Wisconsin - Madison(威斯康星大学麦迪逊分校) University of Cape Coast(科克伍德大学) Carleton University(卡尔顿大学) Stetson University(史蒂文森大学) Northwestern University in Qatar(卡塔尔西北大学) Cornell University(康奈尔大学) Soka University of America(美国立命大学) Columbia University(哥伦比亚大学)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏