arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-14 至 2025-10-14 共收录 120 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 11 篇

2504.09282 2025-10-14 cs.CV 70%

VideoAds for Fast-Paced Video Understanding

Zheyuan Zhang, Monica Dou, Linkai Peng, Hongyi Pan, Ulas Bagci, Boqing Gong

机构 * Northwestern University(西北大学) Boston University(波士顿大学)

专题命中 视频多模态 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

Comments ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11063 2025-10-14 cs.CV 57%

LSVOS 2025 Challenge Report: Recent Advances in Complex Video Object Segmentation

Chang Liu, Henghui Ding, Kaining Ying, Lingyi Hong, Ning Xu, Linjie Yang, Yuchen Fan, Mingqi Gao, Jingkun Chen, Yunqi Miao, Gengshen Wu, Zhijin Qin, Jungong Han, Zhixiong Zhang, Shuangrui Ding, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Jiaqi Wang, Chang Soo Lim, Joonyoung Moon, Donghyeon Cho, Tingmin Li, Yixuan Li, Yang Yang, An Yan, Leilei Cao, Feng Lu, Ran Hong, Youhai Jiang, Fengjie Zhu, Yujie Xie, Hongyang Zhang, Zhihui Liu, Shihai Ruan, Quanzhu Niu, Dengxian Gong, Shihao Chen, Tao Zhang, Yikang Zhou, Haobo Yuan, Lu Qi, Xiangtai Li, Shunping Ji, Ran Hong, Feng Lu, Leilei Cao, An Yan, Alexey Nekrasov, Ali Athar, Daan de Geus, Alexander Hermans, Bastian Leibe

专题命中 视频多模态 :MLLM(abstract);分类 cs.CV

Comments 16 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10976 2025-10-14 cs.AI 57%

Video-STR: Reinforcing MLLMs in Video Spatio-Temporal Reasoning with Relation Graph

Wentao Wang, Heqing Zou, Tianze Luo, Rui Huang, Yutian Zhao, Zhuochen Wang, Hansheng Zhang, Chengwei Qin, Yan Wang, Lin Zhao, Huaijian Zhang

机构 * ByteDance(字节跳动) NUS(国立大学新加坡) HKUST(GZ)(香港科技大学(珠海)) THU(清华大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10422 2025-10-14 cs.CV 57%

Towards Cybersickness Severity Classification from VR Gameplay Videos Using Transfer Learning and Temporal Modeling

Jyotirmay Nag Setu, Kevin Desai, John Quarles

机构 * The University of Texas at San Antonio(德克萨斯大学圣安东尼奥分校)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10022 2025-10-14 cs.CV 57%

Q-Adapter: Visual Query Adapter for Extracting Textually-related Features in Video Captioning

Junan Chen, Trung Thanh Nguyen, Takahiro Komamizu, Ichiro Ide

机构 * Nagoya University(名古屋大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments ACM Multimedia Asia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09981 2025-10-14 cs.CV eess.IV 57%

Scaling Traffic Insights with AI and Language Model-Powered Camera Systems for Data-Driven Transportation Decision Making

Fan Zuo, Donglin Zhou, Jingqin Gao, Kaan Ozbay

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10392 2025-10-14 cs.RO cs.SY eess.SY 50%

MicroRoboScope: A Portable and Integrated Mechatronic Platform for Magnetic and Acoustic Microrobotic Experimentation

Max Sokolich, Yanda Yang, Subrahmanyam Cherukumilli, Fatma Ceren Kirmizitas, Sambeeta Das

机构 * Department of Mechanical Engineering, University of Delaware(机械工程系,德雷克塞尔大学) Departments of Animal & Food Sciences, Biological Sciences, and Medical & Molecular Sciences, University of Delaware(动物与食品科学系、生物科学系和医学与分子科学系,德雷克塞尔大学)

专题命中 视频多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 跨模态检索 7 篇

2510.10828 2025-10-14 cs.IR cs.AI 79%

VeritasFi: An Adaptable, Multi-tiered RAG Framework for Multi-modal Financial Question Answering

Zhenghan Tai, Hanwei Wu, Qingchen Hu, Jijun Chi, Hailin He, Lei Ding, Tung Sum Thomas Kwok, Bohuai Xiao, Yuchen Hua, Suyuchen Wang, Peng Lu, Muzhi Li, Yihong Wu, Liheng Ma, Jerry Huang, Jiayi Zhang, Gonghao Zhang, Chaolong Jiang, Jingrui Tian, Sicheng Lyu, Zeyu Li, Boyu Han, Fengran Mo, Xinyue Yu, Yufei Cui, Ling Zhou, Xinyu Wang

机构 * University of Toronto(多伦多大学) McMaster University(麦马斯特大学) McGill University(麦吉尔大学) University of Manitoba(曼尼托巴大学) University of California, Los Angeles(加州大学洛杉矶分校) University of Montreal(蒙特利尔大学) Mila CUHK(香港中文大学) HKUST(GZ)(香港理工大学(广州)) Nanyang Technological University(南洋理工大学) Stanford University(斯坦福大学) CG Matrix Technology Limited(CG矩阵科技有限公司)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21524 2025-10-14 cs.CV cs.LG stat.ML 70%

Learning Shared Representations from Unpaired Data

Amitai Yacobi, Nir Ben-Ari, Ronen Talmon, Uri Shaham

机构 * Department of Computer Science Bar-Ilan University(巴伊兰大学计算机科学系) Electrical and Computer Engineering Technion(技术学院电子与计算机工程系)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10426 2025-10-14 cs.CV cs.AI 62%

Taming a Retrieval Framework to Read Images in Humanlike Manner for Augmenting Generation of MLLMs

Suyang Xi, Chenxi Yang, Hong Ding, Yiqing Ni, Catherine C. Liu, Yunhao Liu, Chengqi Zhang

机构 * Emory University(埃默里大学) University of Electronic Science and Technology of China(电子科技大学) University of Illinois Chicago(伊利诺伊大学香槟分校) The Hong Kong Polytechnic University(香港理工大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11204 2025-10-14 cs.CV 57%

Class Prototypes based Contrastive Learning for Classifying Multi-Label and Fine-Grained Educational Videos

Rohit Gupta, Anirban Roy, Claire Christensen, Sujeong Kim, Sarah Gerard, Madeline Cincebeaux, Ajay Divakaran, Todd Grindal, Mubarak Shah

机构 * Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,中央佛罗里达大学) SRI International(SRI国际)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Published at CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10787 2025-10-14 cs.CL 57%

Review of Inference-Time Scaling Strategies: Reasoning, Search and RAG

Zhichao Wang, Cheng Wan, Dong Nie

机构 * Inflection AI Georgia Institute of Technology(佐治亚理工学院) ChatAlpha AI

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05970 2025-10-14 cs.CV 57%

Automatic Synthesis of High-Quality Triplet Data for Composed Image Retrieval

Haiwen Li, Delong Liu, Zhaohui Hou, Zhicheng Zhao, Fei Su

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments This paper was originally submitted to ACM MM 2025 on April 12, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10655 2025-10-14 q-bio.OT 56%

Isotropy and Geometry of Pretrained Protein LMs

Sheikh Azizul Hakim, Kowshic Roy, M Saifur Rahman

专题命中 跨模态检索 :multi-modal(abstract,comments)

Comments Published in the Proceedings of the ICML 2025 Workshop on Multi-modal Foun- dation Models and Large Language Models for Life Sciences, Vancouver, Canada. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态生成 16 篇

2510.11096 2025-10-14 cs.CV 83%

CoDefend: Cross-Modal Collaborative Defense via Diffusion Purification and Prompt Optimization

Fengling Zhu, Boshi Liu, Jingyu Hua, Sheng Zhong

专题命中 多模态生成 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10037 2025-10-14 cs.CE 82%

Automated Glaucoma Report Generation via Dual-Attention Semantic Parallel-LSTM and Multimodal Clinical Data Integration

Cheng Huang, Weizheng Xie, Zeyu Han, Tsengdar Lee, Karanjit Kooner, Jui-Ka Wang, Ning Zhang, Jia Zhang

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)

Comments Accepted by IEEE 25th BIBE

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09479 2025-10-14 cs.AI cs.CL 81%

Draw with Thought: Unleashing Multimodal Reasoning for Scientific Diagram Generation

Zhiqing Cui, Jiahao Yuan, Hanqing Wang, Yanshu Li, Chenxu Du, Zhenglong Ding

机构 * Nanjing University of Information Science \& Technology Nanjing China East China Normal University Shanghai China The Hong Kong University of Science Brown University Providence America Southwest Jiaotong University Chengdu China Nanjing University of Information Science \& Technology East China Normal University Brown University Southwest Jiaotong University

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 10 pages, 5 figures, accepted to appear in the Proceedings of the 33rd ACM International Conference on Multimedia (MM '25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10914 2025-10-14 eess.SY cs.SY 78%

Optimal Multi-Modal Transportation and Electric Power Flow: The Value of Coordinated Dynamic Operation

Jiajie Qiu, Dakota Thompson, Kamal Youcef-Toumi, Amro M. Farid

专题命中 多模态生成 :multi-modal(title,abstract)

Comments 31 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10633 2025-10-14 cs.AI 70%

Collaborative Text-to-Image Generation via Multi-Agent Reinforcement Learning and Semantic Fusion

Jiabao Shi, Minfeng Qi, Lefeng Zhang, Di Wang, Yingjie Zhao, Ziying Li, Yalong Xing, Ningran Li

机构 * Minzu University of China(民族大学) City University of Macau(澳门城市大学) Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences)(计算能力网络与信息安全重点实验室,教育部,山东计算机科学中心(济南国家超级计算机中心),齐鲁工业大学(山东省科学院)) The University of Adelaide(阿德莱德大学)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

Comments 16 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07249 2025-10-14 cs.CV 70%

TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation

Jiaben Chen, Zixin Wang, Ailing Zeng, Yang Fu, Xueyang Yu, Siyuan Cen, Julian Tanke, Yihang Chen, Koichi Saito, Yuki Mitsufuji, Chuang Gan

机构 * UMass Amherst(马萨诸塞大学阿姆赫斯特分校) Sony AI(索尼人工智能) UC San Diego(加州大学圣地亚哥分校)

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments Project page: https://talkcuts.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10637 2025-10-14 cs.RO 67%

High-Fidelity Simulated Data Generation for Real-World Zero-Shot Robotic Manipulation Learning with Gaussian Splatting

Haoyu Zhao, Cheng Zeng, Linghao Zhuang, Yaxi Zhao, Shengke Xue, Hao Wang, Xingyue Zhao, Zhongyu Li, Kehan Li, Siteng Huang, Mingxiu Chen, Xin Li, Deli Zhao, Hua Zou

机构 * Wuhan University(武汉大学) DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团) Hupan Lab(虎扑实验室) The Chinese University of Hong Kong(香港中文大学) Tsinghua University(清华大学) Huazhong University of Science and Technology(华中科技大学) Zhejiang University(浙江大学)

专题命中 多模态生成 :multi-modal(abstract);MLLM(abstract)

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09757 2025-10-14 q-bio.OT 67%

A path towards AI-scale, interoperable biological data

Brian Aevermann, Andrea Califano, Chi-Li Chiu, Nathan Clack, William M. Clemons, Jonah Cool Florence D. D'Orazi, Joseph L. DeRisi, Joshua E. Elias, Elizabeth Fahsbender, Scott E. Fraser, Carlos G. Gonzalez, Matthias Haury, Theofanis Karaletsos, Shana O. Kelley, Aly A. Khan, Alan R. Lowe, Emma Lundberg, Ryan A. McClure, Stephani Otte, Evan O. Paull, Loïc A. Royer, Dana Sadgat, Sandra L. Schmid, Samantha Scovanner, Cathy Stolitzka, Jason R. Swedlow, Joan Wong, Garabet Yeretssian, Patricia Brennan, Ambrose J. Carr

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract)

Comments 8 pages, 2 images

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08980 2025-10-14 cs.LG cs.AI cs.CV 62%

Learning Diffusion Models with Flexible Representation Guidance

Chenyu Wang, Cai Zhou, Sharut Gupta, Zongyu Lin, Stefanie Jegelka, Stephen Bates, Tommi Jaakkola

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2025; Also Oral at ICML 2025 FM4LS workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11650 2025-10-14 cs.CV 57%

InfiniHuman: Infinite 3D Human Creation with Precise Control

Yuxuan Xue, Xianghui Xie, Margaret Kostyrko, Gerard Pons-Moll

机构 * University of Tübingen, Tübingen AI Center(图宾根大学,图宾根人工智能中心) University of Tübingen, Tübingen AI Center, MPI for Informatics, SIC(图宾根大学,图宾根人工智能中心,马克斯·普朗克信息研究所,SIC) University of Tübingen(图宾根大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Accepted to ACM SIGGRAPH Asia 2025. Project website: https://yuxuan-xue.com/infini-human

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10158 2025-10-14 cs.NI cs.AI 57%

Multi-Scale Diffusion Transformer for Jointly Simulating User Mobility and Mobile Traffic Pattern

Ziyi Liu, Qingyue Long, Zhiwen Xue, Huandong Wang, Yong Li

机构 * Department of Electronic Engineering, Beijing National Research Center for Information Science and Technology (BNRist), Tsinghua University(电子工程系,信息科学与技术国家研究中心(BNRist),清华大学) International School, Beijing University of Posts and Telecommunications(国际学院,北京邮电大学)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.AI

Comments 9 pages, 4 figures. Code: https://github.com/tsinghua-fib-lab/MSTDiff

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10156 2025-10-14 cs.CV 57%

ReMix: Towards a Unified View of Consistent Character Generation and Editing

Benjia Zhou, Bin Fu, Pei Cheng, Yanru Wang, Jiayuan Fan, Tao Chen

机构 * Tencent GYLab(腾讯GY实验室) Fudan University(复旦大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11346 2025-10-14 cs.CV cs.AI 54%

Uncertainty-Aware ControlNet: Bridging Domain Gaps with Synthetic Image Generation

Joshua Niemeijer, Jan Ehrhardt, Heinz Handels, Hristina Uzunova

机构 * German Aerospace Center (DLR)(德国航空航天中心) University of Lübeck(吕贝克大学) German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)

专题命中 多模态生成 :分类 cs.CV、cs.AI;multimodal(comments);multimodal foundation model(comments)

Comments Accepted for presentation at ICCV Workshops 2025, "The 4th Workshop on What is Next in Multimodal Foundation Models?" (MMFM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11138 2025-10-14 cs.SE 50%

What Slows Down FMware Development? An Empirical Study of Developer Challenges and Resolution Times

Zitao Wang, Zhimin Zhao, Michael W. Godfrey

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15567 2025-10-14 cs.LG 50%

Towards Unified and Lossless Latent Space for 3D Molecular Latent Diffusion Modeling

Yanchen Luo, Zhiyuan Liu, Yi Zhao, Sihang Li, Hengxing Cai, Kenji Kawaguchi, Tat-Seng Chua, Yang Zhang, Xiang Wang

机构 * University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学) DP Technology(DP技术)

专题命中 多模态生成 :multi-modal(abstract)

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15775 2025-10-14 cs.RO 50%

Humanoid Robots and Humanoid AI: Review, Perspectives and Directions

Longbing Cao

机构 * Frontier AI Research Centre, Macquarie University(前沿人工智能研究中心,麦考瑞大学)

专题命中 多模态生成 :multimodal(abstract)

Comments 35 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏