arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4895 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4895 篇

2507.11261 2025-07-29 cs.CV 57%

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition

Ronggang Huang, Haoxin Yang, Yan Cai, Xuemiao Xu, Huaidong Zhang, Shengfeng He

机构 * South China University of Technology(华南理工大学) Guangdong Engineering Center for Large Model and GenAI Technology(广东省大模型与生成式人工智能技术工程中心) State Key Laboratory of Subtropical Building and Urban Science(亚热带建筑科学国家重点实验室) Ministry of Education Key Laboratory of Big Data and Intelligent Robot(教育部大数据与智能机器人重点实验室) Guangdong Provincial Key Lab of Computational Intelligence and Cyberspace Information(广东省计算智能与网络信息重点实验室) Singapore Management University(新加坡国立大学)

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07994 2025-07-29 cs.CV 57%

Doodle Your Keypoints: Sketch-Based Few-Shot Keypoint Detection

Subhajit Maity, Ayan Kumar Bhunia, Subhadeep Koley, Pinaki Nath Chowdhury, Aneeshan Sain, Yi-Zhe Song

机构 * Department of Computer Science, University of Central Florida(中央佛罗里达大学计算机科学系) SketchX, CVSSP, University of Surrey(SketchX、CVSSP、塞雷尔大学)

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted at ICCV 2025. Project Page: https://subhajitmaity.me/DYKp

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16815 2025-07-29 cs.CV 57%

FREE-Merging: Fourier Transform for Efficient Model Merging

Shenghe Zheng, Hongzhi Wang

机构 * Harbin Institute of Technology(哈尔滨工业大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.09880 2025-07-28 cs.CV 57%

Information Extraction from Unstructured data using Augmented-AI and Computer Vision

Aditya Parikh

机构 * Department of Electronics and Telecommunication Engineering(电子与电信工程系) Vishwakarma Institute of Technology(维斯瓦克arma技术学院)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00286 2025-07-22 cs.HC cs.AI cs.ET 57%

"Before, I Asked My Mom, Now I Ask ChatGPT": Visual Privacy Management with Generative AI for Blind and Low-Vision People

Tanusree Sharma, Yu-Yun Tseng, Lotus Zhang, Ayae Ide, Kelly Avery Mack, Leah Findlater, Danna Gurari, Yang Wang

机构 * Pennsylvania State University(宾夕法尼亚州立大学) Computer Science, University of Colorado(计算机科学,科罗拉多大学) Human Centered Design and Engineering, University of Washington(以人为核心的设计与工程,华盛顿大学) Information Sciences, University of Illinois at Urbana-Champaign(信息科学,伊利诺伊大学厄巴纳-香槟分校)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13350 2025-07-18 cs.CV cs.LG 57%

Hierarchical Rectified Flow Matching with Mini-Batch Couplings

Yichi Zhang, Yici Yan, Alex Schwing, Zhizhen Zhao

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

Comments Project Page: https://riccizz.github.io/HRF_coupling

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03217 2025-07-16 eess.IV cs.CV 57%

petBrain: A New Pipeline for Amyloid, Tau Tangles and Neurodegeneration Quantification Using PET and MRI

Pierrick Coupé, Boris Mansencal, Floréal Morandat, Sergio Morell-Ortega, Nicolas Villain, Jose V. Manjón, Vincent Planche

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10547 2025-07-15 cs.CV cs.LG 57%

Quantize-then-Rectify: Efficient VQ-VAE Training

Borui Zhang, Qihang Rao, Wenzhao Zheng, Jie Zhou, Jiwen Lu

机构 * Department of Automation, Tsinghua University(自动化系,清华大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13918 2025-07-09 cs.CV 57%

Quantization without Tears

Minghao Fu, Hao Yu, Jie Shao, Junjie Zhou, Ke Zhu, Jianxin Wu

机构 * National Key Laboratory for Novel Software Technology, Nanjing University, China(国家新型软件技术重点实验室,南京大学,中国) School of Artificial Intelligence, Nanjing University, China(人工智能学院,南京大学,中国)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments CVPR 2025. The code is publicly available at https://github.com/wujx2001/QwT

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02914 2025-07-08 cs.HC cs.AI cs.LG 57%

OAK -- Onboarding with Actionable Knowledge

Steve Devènes, Marine Capallera, Robin Cherix, Elena Mugellini, Omar Abou Khaled, Francesco Carrino

机构 * Institute of Systems Engineering, HEI-VS HES-SO University of Applied Sciences and Arts Western Switzerland(系统工程研究所,HEI-VS HES-SO西部瑞士应用科学与艺术大学) HumanTech Institute, HEIA HES-SO University of Applied Sciences and Arts Western Switzerland(HumanTech研究所,HEIA HES-SO西部瑞士应用科学与艺术大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

Comments This paper is an extended version of the work originally presented at the AI-Days 2024 conference in Lausanne, Switzerland. It builds upon the findings shared during the conference and includes additional results and analysis

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01254 2025-07-03 cs.CV 57%

Robust Brain Tumor Segmentation with Incomplete MRI Modalities Using Hölder Divergence and Mutual Information-Enhanced Knowledge Transfer

Runze Cheng, Xihang Qiu, Ming Li, Ye Zhang, Chun Li, Fei Yu

机构 * MSU-BIT-SMBU Joint Research Center of Applied Mathematics, Shenzhen MSU-BIT University(深圳MSU-BIT大学应用数学联合研究中心) Institute of Control Theory and Control Engineering, School of Automation, Beijing Institute of Technology(北京理工大学控制理论与控制工程研究所) School of Mathematics and Statistics, Beijing Institute of Technology(北京理工大学数学与统计学学院) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济实验室(深圳))

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00087 2025-07-02 cs.LG cs.AI 57%

pUniFind: a unified large pre-trained deep learning model pushing the limit of mass spectra interpretation

Jiale Zhao, Pengzhi Mao, Kaifei Wang, Yiming Li, Yaping Peng, Ranfei Chen, Shuqi Lu, Xiaohong Ji, Jiaxiang Ding, Xin Zhang, Yucheng Liao, Weinan E, Weijie Zhang, Han Wen, Hao Chi

机构 * Key Laboratory of Intelligent Information Processing of Chinese Academy of Sciences (CAS)(中国科学院智能信息处理重点实验室) Institute of Computing Technology(计算技术研究所) DP Technology Co., Ltd.(DP科技有限公司) University of Chinese Academy of Sciences(中国科学院大学) AI for Science Institute(AI for Science研究院) Center for Machine Learning Research(机器学习研究中心) School of Mathematical Sciences(数学科学学院) State Key Laboratory of Medical Proteomics(医学蛋白质组学国家重点实验室)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23944 2025-07-02 cs.RO cs.AI 57%

Adapt Your Body: Mitigating Proprioception Shifts in Imitation Learning

Fuhang Kuang, Jiacheng You, Yingdong Hu, Tong Zhang, Chuan Wen, Yang Gao

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

Comments Need further modification

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02205 2025-07-02 cs.LG cs.AI cs.SY eess.SY 57%

Bregman Centroid Guided Cross-Entropy Method

Yuliang Gu, Hongpeng Cao, Marco Caccamo, Naira Hovakimyan

机构 * Department of Mechanical Science and Engineering, UIUC, United States(伊利诺伊大学厄巴纳-香槟分校机械科学与工程系) School of Engineering and Design, TUM, Germany(慕尼黑工业大学工程与设计学院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14153 2025-07-02 eess.IV cs.CV 57%

De-LightSAM: Modality-Decoupled Lightweight SAM for Generalizable Medical Segmentation

Qing Xu, Jiaxuan Li, Xiangjian He, Chenxin Li, Fiseha B. Tesem, Wenting Duan, Zhen Chen, Rong Qu, Jonathan M. Garibaldi, Chang Wen Chen

机构 * School of Computer Science, University of Nottingham Ningbo China(诺丁汉大学宁波校区计算机科学学院) Department of Electronic Engineering, The Chinese University of Hong Kong(香港中文大学电子工程系) School of School of Engineering & Physical Sciences, University of Lincoln(林肯大学工程与物理科学学院) Hong Kong Institute of Science & Innovation, Chinese Academy of Sciences(中国科学院香港科学与创新研究院) School of Computer Science, University of Nottingham(诺丁汉大学计算机科学学院) The Hong Kong Polytechnic University(香港理工大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22022 2025-07-01 cs.CV 57%

Advancing Facial Stylization through Semantic Preservation Constraint and Pseudo-Paired Supervision

Zhanyi Lu, Yue Zhou

机构 * School of Electrical and Electronic Engineering(电子工程学院) University of Shanghai Jiao Tong University(上海交通大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19367 2025-06-30 cs.CV 57%

VGAT: A Cancer Survival Analysis Framework Transitioning from Generative Visual Question Answering to Genomic Reconstruction

Zizhi Chen, Minghao Han, Xukun Zhang, Shuwei Ma, Tao Liu, Xing Wei, Lihua Zhang

机构 * 1 Academy for Engineering Technology, Fudan University, Shanghai, China 2 Institute of Metaverse \& Intelligent Medicine, Fudan University, Shanghai, China 3 Engineering Research Center of AI 4 Jilin Provincial Key Laboratory of Intelligence Science

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by ICME2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08165 2025-06-27 cs.LG cs.CV 57%

Chain-of-Sketch: Enabling Global Visual Reasoning

Aryo Lotfi, Enrico Fini, Samy Bengio, Moin Nabi, Emmanuel Abbe

机构 * Apple(苹果公司)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

Comments additional experiments added, title changed from "Visual Scratchpads: Enabling Global Reasoning in Vision" to "Chain-of-Sketch: Enabling Global Visual Reasoning"

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19884 2025-06-26 cs.OS cs.AI cs.PF cs.SE 57%

MNN-AECS: Energy Optimization for LLM Decoding on Mobile Devices via Adaptive Core Selection

Zhengxiang Huang, Chaoyue Niu, Zhaode Wang, Jiarui Xue, Hanming Zhang, Yugang Wang, Zewei Xin, Xiaotang Jiang, Chengfei Lv, Fan Wu, Guihai Chen

机构 * Shanghai Jiao Tong University(上海交通大学) Alibaba Group(阿里巴巴集团)

专题命中 其他多模态 :MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07826 2025-06-23 cs.CV eess.IV 57%

Deep Learning in Automated Power Line Inspection: A Review

Md. Ahasan Atick Faisal, Imene Mecheter, Yazan Qiblawey, Javier Hernandez Fernandez, Muhammad E. H. Chowdhury, Serkan Kiranyaz

机构 * Department of Electrical Engineering, Qatar University(卡塔尔大学电气工程系)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

Comments 40 pages, 12 figures

Journal ref Applied Energy. 385 (2025) 125507

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15279 2025-06-19 cs.CV 57%

BCRNet: Enhancing Landmark Detection in Laparoscopic Liver Surgery via Bezier Curve Refinement

Qian Li, Feng Liu, Shuojue Yang, Daiyun Shen, Yueming Jin

机构 * National University of Singapore, Singapore(新加坡国立大学) Harbin Institute of Technology, Harbin, China(哈尔滨工业大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted at MICCAI 2025, 11 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.05249 2025-06-19 cs.LG cs.AI cs.CE eess.IV 57%

Advancing oncology with federated learning: transcending boundaries in breast, lung, and prostate cancer. A systematic review

Anshu Ankolekar, Sebastian Boie, Maryam Abdollahyan, Emanuela Gadaleta, Seyed Alireza Hasheminasab, Guang Yang, Charles Beauville, Nikolaos Dikaios, George Anthony Kastis, Michael Bussmann, Sara Khalid, Hagen Kruger, Philippe Lambin, Giorgos Papanastasiou

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

Comments 5 Figures, 3 Tables, 1 Supplementary Table

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13306 2025-06-17 eess.IV cs.CV 57%

Brain Imaging Foundation Models, Are We There Yet? A Systematic Review of Foundation Models for Brain Imaging and Biomedical Research

Salah Ghamizi, Georgia Kanli, Yu Deng, Magali Perquin, Olivier Keunen

机构 * Luxembourg Institute of Health (LIH)(卢森堡健康研究所) King's College London, UK(伦敦国王学院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12826 2025-06-17 cs.CV 57%

LOP: Learning Optimal Pruning for Efficient On-Demand MLLMs Scaling

Zhihan Zhang, Xiang Pan, Hongchen Wei, Zhenzhong Chen

机构 * School of Remote Sensing and Information Engineering, Wuhan University(武汉大学遥感与信息工程学院) School of Data Science, Lingnan University(岭南大学数据科学学院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08970 2025-06-16 cs.CR cs.AI cs.LG 57%

Self-interpreting Adversarial Images

Tingwei Zhang, Collin Zhang, John X. Morris, Eugene Bagdasarian, Vitaly Shmatikov

机构 * Cornell Tech(康奈尔科技) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

专题命中 其他多模态 :cross-modal(abstract);分类 cs.AI

Comments in USENIX Security 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19407 2025-06-16 cs.CV 57%

YOLO advances to its genesis: a decadal and comprehensive review of the You Only Look Once (YOLO) series

Ranjan Sapkota, Marco Flores Calero, Rizwan Qureshi, Chetan Badgujar, Upesh Nepal, Alwin Poulose, Peter Zeno, Uday Bhanu Prakash Vaddevolu, Sheheryar Khan, Maged Shoman, Hong Yan, Manoj Karkee

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments Published in Artificial Intelligence Review as https://doi.org/10.1007/s10462-025-11253-3

Journal ref Artificial Intelligence Review, SpringerNature, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10910 2025-06-13 cs.CL 57%

Magistral

Mistral-AI, :, Abhinav Rastogi, Albert Q. Jiang, Andy Lo, Gabrielle Berrada, Guillaume Lample, Jason Rute, Joep Barmentlo, Karmesh Yadav, Kartik Khandelwal, Khyathi Raghavi Chandu, Léonard Blier, Lucile Saulnier, Matthieu Dinot, Maxime Darrin, Neha Gupta, Roman Soletskyi, Sagar Vaze, Teven Le Scao, Yihan Wang, Adam Yang, Alexander H. Liu, Alexandre Sablayrolles, Amélie Héliou, Amélie Martin, Andy Ehrenberg, Anmol Agarwal, Antoine Roux, Arthur Darcet, Arthur Mensch, Baptiste Bout, Baptiste Rozière, Baudouin De Monicault, Chris Bamford, Christian Wallenwein, Christophe Renaudin, Clémence Lanfranchi, Darius Dabert, Devon Mizelle, Diego de las Casas, Elliot Chane-Sane, Emilien Fugier, Emma Bou Hanna, Gauthier Delerce, Gauthier Guinet, Georgii Novikov, Guillaume Martin, Himanshu Jaju, Jan Ludziejewski, Jean-Hadrien Chabran, Jean-Malo Delignon, Joachim Studnia, Jonas Amar, Josselin Somerville Roberts, Julien Denize, Karan Saxena, Kush Jain, Lingxiao Zhao, Louis Martin, Luyu Gao, Lélio Renard Lavaud, Marie Pellat, Mathilde Guillaumin, Mathis Felardos, Maximilian Augustin, Mickaël Seznec, Nikhil Raghuraman, Olivier Duchenne, Patricia Wang, Patrick von Platen, Patryk Saffer, Paul Jacob, Paul Wambergue, Paula Kurylowicz, Pavankumar Reddy Muddireddy, Philomène Chagniot, Pierre Stock, Pravesh Agrawal, Romain Sauvestre, Rémi Delacourt, Sanchit Gandhi, Sandeep Subramanian, Shashwat Dalal, Siddharth Gandhi, Soham Ghosh, Srijan Mishra, Sumukh Aithal, Szymon Antoniak, Thibault Schueller, Thibaut Lavril, Thomas Robert, Thomas Wang, Timothée Lacroix, Valeriia Nemychnikova, Victor Paltz, Virgile Richard, Wen-Ding Li, William Marshall, Xuanyu Zhang, Yunhao Tang

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09409 2025-06-10 cs.RO cs.AI cs.CE cs.LG 57%

AI-based Framework for Robust Model-Based Connector Mating in Robotic Wire Harness Installation

Claudius Kienle, Benjamin Alt, Finn Schneider, Tobias Pertlwieser, Rainer Jäkel, Rania Rayyes

机构 * IAS Lab, Computer Science Department, TU Darmstadt(德累斯顿技术大学计算机科学系IAS实验室) AICOR Institute for Artificial Intelligence, University of Bremen(不莱梅大学人工智能研究所) Institute for Material Handling and Logistics (IFL), Karlsruhe Institute of Technology (KIT)(卡尔斯鲁厄理工学院物流与物料搬运研究所) ArtiMinds Robotics(ArtiMinds机器人技术公司)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments 6 pages, 6 figures, 4 tables, presented at the 2025 IEEE 21st International Conference on Automation Science and Engineering (CASE 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16983 2025-05-30 cs.CL 57%

LLM as Effective Streaming Processor: Bridging Streaming-Batch Mismatches with Group Position Encoding

Junlong Tong, Jinlan Fu, Zixuan Lin, Yingqi Fan, Anhao Zhao, Hui Su, Xiaoyu Shen

机构 * Shanghai Jiao Tong University(上海交通大学) Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生关键实验室) Institute of Digital Twin, EIT(数字孪生研究所) National University of Singapore(新加坡国立大学) University of Science and Technology of China(中国科学技术大学) Meituan Inc.(美团公司)

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CL

Comments ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22408 2025-05-29 cs.CV 57%

Frugal Incremental Generative Modeling using Variational Autoencoders

Victor Enescu, Hichem Sahbi

机构 * Sorbonne University, CNRS, LIP6(索邦大学、国家科学研究中心、LIP6)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏