arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-07 至 2025-10-07 共收录 98 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 22 篇

2505.23631 2025-10-07 cs.HC cs.AI cs.CL 62%

Human Empathy as Encoder: AI-Assisted Depression Assessment in Special Education

Boning Zhao, Xinnuo Li, Yutong Hu

机构 * Tandon School of Engineering New York University New York, USA College of Arts \& Science New York University New York, USA

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 7 pages, 6 figures, ACII 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19099 2025-10-07 cs.AI physics.ed-ph physics.pop-ph 57%

SeePhys: Does Seeing Help Thinking? -- Benchmarking Vision-Based Physics Reasoning

Kun Xiang, Heng Li, Terry Jingchen Zhang, Yinya Huang, Zirong Liu, Peixin Qu, Jixi He, Jiaqi Chen, Yu-Jie Yuan, Jianhua Han, Hang Xu, Hanhui Li, Mrinmaya Sachan, Xiaodan Liang

机构 * Sun Yat-sen University(中山大学) ETH Zurich(苏黎世联邦理工学院) Huawei Noah’s Ark Lab(华为诺亚实验室) The University of Hong Kong(香港大学) ETH AI Center(苏黎世人工智能中心)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

Comments 46 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.11334 2025-10-07 cs.AI 57%

Program Synthesis Benchmark for Visual Programming in XLogoOnline Environment

Chao Wen, Jacqueline Staub, Adish Singla

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

Comments ACL'25 paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03874 2025-10-07 cs.CV 57%

DHQA-4D: Perceptual Quality Assessment of Dynamic 4D Digital Human

Yunhao Li, Sijing Wu, Yucheng Zhu, Huiyu Duan, Zicheng Zhang, Guangtao Zhai

机构 * Institute of Image Communication and Network Engineering(图像通信与网络工程研究所) Shanghai Jiao Tong University(上海交通大学) USC-SJTU Institute of Cultural and Creative Industry(USC-SJTU文化产业与创意产业研究院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03863 2025-10-07 cs.AI cs.CR 57%

Spatial CAPTCHA: Generatively Benchmarking Spatial Reasoning for Human-Machine Differentiation

Arina Kharlamova, Bowei He, Chen Ma, Xue Liu

专题命中 多模态评测 :multi-modal(abstract);分类 cs.AI

Comments Submitted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01844 2025-10-07 cs.AI 57%

Towards Generalizable Context-aware Anomaly Detection: A Large-scale Benchmark in Cloud Environments

Xinkai Zou, Xuan Jiang, Ruikai Huang, Haoze He, Parv Kapoor, Hongrui Wu, Yibo Wang, Jian Sha, Xiongbo Shi, Zixun Huang, Jinhua Zhao

机构 * UC San Diego(加州大学圣地亚哥分校) Massachusetts Institute of Technology(麻省理工学院) Georgia Institute of Technology(佐治亚理工学院) Carnegie Mellon University(卡内基梅隆大学) Tongji University(同济大学) Tsinghua University(清华大学) University of Pennsylvania(宾夕法尼亚大学)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09255 2025-10-07 cs.CV 57%

FVQ: A Large-Scale Dataset and an LMM-based Method for Face Video Quality Assessment

Sijing Wu, Yunhao Li, Ziwen Xu, Yixuan Gao, Huiyu Duan, Wei Sun, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学) East China Normal University(华东师范大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments Accepted by ACM MM 2025. Project page: https://github.com/wsj-sjtu/FVQ

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03376 2025-10-07 cs.CV eess.IV 57%

Visual Language Model as a Judge for Object Detection in Industrial Diagrams

Sanjukta Ghosh

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments Pre-review version submitted to IEEE ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04183 2025-10-07 cs.NI 50%

Dynamic Adaptive Federated Learning for mmWave Sector Selection

Lucas Pacheco, Torsten Braun, Kaushik Chowdhury, Denis Rosário, Batool Salehi, Eduardo Cerqueira

专题命中 多模态评测 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04122 2025-10-07 cs.HC 50%

Wrist2Finger: Sensing Fingertip Force for Force-Aware Hand Interaction with a Ring-Watch Wearable

Yingjing Xiao, Zhichao Huang, Junbin Ren, Haichuan Song, Yang Gao, Yuting Bai, Zhanpeng Jin

专题命中 多模态评测 :cross-modal(abstract)

Comments 15 pages, 13 figures. Accepted at UIST 2025 (ACM Symposium on User Interface Software and Technology). Yingjing Xiao and Zhichao Huang contributed equally. Corresponding author: Yang Gao (gaoyang@cs.ecnu.edu.cn)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03815 2025-10-07 eess.SY cs.LG cs.SY eess.SP 50%

A Trustworthy Industrial Fault Diagnosis Architecture Integrating Probabilistic Models and Large Language Models

Yue wu

专题命中 多模态评测 :multimodal(abstract)

Comments 1tables,6 figs,11pages

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 11 篇

2504.09532 2025-10-07 cs.RO cs.AI 86%

Humanoid Agent via Embodied Chain-of-Action Reasoning with Multimodal Foundation Models for Zero-Shot Loco-Manipulation

Congcong Wen, Geeta Chandra Raju Bethala, Yu Hao, Niraj Pudasaini, Hao Huang, Shuaihang Yuan, Baoru Huang, Anh Nguyen, Mengyu Wang, Anthony Tzes, Yi Fang

机构 * Embodied AI and Robotics (AIR) Lab, New York University, New York, USA and NYUAD Center for Artificial Intelligence and Robotics, New York University Abu Dhabi, Abu Dhabi, UAE(纽约大学Embodied AI和机器人实验室及纽约大学阿布扎赫尔人工智能与机器人中心) Harvard AI and Robotics Lab, Harvard University, Boston, USA(哈佛大学人工智能与机器人实验室) Department of Computer Science, University College London, London, UK(伦敦大学学院计算机科学系) Department of Computer Science, University of Liverpool, UK(利物浦大学计算机科学系) NYUAD Center for Artificial Intelligence and Robotics, New York University Abu Dhabi, Abu Dhabi, UAE(纽约大学阿布扎赫尔人工智能与机器人中心)

专题命中 多模态Agent :multimodal(title,abstract);multimodal foundation model(title);分类 cs.AI

Comments website link: https://humanoid-coa.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03612 2025-10-07 cs.AI cs.CR 83%

Cross-Modal Content Optimization for Steering Web Agent Preferences

Tanqiu Jiang, Min Bai, Nikolaos Pappas, Yanjun Qi, Sandesh Swamy

机构 * Stony Brook University(石溪大学) AWS AI Labs(亚马逊人工智能实验室)

专题命中 多模态Agent :cross-modal(title,abstract);multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04791 2025-10-07 cs.SE 82%

GUISpector: An MLLM Agent Framework for Automated Verification of Natural Language Requirements in GUI Prototypes

Kristian Kolthoff, Felix Kretzer, Simone Paolo Ponzetto, Alexander Maedche, Christian Bartelt

专题命中 多模态Agent :MLLM(title,abstract);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04560 2025-10-07 cs.AI 79%

ContextNav: Towards Agentic Multimodal In-Context Learning

Honghao Fu, Yuan Ouyang, Kai-Wei Chang, Yiwei Wang, Zi Huang, Yujun Cai

机构 * The University of Queensland(昆士兰大学) Nanjing University(南京大学) University of California, Los Angeles(加州大学洛杉矶分校) University of California, Merced(加州大学默塞德分校)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07634 2025-10-07 cs.RO cs.AI cs.CV 62%

Neural Brain: A Neuroscience-inspired Framework for Embodied Agents

Jian Liu, Xiongtao Shi, Thai Duy Nguyen, Haitian Zhang, Tianxiang Zhang, Wei Sun, Yanjie Li, Athanasios V. Vasilakos, Giovanni Iacca, Arshad Ali Khan, Arvind Kumar, Jae Won Cho, Ajmal Mian, Lihua Xie, Erik Cambria, Lin Wang

机构 * School of Electrical and Electronic Engineering, Nanyang Technological University(南洋理工大学电子与电气工程学院) School of Artificial Intelligence and Robotics, Hunan University(湖南大学人工智能与机器人学院) School of Intelligence Science and Engineering, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)智能科学与工程学院) Department of Information and Communication Technology, University of Agder(阿格德大学信息与通信技术系) Department of Information Engineering and Computer Science, University of Trento(特伦特大学信息工程与计算机科学系) Elm Company(Elm公司) Division of Computational Science and Technology, KTH Royal Institute of Technology(皇家理工学院计算科学与技术系) School of Artificial Intelligence and Data Science, Sejong University(世宗大学人工智能与数据科学学院) Department of Computer Science of the University of Western Australia(西澳大学计算机科学系) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

Comments 51 pages, 17 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04039 2025-10-07 cs.CV cs.AI 62%

\textsc{GUI-Spotlight}: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding

Bin Lei, Nuo Xu, Ali Payani, Mingyi Hong, Chunhua Liao, Yu Cao, Caiwen Ding

机构 * University of Minnesota(明尼苏达大学) Cisco Research(思科研究) Lawrence Livermore National Labs(劳伦斯利弗莫尔国家实验室)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04023 2025-10-07 cs.AI cs.CL 62%

LLM-Based Data Science Agents: A Survey of Capabilities, Challenges, and Future Directions

Mizanur Rahman, Amran Bhuiyan, Mohammed Saidul Islam, Md Tahmid Rahman Laskar, Ridwan Mahbub, Ahmed Masry, Shafiq Joty, Enamul Hoque

机构 * York University(约克大学) Vector Institute for AI(人工智能矢量研究所) Dialpad Inc.(Dialpad公司) Nanyang Technological University(南洋理工大学) Salesforce AI Research(Salesforce人工智能研究)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI

Comments Survey paper; 45 data science agents; under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03829 2025-10-07 cs.NI cs.AI 57%

A4FN: an Agentic AI Architecture for Autonomous Flying Networks

André Coelho, Pedro Ribeiro, Helder Fontes, Rui Campos

机构 * FCT – Fundação para a Ciência e a Tecnologia, I.P.(葡萄牙科学与技术基金会)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments This paper has been accepted for presentation in the Auto ML for Zero-Touch Network Management Workshop (WS04-01) at the IEEE International Symposium on Personal, Indoor and Mobile Radio Communications (PIMRC) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00413 2025-10-07 cs.CV 57%

PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents

Zikang Liu, Junyi Li, Wayne Xin Zhao, Dawei Gao, Yaliang Li, Ji-rong Wen

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学全球化人工智能学院) Department of Data Science, City University of Hong Kong(香港城市大学数据科学系) Alibaba Group(阿里巴巴集团)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07778 2025-10-07 cs.CV 57%

A Neurosymbolic Agent System for Compositional Visual Reasoning

Yichang Xu, Gaowen Liu, Ramana Rao Kompella, Sihao Hu, Fatih Ilhan, Selim Furkan Tekin, Zachary Yahn, Ling Liu

机构 * Georgia Institute of Technology(佐治亚理工学院) Cisco Systems(思科系统)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20140 2025-10-07 cs.AI 57%

MAD-Sherlock: Multi-Agent Debate for Visual Misinformation Detection

Kumud Lakara, Georgia Channing, Christian Rupprecht, Juil Sock, Philip Torr, John Collomosse, Christian Schroeder de Witt

机构 * University of Oxford, Oxford, UK(牛津大学) BBC AI Research, London, UK(BBC人工智能研究) University of Surrey, Guildford, UK(萨里大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 8 篇

2507.09747 2025-10-07 cs.NE 82%

BrainFLORA: Uncovering Brain Concept Representation via Multimodal Neural Embeddings

Dongyang Li, Haoyang Qin, Mingyang Wu, Chen Wei, Quanying Liu

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

Comments ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22962 2025-10-07 cs.LG cond-mat.mtrl-sci physics.chem-ph 78%

Multimodal machine learning with large language embedding model for polymer property prediction

Tianren Zhang, Dai-Bei Yang

机构 * Department of Materials Science and Engineering, University of Delaware, Newark, Delaware 19716, United States(材料科学与工程系,德雷克塞尔大学) Department of Chemistry, University of Pennsylvania, Philadelphia, Pennsylvania 19104, United States(化学系,宾夕法尼亚大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Journal ref Chem. Mater. 2025, 37, 7002-7013

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.16357 2025-10-07 cs.CV 77%

Law of Vision Representation in MLLMs

Shijia Yang, Bohan Zhai, Quanzeng You, Jianbo Yuan, Hongxia Yang, Chenfeng Xu

机构 * Stanford University(斯坦福大学) UC Berkeley(加州大学伯克利分校) The Hong Kong Polytechnic University(香港理工大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV

Comments The code is available at https://github.com/bronyayang/Law_of_Vision_Representation_in_MLLMs

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10666 2025-10-07 astro-ph.SR astro-ph.GA astro-ph.IM 75%

Machine-learning inference of stellar properties using integrated photometric and spectroscopic data

Ilay Kamai, Alex M. Bronstein, Hagai B. Perets

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);cross-modal(abstract)

Comments Accepted to ApJ

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03455 2025-10-07 cs.CV 70%

PEaRL: Pathway-Enhanced Representation Learning for Gene and Pathway Expression Prediction from Histology

Sejuti Majumder, Saarthak Kapse, Moinak Bhattacharya, Xuan Xu, Alisa Yurovsky, Prateek Prasanna

机构 * Department of Biomedical Informatics, Stony Brook University, NY, USA(生物医学信息学系,石溪大学,纽约,美国) Department of Computer Science, Stony Brook University, NY, USA(计算机科学系,石溪大学,纽约,美国)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04417 2025-10-07 cs.LG cs.AI cs.CL cs.CV cs.IT math.IT 67%

Partial Information Decomposition via Normalizing Flows in Latent Gaussian Distributions

Wenyuan Zhao, Adithya Balachandran, Chao Tian, Paul Pu Liang

机构 * Texas A&M University(德克萨斯大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20752 2025-10-07 cs.CV cs.AI 62%

Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language Models

Huajie Tan, Yuheng Ji, Xiaoshuai Hao, Xiansheng Chen, Pengwei Wang, Zhongyuan Wang, Shanghang Zhang

机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机科学学院,北京大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院自动化研究所人工智能学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 51 pages, 23 figures, NeurIPS'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04196 2025-10-07 cs.AI cs.LG 57%

COSMO-RL: Towards Trustworthy LMRMs via Joint Safety and Stability

Yizhuo Ding, Mingkang Chen, Qiuhua Liu, Fenghua Weng, Wanying Qu, Yue Yang, Yugang Jiang, Zuxuan Wu, Yanwei Fu, Wenqi Shao

机构 * Fudan University(复旦大学) Shanghai AI Laboratory(上海人工智能实验室) The University of Hong Kong(香港大学) Shenzhen University(深圳大学) ShanghaiTech University(上海交通大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏