arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-11 至 2025-08-11 共收录 44 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 5 篇

2502.00342 2025-08-11 cs.CV 57%

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering

Zechuan Li, Hongshan Yu, Yihao Ding, Yan Li, Yong He, Naveed Akhtar

机构 * organization= College of Electrical Information Engineering,Hunan University , city= Changsha , postcode= 410082 , state= Hunan , country= China organization= School of Computing \& Information Systems ,The University of Melbourne , city= Melbourne , postcode= VIC 3053 , state= VIC , country= Australia organization= School of Computer Science,The University of Sydney , city= Sydney , postcode= NSW 2006 , state= NSW , country= Australia organization= School of Artificial Intelligence ,Anhui University , city= Hefei , postcode= 230601 , state= Anhui , country= China

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

Comments This is a submitted version of a paper accepted by Information Fusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05637 2025-08-11 cs.HC cs.AI 57%

Automated Visualization Makeovers with LLMs

Siddharth Gangwar, David A. Selby, Sebastian J. Vollmer

机构 * University of Kaiserslautern–Landau (RPTU)(凯撒斯劳滕-兰道大学(RPTU)) Department of Data Science and its Applications, German Research Center for Artificial Intelligence (DFKI)(数据科学及其应用系,德国人工智能研究中心(DFKI))

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05646 2025-08-11 cs.HC cs.RO 50%

A Humanoid Social Robot as a Teaching Assistant in the Classroom

Thomas Sievers

机构 * Institute of Information Systems, University of Lübeck(吕贝克大学信息系统研究所)

专题命中 多模态Agent :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态训练与对齐 7 篇

2508.05991 2025-08-11 cs.CV cs.AI cs.CY 87%

ECMF: Enhanced Cross-Modal Fusion for Multimodal Emotion Recognition in MER-SEMI Challenge

Juewen Hu, Yexin Li, Jiulin Li, Shuo Chen, Pring Wong

机构 * State Key Laboratory of General Artificial Intelligence, BIGAI(人工智能通用基础理论国家重点实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16658 2025-08-11 cs.CL cs.AI 84%

Contextual Reinforcement in Multimodal Token Compression for Large Language Models

Naderdel Piero, Zacharias Cromwell, Nathaniel Wainwright, Matthias Nethercott

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05934 2025-08-11 cs.HC cs.AI cs.LG 79%

ASLSL: Adaptive shared latent structure learning with incomplete multi-modal physiological data for multi-dimensional emotional feature selection

Xueyuan Xu, Tianze Yu, Wenjia Dong, Fulin Wei, Li Zhuo

机构 * School of Information Science and Technology, Beijing University of Technology, Beijing 100124, China(信息科学与技术学院,北京理工大学,北京) School of Artificial Intelligence, Anhui University, Beijing 100124, China(人工智能学院,安徽大学,北京)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03155 2025-08-11 cs.LG cs.AI 79%

Fusing Cross-Domain Knowledge from Multimodal Data to Solve Problems in the Physical World

Yu Zheng

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06836 2025-08-11 cs.LG cond-mat.mtrl-sci cs.AI 79%

CAST: Cross Attention based multimodal fusion of Structure and Text for materials property prediction

Jaewan Lee, Changyoung Park, Hongjun Yang, Sungbin Lim, Woohyung Lim, Sehui Han

机构 * LG AI Research(LG人工智能研究)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06146 2025-08-11 cs.CV 70%

Text-guided Visual Prompt DINO for Generic Segmentation

Yuchen Guan, Chong Sun, Canmiao Fu, Zhipeng Huang, Chun Yuan, Chen Li

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) WeChat AI, Tencent Inc.(微信AI,腾讯公司)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06163 2025-08-11 cs.CL cs.AI cs.LG 62%

One Size Does Not Fit All: A Distribution-Aware Sparsification for More Precise Model Merging

Yingfeng Luo, Dingyang Lin, Junxin Wang, Ziqiang Xu, Kaiyan Chang, Tong Zheng, Bei Li, Anxiang Ma, Tong Xiao, Zhengtao Yu, Jingbo Zhu

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他多模态 4 篇

2508.02622 2025-08-11 cs.AI cs.CL cs.CY 62%

Noosemia: toward a Cognitive and Phenomenological Account of Intentionality Attribution in Human-Generative AI Interaction

Enrico De Santis, Antonello Rizzi

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments This version has been extensively revised and revisited in light of feedback and further research. Several sections have been expanded or improved for greater clarity and completeness. Specifically, new clarification on complex system foundation related to Noosemia has been added (Secs. "2.4 and "2.5")

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14956 2025-08-11 cs.LO 50%

Recursive windows for grammar logics of bounded density

Olivier Gasquet

专题命中 其他多模态 :multi-modal(abstract)

Comments This paper is still under construction, next versions will be uploaded

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05786 2025-08-11 cs.NE 50%

Functional Connectivity Graph Neural Networks

Yang Li, Luopeiwen Yi, Tananun Songdechakraiwut

专题命中 其他多模态 :multi-modal(abstract)

Comments 26 pages, 5 figures, 24 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16321 2025-08-11 cs.HC 50%

XR for All: Understanding Developers' Perspectives on Accessibility Integration in Extended Reality

Daniel Killough, Tiger F. Ji, Kexin Zhang, Yaxin Hu, Yu Huang, Ruofei Du, Yuhang Zhao

专题命中 其他多模态 :multimodal(abstract)

Comments 20 pages, 1 figure, 3 tables, LaTeX

详情

展开后加载摘要…

URL PDF HTML 收藏