arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-11 至 2025-11-11 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 6 篇

2511.05952 2025-11-11 cs.HC cs.CV cs.MM 84%

Pinching Visuo-haptic Display: Investigating Cross-Modal Effects of Visual Textures on Electrostatic Cloth Tactile Sensations

Takekazu Kitagishi, Chun-Wei Ooi, Yuichi Hiroi, Jun Rekimoto

机构 * The University of Tokyo(东京大学) ZOZO Research(ZOZO研究) Cluster Metaverse Lab(集群元宇宙实验室) Sony CSL Kyoto(索尼 CSL京都)

专题命中 其他多模态 :cross-modal(title,abstract);multimodal(abstract,comments);分类 cs.CV、cs.MM

Comments 10 pages, 8 figures, 3 tables. Presented at ACM International Conference on Multimodal Interaction (ICMI) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06749 2025-11-11 cs.RO cs.CV 79%

Semi-distributed Cross-modal Air-Ground Relative Localization

Weining Lu, Deer Bin, Lian Ma, Ming Ma, Zhihao Ma, Xiangyang Chen, Longfei Wang, Yixiao Feng, Zhouxian Jiang, Yongliang Shi, Bin Liang

机构 * Beijng National Research Center for Information Science and Technology(北京国家信息科学与技术研究中心) Qiyuan Lab(启元实验室) JiangHuai Advanced Technology Center(江淮先进技术中心)

专题命中 其他多模态 :cross-modal(title,abstract);分类 cs.CV

Comments 7 pages, 3 figures. Accepted by IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07004 2025-11-11 cs.CV cs.HC 57%

Exploring the "Great Unseen" in Medieval Manuscripts: Instance-Level Labeling of Legacy Image Collections with Zero-Shot Models

Christofer Meinecke, Estelle Guéville, David Joseph Wrisley

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06744 2025-11-11 cs.CV 57%

PointCubeNet: 3D Part-level Reasoning with 3x3x3 Point Cloud Blocks

Da-Yeong Kim, Yeong-Jun Cho

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06297 2025-11-11 cs.HC cs.AI 57%

Decomate: Leveraging Generative Models for Co-Creative SVG Animation

Jihyeon Park, Jiyoon Myung, Seone Shin, Jungki Son, Joohyung Han

机构 * MODULABS

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted at the 1st Workshop on Generative and Protective AI for Content Creation (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06256 2025-11-11 cs.CV 57%

VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving

Ruifei Zhang, Wei Zhang, Xiao Tan, Sibei Yang, Xiang Wan, Xiaonan Luo, Guanbin Li

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shenzhen Research Institute of Big Data(深圳大数据研究院) Sun Yat-sen University(中山大学) Baidu Inc.(百度公司) Guilin University of Electronic Technology(桂林电子科技大学) Guangdong Key Laboratory of Big Data Analysis and Processing(广东大数据分析与处理重点实验室)

专题命中 其他多模态 :MLLM(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏