arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-01-21 至 2026-01-21 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 9 篇

2601.12981 2026-01-21 cs.CV cs.LG 79%

Early Prediction of Type 2 Diabetes Using Multimodal data and Tabular Transformers

利用多模态数据和表格变压器进行2型糖尿病早期预测

Sulaiman Khan, Md. Rafiul Biswas, Zubair Shah

机构 * College of Science and Engineering(科学与工程学院)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

AI总结 利用TabTrans分析多模态数据,通过预测T2DM风险,提升糖尿病管理的前瞻性与个性化水平。

Comments 08 pages, 06 figures, accepted for publication in FLLM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12037 2026-01-21 cs.HC cs.CV cs.SY eess.SY 79%

Multimodal Feedback for Handheld Tool Guidance: Combining Wrist-Based Haptics with Augmented Reality

多模态反馈用于手持工具引导:结合基于手腕的触觉反馈与增强现实

Yue Yang, Christoph Leuze, Brian Hargreaves, Bruce Daniel, Fred M Baik

机构 * Stanford University(斯坦福大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

AI总结 本研究通过结合增强现实与触觉反馈,提升手术中手持工具的精确引导与操作效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08434 2026-01-21 cs.RO cs.AI 79%

Large Multimodal Models for Embodied Intelligent Driving: The Next Frontier in Self-Driving?

大规模多模态模型用于具身智能驾驶:自我驾驶的下一个前沿?

Long Zhang, Yuchen Xia, Bingqing Wei, Zhen Liu, Shiwen Mao, Zhu Han, Mohsen Guizani

机构 * School of Information and Electrical Engineering, Hebei University of Engineering(河北工程大学信息与电子工程学院) School of Information Science and Engineering, Lanzhou University(兰州大学信息科学与工程学院) Department of Electrical and Computer Engineering, Auburn University(阿肯色大学电气与计算机工程系) Department of Electrical and Computer Engineering, University of Houston(休斯顿大学电气与计算机工程系) Department of Computer Science and Engineering, Kyung Hee University(庆熙大学计算机科学与工程系) Machine Learning Department, Mohamed Bin Zayed University of Artificial Intelligence(Mohamed Bin Zayed人工智能大学机器学习系)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出了一种语义和策略双驱动的混合决策框架,用于提升具身智能驾驶中的持续学习与联合决策能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18177 2026-01-21 cs.CV cs.LG cs.MA 79%

Scene-Aware Vectorized Memory Multi-Agent Framework with Cross-Modal Differentiated Quantization VLMs for Visually Impaired Assistance

具有跨模态差异化量化功能的场景感知向量内存多智能体框架用于视障辅助

Xiangxiang Wang, Xuanyu Wang, YiJia Luo, Yongbin Yu, Manping Fan, Jingtao Zhang, Liyong Ren

机构 * School of Information Software Engineering, University of Electronic Science Sichuan Provincial Key Laboratory for Human Disease Gene Study, Sichuan Academy of Medical Sciences \& Sichuan Provincial People's Hospital, University of Electronic Science Faculty of Computing, Harbin Institute of Technology, Harbin, China

专题命中 多模态Agent :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出了一种具有跨模态差异化量化的多智能体框架,通过减少内存消耗和提升处理效率,为视障人士提供更高效的环境感知与辅助导航支持。

Comments 28 pages,9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12582 2026-01-21 cond-mat.mtrl-sci cs.AI 74%

Ontology-aligned structuring and reuse of multimodal materials data and workflows towards automatic reproduction

面向多模态材料数据和工作流的本体对齐结构化与重用,以实现自动重现

Sepideh Baghaee Ravari, Abril Azocar Guzman, Sarath Menon, Stefan Sandfeld, Tilmann Hickel, Markus Stricker

专题命中 多模态Agent :multimodal(title);分类 cs.AI

AI总结 本文提出一种基于本体驱动和大型语言模型的框架,用于自动提取和结构化多模态材料数据和工作流,以提高计算结果的可重现性和重用性。

Comments 39 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10672 2026-01-21 cs.RO cs.CV 57%

Vision Language Action Models in Robotic Manipulation: A Systematic Review

视觉语言动作模型在机器人操作中的应用:系统综述

Muhayy Ud Din, Waseem Akram, Lyes Saad Saoud, Jan Rosell, Irfan Hussain

机构 * Khalifa University Center for Autonomous Robotic Systems (KUCARS), Khalifa University, United Arab Emirates(卡利法大学自主机器人系统中心(KUCARS)、卡利法大学、阿拉伯联合酋长国) Institute of Industrial and Control Engineering (IOC), Universitat Politecnica de Catalunya, Spain(工业与控制工程研究所(IOC)、巴塞罗那技术大学、西班牙)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 本文综述了视觉语言动作模型在机器人操作中的应用,分析了102个模型、26个数据集和12个模拟平台,探讨了多模态对齐和可扩展预训练等关键挑战。

Comments submitted to annual review in control

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13801 2026-01-21 cs.RO 50%

HoverAI: An Embodied Aerial Agent for Natural Human-Drone Interaction

HoverAI: 一种用于自然人-无人机交互的具身空中代理

Yuhua Jin, Nikita Kuzmin, Georgii Demianchuk, Mariya Lezina, Fawad Mehboob, Issatay Tokmurziyev, Miguel Altamirano Cabrera, Muhammad Ahsan Mustafa, Dzmitry Tsetserukou

机构 * Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Skolkovo Institute of Science and Technology(斯克尔科沃信息科技研究所)

专题命中 多模态Agent :multimodal(abstract)

AI总结 HoverAI通过结合无人机移动、视觉投影和对话式AI,实现了人-无人机自然交互的具身代理,提升了空间感知与社交响应能力。

Comments This paper has been accepted for publication at LBR HRI 2026 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12993 2026-01-21 cs.RO 50%

Being-H0.5: Scaling Human-Centric Robot Learning for Cross-Embodiment Generalization

Being-H0.5:面向跨躯体泛化的以人为本的机器人学习规模化

Hao Luo, Ye Wang, Wanpeng Zhang, Sipeng Zheng, Ziheng Xi, Chaoyi Xu, Haiweng Xu, Haoqi Yuan, Chi Zhang, Yiqing Wang, Yicheng Feng, Zongqing Lu

机构 * BeingBeyond Team(BeingBeyond 团队)

专题命中 多模态Agent :multimodal(abstract)

AI总结 Being-H0.5通过以人为本的学习范式和统一动作空间,实现了在不同机器人平台间的跨躯体泛化,结合混合流框架和流形保持门控,达到模拟和现实中的高性能表现。

Comments 44 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01147 2026-01-21 cs.RO 50%

Astra: Efficient Transformer Architecture and Contrastive Dynamics Learning for Embodied Instruction Following

Astra:面向具身指令跟随的高效Transformer架构与对比动态学习

Yueen Ma, Dafeng Chi, Shiguang Wu, Yuecheng Liu, Yuzheng Zhuang, Irwin King

机构 * Department of Computer Science and Engineering, The Chinese University of Hong Kong(计算机科学与工程系,香港中文大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 多模态Agent :multimodal(abstract)

AI总结 Astra通过引入轨迹注意力和对比动态学习目标,提升了具身指令跟随任务中多模态序列处理的效率与准确性。

Comments Accepted to EMNLP 2025 (main). Published version: https://aclanthology.org/2025.emnlp-main.688/ Code available at: https://github.com/yueen-ma/Astra

详情

展开后加载摘要…

URL PDF HTML 收藏