arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-16 至 2025-12-16 共收录 7 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 7 篇

2512.12799 2025-12-16 cs.CV 83%

DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planning

DrivePI: 基于空间感知的4D MLLM用于统一自动驾驶理解、感知、预测与规划

Zhe Liu, Runhui Huang, Rui Yang, Siming Yan, Zining Wang, Lu Hou, Di Lin, Xiang Bai, Hengshuang Zhao

机构 * The University of Hong Kong(香港大学) Yinwang Intelligent Technology Co. Ltd.(英维智能科技有限公司) Tianjin University(天津大学) Huazhong University of Science and Technology(华中科技大学)

专题命中 多模态Agent :MLLM(title,abstract);multi-modal(abstract);分类 cs.CV

AI总结 DrivePI是一种基于空间感知的4D MLLM,用于统一自动驾驶的理解、感知、预测和规划,通过端到端优化实现多任务并行,提升性能并减少碰撞率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20164 2025-12-16 cs.CL 79%

Thinking with Visual Abstract: Enhancing Multimodal Reasoning via Visual Abstraction

通过视觉抽象思考:通过视觉抽象增强多模态推理

Dairu Liu, Ziyue Wang, Minyuan Ruan, Fuwen Luo, Chi Chen, Peng Li, Yang Liu

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL

AI总结 通过引入视觉抽象思考范式,提升多模态大语言模型在视觉感知和推理任务中的性能,实现更高效的视觉推理机制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11876 2025-12-16 cs.RO cs.SY eess.SY 78%

Traversability Aware Autonomous Navigation for Multi-Modal Mobility Morphobot (M4)

多模态移动形变机器人(M4)的可 traversability 自主导航

Hrigved Mahesh Suryawanshi

机构 * SiliconSynapse Lab(硅合成实验室)

专题命中 多模态Agent :multi-modal(title,abstract)

AI总结 本研究提出了一种基于LiDAR的多模态移动形变机器人M4的可 traversability 自主导航框架,通过学习地形分析生成节能路径,提升地形适应能力。

Comments Master's thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23046 2025-12-16 cs.CL cs.AI cs.CV cs.RO 67%

SoMi-ToM: Evaluating Multi-Perspective Theory of Mind in Embodied Social Interactions

SoMi-ToM:评估具身社会互动中的多视角理论之心假设

Xianzhe Fan, Xuhui Zhou, Chuanyang Jin, Kolby Nottingham, Hao Zhu, Maarten Sap

机构 * The University of Hong Kong(香港大学) Carnegie Mellon University(卡内基梅隆大学) Johns Hopkins University(约翰霍普金斯大学) University of California Irvine(加州大学尔湾分校) Stanford University(斯坦福大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 SoMi-ToM基准通过多视角评估人类与模型在具身社会互动中的理论之心能力,揭示大型视觉-语言模型在复杂社交场景中的不足。

Comments 24 pages, 6 figures

Journal ref Proceedings of the 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12921 2025-12-16 cs.CR cs.AI 57%

Cisco Integrated AI Security and Safety Framework Report

思科集成AI安全与安全框架报告

Amy Chang, Tiffany Saade, Sanket Mendapara, Adam Swanda, Ankit Garg

机构 * Cisco AI Threat and Security Research(思科人工智能威胁与安全研究)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 本文提出思科集成AI安全与安全框架,旨在统一分类和操作化AI风险,涵盖安全与安全,适用于威胁识别、风险优先级排序等,具有全面性和扩展性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12630 2025-12-16 cs.HC cs.AI 57%

ORIBA: Exploring LLM-Driven Role-Play Chatbot as a Creativity Support Tool for Original Character Artists

ORIBA:探索基于大语言模型的对话机器人作为原创角色艺术家创造力支持工具

Yuqian Sun, Xingyu Li, Shunyu Yao, Noura Howell, Tristan Braud, Chang Hee Lee, Ali Asadipour

机构 * Computer Science Research Centre, Royal College of Art(皇家艺术学院计算机科学研究中心) Digital Media, School of Literature, Media, and Communication, Georgia Institute of Technology(佐治亚理工学院数字媒体系) Princeton University(普林斯顿大学) Digital Media, Georgia Institute of Technology(佐治亚理工学院数字媒体系) Division of Integrative Systems and Design, The Hong Kong University of Science and Technology(香港科学大学整合系统与设计 division) Industrial Design Department, College of Engineering, KAIST(韩国科学技术院工程学院工业设计系)

专题命中 多模态Agent :cross-modal(abstract);分类 cs.AI

AI总结 ORIBA通过大语言模型支持原创角色艺术家的创意过程,平衡AI辅助与创意自主权。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02580 2025-12-16 cs.CL 57%

From Imitation to Discrimination: Toward A Generalized Curriculum Advantage Mechanism Enhancing Cross-Domain Reasoning Tasks

从模仿到辨别:一种通用的课程优势机制,增强跨领域推理任务

Changpeng Yang, Jinyang Wu, Yuchen Liu, Shuai Zhang, Yang Li, Qiliang Liang, Hongzhen Wang, Shuai Nie, Jiaming Xu, Runyu Shi, Ying Huang, Guoquan Zhang

机构 * \equalcontrib(1号机构)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

AI总结 CAPO是一种基于优势信号的自适应课程机制,通过引导模仿学习和引入负信号提升跨领域推理任务的泛化能力。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏