arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-16 至 2025-12-16 共收录 87 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 8 篇

2512.13072 2025-12-16 cs.CV 70%

Forging a Dynamic Memory: Retrieval-Guided Continual Learning for Generalist Medical Foundation Models

锻造动态记忆:基于检索的持续学习用于通用医学基础模型

Zizhi Chen, Yizhen Gao, Minghao Han, Yizhou Liu, Zhaoyu Chen, Dingkang Yang, Lihua Zhang

机构 * College of Intelligent Robotics and Advanced Manufacturing(智能机器人与先进制造学院) Fudan University(复旦大学) Fysics Intelligence Technologies Co., Ltd. (Fysics AI)(Fysics智能技术有限公司(Fysics AI)) School of Computer Science and Engineering(计算机科学与工程学院) Central South University(中南大学)

专题命中 跨模态检索 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

AI总结 本文提出基于检索的持续学习方法,通过动态知识蒸馏和RAG技术,解决多模态医学模型在领域迁移和细粒度特征保留中的核心难题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12309 2025-12-16 cs.CV 57%

WeDetect: Fast Open-Vocabulary Object Detection as Retrieval

WeDetect:快速开放词汇物体检测作为检索

Shenghao Fu, Yukun Su, Fengyun Rao, Jing Lyu, Xiaohua Xie, Wei-Shi Zheng

机构 * WeChat Vision, Tencent Inc.(腾讯微信视觉部) Pazhou Laboratory (Huangpu)(琶洲实验室(黄埔)) Shenzhen Loop Area Institute(深圳河套学院)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 WeDetect通过检索框架实现快速开放词汇物体检测,结合双塔架构和LMMs,提升检测效率和通用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13102 2025-12-16 cs.CV 57%

CapeNext: Rethinking and Refining Dynamic Support Information for Category-Agnostic Pose Estimation

CapeNext: 重新思考和细化用于类别无关姿态估计的动态支持信息

Yu Zhu, Dan Zeng, Shuiwang Li, Qijun Zhao, Qiaomu Shen, Bo Tang

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 CapeNext通过整合层次交叉模态交互与双流特征细化,提升了类别无关姿态估计中动态支持信息的鲁棒性和鉴别能力。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态生成 7 篇

2512.12756 2025-12-16 cs.CV 88%

FysicsWorld: A Unified Full-Modality Benchmark for Any-to-Any Understanding, Generation, and Reasoning

FysicsWorld: 一个统一的全模态基准用于任意到任意的理解、生成和推理

Yue Jiang, Dingkang Yang, Minghao Han, Jinghang Han, Zizhi Chen, Yizhou Liu, Mingcheng Li, Peng Zhai, Lihua Zhang

机构 * College of Intelligent Robotics and Advanced Manufacturing(智能机器人与先进制造学院) Fudan University(复旦大学) Fysics Intelligence Technologies Co., Ltd.(菲茨智能技术有限公司)

专题命中 多模态生成 :any-to-any(title,abstract);omni-modal(abstract,comments);multimodal(abstract);cross-modal(abstract)

AI总结 FysicsWorld是一个统一的全模态基准,支持图像、视频、音频和文本之间的双向交互,通过16个主要任务和3,268个样本全面评估理解、生成和推理能力,揭示模型在多模态任务中的性能差异。

Comments The omni-modal benchmark report from Fysics AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04810 2025-12-16 cs.CV 79%

EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture

EMMA: 一种高效的多模态理解、生成与编辑统一架构

Xin He, Longhui Wei, Jianbo Ouyang, Minghui Liao, Lingxi Xie, Qi Tian

机构 * Huawei Inc.(华为公司)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 EMMA提出了一种高效的统一多模态架构,通过高效自编码器、通道级拼接、共享解耦网络和专家混合机制,在效率和性能上超越现有方法。

Comments Project Page: https://emma-umm.github.io/emma/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11906 2025-12-16 cs.CV cs.LG 79%

MPath: Multimodal Pathology Report Generation from Whole Slide Images

MPath:从全切片图像生成多模态病历报告

Noorul Wahab, Nasir Rajpoot

机构 * TIA Centre, Department of Computer Science, University of Warwick, UK(沃里克大学计算机科学系TIA研究中心)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 MPath通过多模态条件化方法,利用预训练的生物医学语言模型生成病理报告,展示了在病理报告生成中的有效性与可扩展性。

Comments Pages 4, Figures 1, Table 1

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12196 2025-12-16 cs.MM cs.CV cs.SD eess.AS 67%

AutoMV: An Automatic Multi-Agent System for Music Video Generation

AutoMV: 一种自动多智能体系统用于音乐视频生成

Xiaoxuan Tang, Xinping Lei, Chaoran Zhu, Shiyun Chen, Ruibin Yuan, Yizhi Li, Changjae Oh, Ge Zhang, Wenhao Huang, Emmanouil Benetos, Yang Liu, Jiaheng Liu, Yinghao Ma

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Nanjing University(南京大学) Queen Mary University of London(伦敦大学女王学院) Hong Kong University of Science and Technology(香港科技大学) University of Manchester(曼彻斯特大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.MM、eess.AS

AI总结 AutoMV通过多智能体协作生成完整音乐视频,优于现有方法并接近专业水准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13690 2025-12-16 cs.CV cs.AI cs.GR cs.LG 62%

DiffusionBrowser: Interactive Diffusion Previews via Multi-Branch Decoders

DiffusionBrowser: 通过多分支解码器实现交互式扩散预览

Susung Hong, Chongjian Ge, Zhifei Zhang, Jui-Hsien Wang

机构 * University of Washington(华盛顿大学) Adobe Research(Adobe研究)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 DiffusionBrowser通过多分支解码器实现交互式视频生成预览,提升生成效率与可控性。

Comments Project page: https://susunghong.github.io/DiffusionBrowser

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13276 2025-12-16 cs.CV 57%

CogniEdit: Dense Gradient Flow Optimization for Fine-Grained Image Editing

CogniEdit: 密集梯度流优化用于细粒度图像编辑

Yan Li, Lin Liu, Xiaopeng Zhang, Wei Xue, Wenhan Luo, Yike Guo, Qi Tian

机构 * Hongkong University of Science and Technology(香港科学与技术大学) Huawei Company(华为公司)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 CogniEdit通过密集梯度流优化实现细粒度图像编辑,结合多模态推理与动态令牌焦点重新定位,提升指令遵循与视觉质量的平衡

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21329 2025-12-16 cs.AI physics.soc-ph 57%

Active Inference AI Systems for Scientific Discovery

主动推断AI系统用于科学发现

Karthik Duraisamy

机构 * Michigan Institute for Computational Discovery & Engineering(密歇根计算发现与工程研究所) University of Michigan(密歇根大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 本文提出主动推断AI系统,通过结合缓慢假设生成与快速验证推理,利用因果和多模态模型促进科学发现,强调人类判断在处理不确定性中的关键作用。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态评测 16 篇

2503.06940 2025-12-16 cs.CV 85%

CineBrain: A Large-Scale Multi-Modal Brain Dataset During Naturalistic Audiovisual Narrative Processing

CineBrain: 一种大规模多模态大脑数据集用于自然视听叙事处理

Jianxiong Gao, Yichang Liu, Baofeng Yang, Jianfeng Feng, Yanwei Fu

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 CineBrain通过多模态融合编码器和神经潜在解码器实现动态视频重建,利用fMRI和EEG互补优势提升视觉保真度,并揭示听觉输入对视觉感知的影响。

Comments Project Page: https://jianxgao.github.io/CineBrain

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10576 2025-12-16 cs.CV 83%

HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs

HumanSense: 从多模态感知到通过推理的 empathetic 上下文感知响应

Zheng Qin, Ruobing Zheng, Yabing Wang, Tianqi Li, Yi Yuan, Jingdong Chen, Le Wang

专题命中 多模态评测 :multimodal(title,abstract);omni-modal(abstract);分类 cs.CV

AI总结 HumanSense通过多阶段模态渐进强化学习提升MLLMs在多模态感知和交互任务中的性能,强调推理能力对共情反馈的重要性。

Comments Accepted by AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13297 2025-12-16 cs.AI cs.LG 79%

MedInsightBench: Evaluating Medical Analytics Agents Through Multi-Step Insight Discovery in Multimodal Medical Data

MedInsightBench: 通过多步骤洞察发现评估医疗分析代理在多模态医疗数据中的能力

Zhenghao Zhu, Chuxue Cao, Sirui Han, Yuanfeng Song, Xing Chen, Caleb Chen Cao, Yike Guo

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) ByteDance(字节跳动)

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.AI

AI总结 MedInsightBench通过多步骤洞察发现评估医疗分析代理在多模态医疗数据中的能力,提出MedInsightAgent框架提升医疗数据洞察发现性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12503 2025-12-16 cs.AI 57%

KidsArtBench: Multi-Dimensional Children's Art Evaluation with Attribute-Aware MLLMs

KidsArtBench: 多维度儿童艺术评估与属性感知的大语言模型

Mingrui Ye, Chanjin Zheng, Zengyi Yu, Chenyu Xiang, Zhixue Zhao, Zheng Yuan, Helen Yannakoudakis

机构 * King’s College London(伦敦国王学院) East China Normal University(东华大学) University of Sheffield(谢菲尔德大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

AI总结 KidsArtBench通过属性感知的多LoRA方法提升儿童艺术评估的准确性与教育意义。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12500 2025-12-16 cs.HC cs.AI 57%

Explainable AI as a Double-Edged Sword in Dermatology: The Impact on Clinicians versus The Public

可解释AI在皮肤科中的双刃剑作用:对医生与公众的影响

Xuhai Xu, Haoyu Hu, Haoran Zhang, Will Ke Wang, Reina Wang, Luis R. Soenksen, Omar Badri, Sheharbano Jafry, Elise Burger, Lotanna Nwandu, Apoorva Mehta, Erik P. Duhaime, Asif Qasim, Hause Lin, Janis Pereira, Jonathan Hershon, Paulius Mui, Alejandro A. Gru, Noémie Elhadad, Lena Mamykina, Matthew Groh, Philipp Tschandl, Roxana Daneshjou, Marzyeh Ghassemi

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

AI总结 研究探讨了可解释AI在皮肤科诊断中的影响,发现不同专业背景的人对XAI的反应不同,LLM在医疗AI中具有双刃剑效应。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12444 2025-12-16 cs.CL 57%

Can GPT replace human raters? Validity and reliability of machine-generated norms for metaphors

GPT能否替代人类评分者?机器生成的隐喻规范的有效性和可靠性

Veronica Mangiaterra, Hamad Al-Azary, Chiara Barattieri di San Pietro, Paolo Canal, Valentina Bambini

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

AI总结 本文研究了GPT在隐喻评分中的有效性与可靠性,发现较大模型能有效替代人类评分,但需注意隐喻惯例性和多模态因素的影响。

Comments 30 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12443 2025-12-16 cs.AI cs.SE 57%

AI Transparency Atlas: Framework, Scoring, and Real-Time Model Card Evaluation Pipeline

AI透明度地图:框架、评分与实时模型卡片评估流程

Akhmadillo Mamirov, Faiaz Azmain, Hanyu Wang

机构 * Department of Computer Science, The College of Wooster, Wooster, OH, USA(计算机科学系,沃斯特学院,沃斯特,OH,USA) Robert F. Wagner Graduate School of Public Service, New York University, New York, NY, USA(罗伯特·F·瓦格纳公共事务研究生学院,纽约大学,纽约,NY,USA)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

AI总结 本文提出AI透明度框架和实时评估流程,评估50个模型发现前沿实验室合规率约80%,而多数供应商低于60%,安全关键领域存在显著缺口。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12424 2025-12-16 cs.CV cs.LG 57%

ViInfographicVQA: A Benchmark for Single and Multi-image Visual Question Answering on Vietnamese Infographics

ViInfographicVQA:单图和多图越南信息图视觉问答基准测试

Tue-Thu Van-Dinh, Hoang-Duy Tran, Truong-Binh Duong, Mai-Hanh Pham, Binh-Nam Le-Nguyen, Quoc-Thai Nguyen

机构 * Neu.edu.vn

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 ViInfographicVQA是首个针对越南信息图VQA的基准测试,通过单图和多图任务评估模型在复杂信息图中的视觉问答能力。

Comments 10 pages, 4 figures, Accepted to AI4Research @ AAAI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04457 2025-12-16 cs.CL 57%

Do MLLMs Really Understand the Charts?

多模态大语言模型真的能理解图表吗?

Xiao Zhang, Dongyuan Li, Liuyu Xiang, Yao Zhang, Cheng Zhong, Zhaofeng He

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) AAITC, CTO Organization, Lenovo(AAITC、CTO组织、联想) Lenovo(联想)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

AI总结 本文提出 ChartVR-3B/7B 模型,通过 VR-RFT 策略提升图表视觉推理能力,并在 ChartVRBench 上取得优异成绩。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13405 2025-12-16 cs.CL 57%

RealHiTBench: A Comprehensive Realistic Hierarchical Table Benchmark for Evaluating LLM-Based Table Analysis

RealHiTBench: 一个全面的现实感分层表格基准,用于评估基于LLM的表格分析

Pengzuo Wu, Yuhang Yang, Guangcheng Zhu, Chao Ye, Hong Gu, Xu Lu, Ruixuan Xiao, Bowen Bao, Yijing He, Liangyu Zha, Wentao Ye, Junbo Zhao, Haobo Wang

机构 * Zhejiang University(浙江大学) vivo Mobile Communication Co., Ltd(vivo移动通信有限公司) Institute of Computing Innovation, Zhejiang University(浙江大学计算机创新研究院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

AI总结 RealHiTBench是一个用于评估基于LLM的表格分析能力的综合现实感分层表格基准,通过多种输入格式和复杂结构的表格测试,验证了改进LLMs对表格层次结构感知的重要性。

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17159 2025-12-16 cs.CV 57%

RobustMerge: Parameter-Efficient Model Merging for MLLMs with Direction Robustness

RobustMerge: 用于MLLMs的参数高效模型合并的鲁棒性

Fanhu Zeng, Haiyang Guo, Fei Zhu, Li Shen, Hao Tang

机构 * MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS) Centre for Artificial Intelligence and Robotics, HKISI-CAS(香港科学院智能机器人中心) Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) Shenzhen Loop Area Institute(深圳河套学院) State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机科学学院多媒体信息处理国家重点实验室)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 RobustMerge提出一种无需训练的参数高效模型合并方法,通过参数修剪和跨任务归一化提升多任务泛化能力。

Comments NeurIPS 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12059 2025-12-16 cs.AI 57%

The Forecast Critic: Leveraging Large Language Models for Poor Forecast Identification

预测批评者:利用大语言模型进行差预测识别

Luke Bhan, Hanyu Zhang, Andrew Gordon Wilson, Michael W. Mahoney, Chuck Arvin

机构 * Amazon(亚马逊)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.AI

AI总结 利用大语言模型进行预测监控,通过识别不合理预测提升预测质量。

Comments Presented at AAAI 2026 AI4TS workshop and AABA4ET workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17699 2025-12-16 cs.CV 57%

Establishing Reality-Virtuality Interconnections in Urban Digital Twins for Superior Intelligent Road Inspection and Simulation

在城市数字孪生中建立现实-虚拟联系以实现更智能的道路检查与模拟

Yikang Zhang, Chuang-Wei Liu, Jiahang Li, Yingbing Chen, Jie Cheng, Rui Fan

机构 * College of Electronics & Information Engineering, Shanghai Research Institute for Intelligent Autonomous Systems, the State Key Laboratory of Intelligent Autonomous Systems, and Frontiers Science Center for Intelligent Autonomous Systems, Tongji University, Shanghai 201804, China(电子与信息工程学院,上海智能自主系统研究院,智能自主系统国家重点实验室,智能自主系统前沿科学中心,同济大学,上海)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出了一种多模态传感器平台与城市数字孪生结合的智能道路检查系统,通过高保真的道路缺陷场景提升驾驶任务的感知和决策性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13010 2025-12-16 cs.LG q-bio.TO 50%

Deep Learning-Driven Inversion Framework for Shear Modulus Estimation in Magnetic Resonance Elastography (DIME)

基于深度学习的磁共振弹性成像中剪切模量估计反演框架(DIME)

Hassan Iftikhar, Rizwan Ahmad, Arunark Kolipaka

机构 * Biomedical Engineering, The Ohio State University(生物医学工程,俄亥俄州立大学) Department of Radiology, The Ohio State University(放射学系,俄亥俄州立大学) Davis Heart & Lung Research Institute, The Ohio State University(戴维斯心脏与肺研究所,俄亥俄州立大学)

专题命中 多模态评测 :multimodal(abstract)

AI总结 本文提出DIME框架,利用深度学习提升MRE中剪切模量估计的鲁棒性和精度,实验显示其在刚度图生成和真实值匹配方面优于传统MMDI方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15497 2025-12-16 eess.SP 50%

A Review of Machine Learning for Cavitation Intensity Recognition in Complex Industrial Systems

机械系统中空化强度识别的机器学习综述

Yu Sha, Ningtao Liu, Haofeng Liu, Junqi Tao, Zhenxing Niu, Guojun Huang, Yao Yao, Jiaqi Liang, Moxian Qian, Horst Stoecker, Domagoj Vnucec, Andreas Widl, Kai Zhou

专题命中 多模态评测 :multi-modal(abstract)

AI总结 本文综述了机械系统中空化强度识别的机器学习发展,强调传统方法与深度学习的演变,并展望未来在多源数据处理和工业应用中的发展方向。

Comments 43 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15562 2025-12-16 cs.RO 50%

Transformer-XL for Long Sequence Tasks in Robotic Learning from Demonstration

Transformer-XL在机器人学习从示范中的长序列任务应用

Gao Tianci

机构 * Bauman Moscow State Technical University(巴甫洛夫莫斯科国立技术大学)

专题命中 多模态评测 :multi-modal(abstract)

AI总结 本文提出利用Transformer-XL处理机器人学习从示范中的长序列任务,通过多模态传感器输入提升任务成功率和计算效率。

Comments 11 pages,4 figures

Journal ref Tianci G. Transformer-xl for long sequence tasks in robotic learning from demonstrations[C]//International Conference on Intelligent Systems. Cham: Springer Nature Switzerland, 2024: 25-36

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 多模态Agent 7 篇

2512.12799 2025-12-16 cs.CV 83%

DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planning

DrivePI: 基于空间感知的4D MLLM用于统一自动驾驶理解、感知、预测与规划

Zhe Liu, Runhui Huang, Rui Yang, Siming Yan, Zining Wang, Lu Hou, Di Lin, Xiang Bai, Hengshuang Zhao

机构 * The University of Hong Kong(香港大学) Yinwang Intelligent Technology Co. Ltd.(英维智能科技有限公司) Tianjin University(天津大学) Huazhong University of Science and Technology(华中科技大学)

专题命中 多模态Agent :MLLM(title,abstract);multi-modal(abstract);分类 cs.CV

AI总结 DrivePI是一种基于空间感知的4D MLLM,用于统一自动驾驶的理解、感知、预测和规划,通过端到端优化实现多任务并行,提升性能并减少碰撞率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20164 2025-12-16 cs.CL 79%

Thinking with Visual Abstract: Enhancing Multimodal Reasoning via Visual Abstraction

通过视觉抽象思考:通过视觉抽象增强多模态推理

Dairu Liu, Ziyue Wang, Minyuan Ruan, Fuwen Luo, Chi Chen, Peng Li, Yang Liu

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL

AI总结 通过引入视觉抽象思考范式,提升多模态大语言模型在视觉感知和推理任务中的性能,实现更高效的视觉推理机制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11876 2025-12-16 cs.RO cs.SY eess.SY 78%

Traversability Aware Autonomous Navigation for Multi-Modal Mobility Morphobot (M4)

多模态移动形变机器人(M4)的可 traversability 自主导航

Hrigved Mahesh Suryawanshi

机构 * SiliconSynapse Lab(硅合成实验室)

专题命中 多模态Agent :multi-modal(title,abstract)

AI总结 本研究提出了一种基于LiDAR的多模态移动形变机器人M4的可 traversability 自主导航框架,通过学习地形分析生成节能路径,提升地形适应能力。

Comments Master's thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23046 2025-12-16 cs.CL cs.AI cs.CV cs.RO 67%

SoMi-ToM: Evaluating Multi-Perspective Theory of Mind in Embodied Social Interactions

SoMi-ToM:评估具身社会互动中的多视角理论之心假设

Xianzhe Fan, Xuhui Zhou, Chuanyang Jin, Kolby Nottingham, Hao Zhu, Maarten Sap

机构 * The University of Hong Kong(香港大学) Carnegie Mellon University(卡内基梅隆大学) Johns Hopkins University(约翰霍普金斯大学) University of California Irvine(加州大学尔湾分校) Stanford University(斯坦福大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 SoMi-ToM基准通过多视角评估人类与模型在具身社会互动中的理论之心能力,揭示大型视觉-语言模型在复杂社交场景中的不足。

Comments 24 pages, 6 figures

Journal ref Proceedings of the 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏