arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-10 至 2026-02-10 共收录 127 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 18 篇

2602.08202 2026-02-10 cs.CV 57%

Generative Regression for Left Ventricular Ejection Fraction Estimation from Echocardiography Video

基于生成回归的超声心动图视频左室射血分数估计

Jinrong Lv, Xun Gong, Zhaohuan Li, Weili Jiang

机构 * School of Computing and Artificial Intelligence, Southwest Jiaotong University(计算机与人工智能学院,西南交通大学) Manufacturing Industry Chain Collaboration Industrial Software Key Laboratory of Sichuan Province(四川省制造业产业链协同工业软件重点实验室) School of Medicine, University of Electronic Science and Technology(医学院,电子科技大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出MCSDR模型,通过生成回归方法提升超声心动图视频中左室射血分数估计的准确性与解释性。

Comments 11 pages, 5 tables, 10 figures. Under peer review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07915 2026-02-10 cs.CV 57%

MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding

MARC: 用于高效视频理解的记忆增强强化学习令牌压缩

Peiran Wu, Zhuorui Yu, Yunze Liu, Chi-Hao Wu, Enmin Zhou, Junxiao Shen

机构 * University of Bristol(布里斯托大学) Memories.ai Research(Memories.ai研究)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 MARC通过结合结构化检索和强化学习知识蒸馏,实现高效视频理解,显著减少视觉令牌、GPU内存和延迟。

Comments Accepted at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07034 2026-02-10 cs.AI 57%

ST-Raptor: An Agentic System for Semi-Structured Table QA

ST-Raptor:一个用于半结构化表格问答的智能系统

Jinxiu Qu, Zirui Tang, Hongzhang Huang, Boyu Niu, Wei Zhou, Jiannan Wang, Yitong Song, Guoliang Li, Xuanhe Zhou, Fan Wu

机构 * Tsinghua University(清华大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

AI总结 ST-Raptor通过结合视觉编辑、树形结构建模和代理驱动查询解决,提升半结构化表格问答的准确性和易用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06997 2026-02-10 eess.SP cs.AI cs.LG cs.NE 57%

Adaptive Temporal Dynamics for Personalized Emotion Recognition: A Liquid Neural Network Approach

自适应时间动态用于个性化情绪识别:一种液态神经网络方法

Anindya Bhattacharjee, Nittya Ananda Biswas, K. A. Shahriar, Adib Rahman

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文提出了一种基于液态神经网络的多模态框架,用于提升基于EEG的情绪识别准确率,通过结合卷积特征提取、可学习时间常数和注意力引导融合,实现高准确率和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08090 2026-02-10 cs.HC 50%

A Collaborative Crowdsourcing Method for Designing External Interfaces for Autonomous Vehicles

为自动驾驶车辆设计外部接口的协作众包方法

Ronald Cumbal, Marcus Göransson, Alexandros Rouchitsas, Didem Gürdür Broo, Ginevra Castellano

专题命中 视频多模态 :multimodal(abstract)

AI总结 本文提出一种协作众包方法,通过结合大众创意、结构化原则和专家反馈,提升自动驾驶车辆接口设计的可扩展性和用户体验。

Comments Paper accepted for publication at the 2026 CHI Conference on Human Factors in Computing Systems (CHI'26), April 13-17, 2026, Barcelona, Spain

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07521 2026-02-10 cs.LG 50%

Pareto-guided Pipeline for Distilling Featherweight AI Agents in Mobile MOBA Games

基于帕累托最优的流水线:在移动MOBA游戏中蒸馏轻量级AI代理

Xionghui Yang, Bozhou Chen, Yunlong Lu, Yongyi Wang, Lingfeng Li, Lanxiao Huang, Lin Liu, Wenjun Wang, Meng Meng, Xia Lin, Wenxin Li

机构 * School of Computer Science Peking University Beijing China(计算机科学系 首都大学 北京 中国) TiMi L1 Studio Tencent Chengdu China(TiMi L1工作室 腾讯 成都 中国) Peking University(首都大学) Tencent(腾讯)

专题命中 视频多模态 :multi-modal(abstract)

AI总结 本文提出基于帕累托最优的流水线,设计高效学生架构搜索空间,在移动MOBA游戏中实现轻量级AI代理的蒸馏,提升推断速度与能效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20030 2026-02-10 cs.SI 50%

Counting How the Seconds Count: Understanding Algorithm-User Interplay in TikTok via ML-driven Analysis of Video Content

秒数如何计数:通过机器学习分析视频内容理解TikTok上的算法-用户互动

Maleeha Masood, Shreya Kannan, Zikun Liu, Deepak Vasisht, Indranil Gupta

专题命中 视频多模态 :multimodal(abstract)

AI总结 通过机器学习分析TikTok视频内容,研究算法与用户互动的时间演变及用户体验影响。

Comments To contact, email maleeha2@illinois.edu

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 跨模态检索 9 篇

2305.04195 2026-02-10 cs.CV cs.CL 84%

Cross-Modal Retrieval for Motion and Text via DropTriple Loss

通过DropTriple损失实现动与文本的跨模态检索

Sheng Yan, Yang Liu, Haoqiang Wang, Xin Du, Mengyuan Liu, Hong Liu

机构 * School of Artificial Intelligence, Chongqing University of Technology, China(重庆理工大学人工智能学院) College of Computer Science, Sichuan University, China(四川大学计算机学院) Key Laboratory of Machine Perception, Shenzhen Graduate School, Peking University, China(北京大学深圳研究生院机器感知重点实验室)

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

AI总结 本文提出DropTriple损失,用于提升人类动作与文本之间的跨模态检索性能,实验表明在HumanML3D数据集上实现了较高的检索准确率。

Comments This paper has been accepted by ACM MM Asia 2023 (Best Paper Candidate)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08099 2026-02-10 cs.CV cs.AI 84%

VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval

VidVec:解锁视频MLLM嵌入用于视频-文本检索

Issar Tzachor, Dvir Samuel, Rami Ben-Ari

专题命中 跨模态检索 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 VidVec通过利用预训练MLLM的中间层嵌入和校准头部,实现无需训练的视频-文本检索,超越现有方法,达到最佳性能。

Comments Project page: https://iyttor.github.io/VidVec/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07125 2026-02-10 cs.IR cs.AI cs.CV cs.LG 84%

Reasoning-Augmented Representations for Multimodal Retrieval

增强推理的表示用于多模态检索

Jianrui Zhang, Anirudh Sundara Rajan, Brandon Han, Soochahn Lee, Sukanta Ganguly, Yong Jae Lee

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Kookmin University(韩国高丽大学)

专题命中 跨模态检索 :multimodal(title,abstract);any-to-any(abstract);分类 cs.CV、cs.AI

AI总结 本研究提出一种增强推理的多模态检索方法,通过外部化推理和语义密集表示提升检索性能,尤其在知识密集型查询和组合修改请求中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08342 2026-02-10 cs.CV cs.AI 81%

UrbanGraphEmbeddings: Learning and Evaluating Spatially Grounded Multimodal Embeddings for Urban Science

UrbanGraphEmbeddings: 学习和评估空间导向的多模态嵌入用于城市科学

Jie Zhang, Xingtong Yu, Yuan Fang, Rudi Stouffs, Zdravko Trivic

机构 * National University of Singapore Department of Architecture Singapore The Chinese University of Hong Kong Dept of Systems Eng. \& Eng. Mgmt. China Singapore Management University School of Computing \& Info. Systems Singapore National University of Singapore Department of Architecture The Chinese University of Hong Kong Singapore Management University

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出UGE框架,通过空间导向的多模态嵌入提升城市理解任务性能,实验显示在图像检索和地理位置排名上取得显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08741 2026-02-10 cs.CL 79%

From Rows to Reasoning: A Retrieval-Augmented Multimodal Framework for Spreadsheet Understanding

从行到推理:一种增强检索的多模态框架用于电子表格理解

Anmol Gulati, Sahil Sen, Waqar Sarguroh, Kevin Paul

机构 * Commercial Technology and Innovation Office, PricewaterhouseCoopers U.S.(普华永道美国商业技术与创新办公室)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

AI总结 FRTR提出了一种多模态检索增强生成框架,通过分解电子表格为细粒度嵌入并整合多模态信息,提升了对复杂电子表格的推理能力,在基准测试中实现了显著的准确率提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07642 2026-02-10 cs.AI cs.LG 79%

Efficient Table Retrieval and Understanding with Multimodal Large Language Models

基于多模态大语言模型的高效表格检索与理解

Zhuoyan Xu, Haoyang Fang, Boran Han, Bonan Min, Bernie Wang, Cuixiong Hu, Shuai Zhang

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) AWS(亚马逊网络服务)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 TabRAG通过多模态大语言模型实现高效表格检索与理解,显著提升检索召回率和答案准确率。

Comments Published at EACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07208 2026-02-10 cs.IR cs.AI 79%

Sequences as Nodes for Contrastive Multimodal Graph Recommendation

序列作为节点的对比多模态图推荐

Bucher Sahyouni, Matthew Vowels, Liqun Chen, Simon Hadfield

机构 * University of Surrey(萨里大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 MuSICRec通过多视图图方法结合协同、序列和多模态信号,提升推荐系统在冷启动和数据稀疏问题上的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08700 2026-02-10 cs.CL cs.HC cs.IR 57%

Do Images Clarify? A Study on the Effect of Images on Clarifying Questions in Conversational Search

图像能澄清问题吗?一种研究图像在会话搜索中澄清问题效果的探讨

Clemencia Siro, Zahra Abbasiantaeb, Yifei Yuan, Mohammad Aliannejadi, Maarten de Rijke

机构 * University of Amsterdam(阿姆斯特丹大学) University of Copenhagen(哥本哈根大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

AI总结 研究探讨图像在会话搜索中澄清问题的效果,发现多模态问题在回答澄清问题时更受青睐,但查询重述任务中效果更平衡,且图像影响因任务类型和用户专业知识而异。

Comments Accepted at CHIIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01317 2026-02-10 cs.HC cs.CY 50%

DietGlance: Dietary Monitoring and Personalized Analysis at a Glance with Knowledge-Empowered AI Assistant

DietGlance: 通过知识增强的AI助手实现饮食监控与个性化分析

Zhihan Jiang, Running Zhao, Lin Lin, Yue Yu, Handi Chen, Xinchen Zhang, Xuhai Xu, Yifang Wang, Xiaojuan Ma, Edith C. H. Ngai

专题命中 跨模态检索 :multimodal(abstract)

AI总结 DietGlance利用知识增强的AI助手,通过眼镜实现日常饮食行为的自动监控和个性化分析,提供营养分析和饮食建议。

Comments 47 pages, 14 figures. Accepted by ACM Transactions on Computing for Healthcare

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态生成 18 篇

2602.07993 2026-02-10 cs.CV cs.AI 84%

MCIE: Multimodal LLM-Driven Complex Instruction Image Editing with Spatial Guidance

MCIE: 多模态大语言模型驱动的复杂指令图像编辑与空间引导

Xuehai Bai, Xiaoling Gu, Akide Liu, Hangjie Yuan, YiFan Zhang, Jack Ma

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 MCIE通过空间引导和背景一致性模块,提升复杂指令图像编辑的指令合规性,实现23.96%的性能提升。

Comments Accepted by AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08249 2026-02-10 eess.IV cs.CV 79%

A Unified Framework for Multimodal Image Reconstruction and Synthesis using Denoising Diffusion Models

基于去噪扩散模型的多模态图像重建与合成统一框架

Weijie Gan, Xucheng Wang, Tongyao Wang, Wenshang Wang, Chunwei Ying, Yuyang Hu, Yasheng Chen, Hongyu An, Ulugbek S. Kamilov

机构 * Department of Computer Science and Engineering, Washington University in St. Louis(华盛顿大学圣路易斯分校计算机科学与工程系) Mallinckrodt Institute of Radiology, Washington University in St. Louis(华盛顿大学圣路易斯分校马林克罗德特放射医学研究所) Department of Electrical and Systems Engineering, Washington University in St. Louis(华盛顿大学圣路易斯分校电气与系统工程系) Department of Neurology, Washington University in St. Louis(华盛顿大学圣路易斯分校神经病学系) Department of Biomedical Engineering, Washington University in St. Louis(华盛顿大学圣路易斯分校生物医学工程系) Division of Biology and Biomedical Sciences, Washington University in St. Louis(华盛顿大学圣路易斯分校生物学与生物医学科学 division) Department of Electrical and Computer Engineering, University of Wisconsin–Madison(威斯康星大学麦迪逊分校电气与计算机工程系)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 Any2all通过统一框架实现多模态图像重建与合成,利用去噪扩散模型提升性能与质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09851 2026-02-10 cs.RO cs.CV 79%

Simultaneous Tactile-Visual Perception for Learning Multimodal Robot Manipulation

同时触觉-视觉感知用于学习多模态机器人操作

Yuyang Li, Yinghan Chen, Zihang Zhao, Puhao Li, Tengyu Liu, Siyuan Huang, Yixin Zhu

机构 * Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院) School of Psychological and Cognitive Sciences, Peking University(北京大学心理与认知科学学院) Beijing Key Lab of Behavior and Mental Health, Peking University(北京大学行为与心理健康北京市重点实验室) Beijing Institute for General Artificial Intelligence(北京一般人工智能研究院) State Key Lab for General Artificial Intelligence(一般人工智能国家重点实验室) Embodied Intelligence Lab, PKU-Wuhan Institute for Artificial Intelligence(具身智能实验室,北京大学武汉人工智能研究院) Department of Computer Science and Technology, University of Cambridge(剑桥大学计算机科学与技术系)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 TacThru-UMI通过结合同时触觉-视觉感知与现代学习框架,实现了高精度多模态机器人操作。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23278 2026-02-10 cs.CV 79%

UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing

UniLiP: 适配CLIP以实现统一的多模态理解、生成与编辑

Hao Tang, Chenwei Xie, Xiaoyi Bao, Tingyu Weng, Pandeng Li, Yun Zheng, Liwei Wang

机构 * Center for Data Science, Peking University(北京大学数据科学中心) Alibaba Group(阿里巴巴集团) CASIA Center for Machine Learning Research, Peking University(北京大学机器学习研究中心) State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 UniLIP通过两阶段训练和双条件架构,提升CLIP在多模态理解、生成和编辑任务中的性能,实现高效参数下的高表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25682 2026-02-10 cs.CL 79%

PairUni: Pairwise Training for Unified Multimodal Language Models

PairUni: 用于统一多模态语言模型的成对训练

Jiani Zheng, Zhiyang Teng, Kunpeng Qiu, Xiangtai Li, Anran Wang, Yu Tian, Ye Tian, Haochen Wang, Zhuochen Wang

机构 * ByteDance(字节跳动)

专题命中 多模态生成 :multimodal(title);cross-modal(abstract);分类 cs.CL

AI总结 PairUni通过成对训练提升统一多模态语言模型的理解与生成能力,采用配对数据集和PairGRPO算法实现更有效的策略学习。

Comments 22 pages, 11 figures, and 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08830 2026-02-10 cs.HC 78%

Enhancing Generative AI Image Refinement with Scribbles and Annotations: A Comparative Study of Multimodal Prompts

通过草图和注释增强生成式AI图像细化:多模态提示的比较研究

Hyerim Park, Phuong Thao Tran, Andre Luckow, Ceenu George, Michael Sedlmair, Malin Eiband

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 本研究通过比较多模态提示,探讨草图和注释如何提升生成式AI图像细化,提出原型并揭示设计师在多模态策略中的偏好。

Comments 22 pages, 14 figures. Preprint of an accepted IUI '26 paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04780 2026-02-10 cs.LG cond-mat.dis-nn 78%

Dynamical Regimes of Multimodal Diffusion Models

多模态扩散模型的动力学区域

Emil Albrychiewicz, Andrés Franco Valiente, Li-Ching Chen

机构 * Leinweber Institute for Theoretical Physics(莱因韦伯理论物理研究所) Department of Physics, University of California, Berkeley, CA, 94720-7300, USA(加州大学伯克利分校物理系) Theoretical Physics Group, Lawrence Berkeley National Laboratory(劳伦斯伯克利国家实验室理论物理组) Department of Radiation Oncology, University of California, San Francisco(加州大学旧金山分校放射肿瘤科) Computational Precision Health, University of California San Francisco(加州大学旧金山分校计算精准健康)

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 多模态扩散模型通过耦合动态机制,揭示生成过程中的时间层次结构及同步间隙现象,为生成模型的稳定性提供理论框架。

Comments 40 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08145 2026-02-10 cs.LG cs.AI cs.CL cs.CV cs.CY 67%

Reliable and Responsible Foundation Models: A Comprehensive Survey

可靠且负责任的基础模型:全面综述

Xinyu Yang, Junlin Han, Rishi Bommasani, Jinqi Luo, Wenjie Qu, Wangchunshu Zhou, Adel Bibi, Xiyao Wang, Jaehong Yoon, Elias Stengel-Eskin, Shengbang Tong, Lingfeng Shen, Rafael Rafailov, Runjia Li, Zhaoyang Wang, Yiyang Zhou, Chenhang Cui, Yu Wang, Wenhao Zheng, Huichi Zhou, Jindong Gu, Zhaorun Chen, Peng Xia, Tony Lee, Thomas Zollo, Vikash Sehwag, Jixuan Leng, Jiuhai Chen, Yuxin Wen, Huan Zhang, Zhun Deng, Linjun Zhang, Pavel Izmailov, Pang Wei Koh, Yulia Tsvetkov, Andrew Wilson, Jiaheng Zhang, James Zou, Cihang Xie, Hao Wang, Philip Torr, Julian McAuley, David Alvarez-Melis, Florian Tramèr, Kaidi Xu, Suman Jana, Chris Callison-Burch, Rene Vidal, Filippos Kokkinos, Mohit Bansal, Beidi Chen, Huaxiu Yao

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文综述了基础模型的可靠和负责任发展,探讨了偏见、安全、不确定性等关键问题,并提出了未来研究方向。

Comments TMLR camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07983 2026-02-10 cs.AI cs.CL 62%

Accelerating Social Science Research via Agentic Hypothesization and Experimentation

通过代理假设和实验加速社会科学研究

Jishu Sen Gupta, Harini SI, Somesh Kumar Singh, Syed Mohamad Tawseeq, Yaman Kumar Singla, David Doermann, Rajiv Ratn Shah, Balaji Krishnamurthy

机构 * Adobe Media and Data Science Research (MDSR)(Adobe媒体与数据科学研究所)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 EXPERIGEN通过代理框架实现端到端的科学发现,发现更多显著且预测性强的假设,并通过A/B测试验证其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08785 2026-02-10 cs.CV cs.AI 62%

DeltaSpace: A Semantic-aligned Feature Space for Flexible Text-guided Image Editing

DeltaSpace: 一种语义对齐的特征空间用于灵活的文本引导图像编辑

Yueming Lyu, Kang Zhao, Bo Peng, Huafeng Chen, Yue Jiang, Yingya Zhang, Jing Dong, Caifeng Shan

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI

AI总结 DeltaSpace通过语义对齐的特征空间实现文本引导图像编辑的灵活训练和推理,支持零样本推理和无需文本的训练。

Comments 18 pages. arXiv admin note: text overlap with arXiv:2303.06285

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07176 2026-02-10 cs.CL cs.AI cs.ET cs.HC 62%

Open TutorAI: An Open-source Platform for Personalized and Immersive Learning with Generative AI

Open TutorAI: 一个基于生成AI的开源平台,用于个性化和沉浸式学习

Mohamed El Hajji, Tarek Ait Baha, Aicha Dakir, Hammou Fadili, Youssef Es-Saady

机构 * IRF-SIC Laboratory, Ibnou Zohr University(IRF-SIC实验室,伊本·扎赫尔大学) Regional Center for Education(教育与培训专业地区中心) Polydisciplinary Faculty of Taroudant, Ibnou Zohr University(塔鲁旦多学科学院,伊本·扎赫尔大学) Higher School of Technology of Guelmim, Ibnou Zohr University(盖尔米姆技术高等学校,伊本·扎赫尔大学) Paragraphe laboratory, Paris 8 and CY Cergy Paris Universities(Paragraphe实验室,巴黎8大学和CY塞克巴黎大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 Open TutorAI 是一个基于生成AI的开源平台,通过个性化和沉浸式学习体验提升教育效果。

Comments 19 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06133 2026-02-10 cs.LG cs.AI cs.RO 57%

A Review of Online Diffusion Policy RL Algorithms for Scalable Robotic Control

在线扩散策略强化学习算法综述:可扩展机器人控制

Wonhyeok Choi, Shutong Ding, Minwoo Choi, Jungwan Woo, Kyumin Hwang, Jaeyeul Kim, Ye Shi, Sunghoon Im

机构 * Daegu Gyeongbuk Institute of Science and Technology(大邱庆尚科学技术院) ShanghaiTech University(上海科技大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 本文综述了在线扩散策略强化学习算法,分析了其在可扩展机器人控制中的性能、挑战及未来发展方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09777 2026-02-10 cs.CV cs.RO 57%

EgoFSD: Ego-Centric Fully Sparse Paradigm with Uncertainty Denoising and Iterative Refinement for Efficient End-to-End Self-Driving

EgoFSD:面向端到端自动驾驶的以自我为中心的完全稀疏范式,结合不确定性去噪和迭代细化

Haisheng Su, Wei Wu, Zhenjie Yang, Isabel Guan

机构 * School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院) SenseAuto The Hong Kong University of Science and Technology(香港理工大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 EgoFSD通过引入稀疏感知、分层交互和迭代运动规划,提升端到端自动驾驶的效率和性能,减少误差和碰撞,提高训练稳定性。

Comments Accepted to ICRA2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07564 2026-02-10 cs.CV 57%

SIGMA: Selective-Interleaved Generation with Multi-Attribute Tokens

SIGMA: 带多属性标记的选择性交错生成

Xiaoyan Zhang, Zechen Bai, Haofan Wang, Yiren Song

机构 * Creatly AI University of Michigan(密歇根大学) Lovart AI National University of Singapore(新加坡国立大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 SIGMA通过引入多属性标记实现多条件生成,提升可控性、一致性和视觉质量,优于Bagel在组合任务中的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏