arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

2026-01-30 至 2026-01-30 共收录 55 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 具身导航 10 篇

2507.22188 2026-01-30 cs.RO cs.SY eess.SY 61%

Transport and Delivery of Objects with a Soft Everting Robot

具有软翻转机器人的物体运输与交付

Ethan DeVries, Jack Ferlazzo, Mustafa Ugur, Laura H. Blumenschein

机构 * School of Mechanical Engineering, Purdue University(机械工程学院,普渡大学)

专题命中 具身导航 :navigation(abstract);分类 cs.RO;robotics(comments)

AI总结 本研究提出了一种利用软翻转机器人通过内部运输和部署各类载荷的方法,展示了其在危险环境中的运输能力和灵活性。

Comments 8 pages, 11 figures, Published in IEEE Robotics and Automation Letters Link to publication: https://ieeexplore.ieee.org/abstract/document/11359009 Citation: E. M. DeVries, J. Ferlazzo, M. Ugur and L. H. Blumenschein, "Transport and Delivery of Objects With a Soft Everting Robot," in IEEE Robotics and Automation Letters, vol. 11, no. 3, pp. 2935-2942, March 2026, doi: 10.1109/LRA.2026.3655537

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21976 2026-01-30 cs.RO physics.app-ph 57%

Macro-Scale Electrostatic Origami Motor

宏观尺度静电折纸电机

Alex S. Miller, Leo McElroy, Jeffrey H. Lang

机构 * Dept. of Aeronautics and Astronautics(航空与宇航系) Massachusetts Institute of Technology(麻省理工学院) Independent Researcher(独立研究者) Department of Electrical Eng. and Computer Sci.(电气工程与计算机科学系)

专题命中 具身导航 :robotics(abstract);分类 cs.RO

AI总结 本文提出了一种可折叠的宏观尺度静电折纸电机,通过电晕放电实现连续旋转运动,展示了2.5:1的扩展比、1440 rpm的速度和0.04 Nm/kg的扭矩密度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13421 2026-01-30 cs.AI cs.ET 57%

Virtuous Machines: Towards Artificial General Science

美德机器:迈向人工智能通用科学

Gabrielle Wehr, Reuben Rideaux, Amaya J. Fox, David R. Lightfoot, Jason Tangen, Jason B. Mattingley, Shane E. Ehrhardt

机构 * Explore Science School of Psychology(心理学学院) Queensland Brain Institute(昆士兰脑研究所) Canadian Institute of Advanced Research (CIFAR)(加拿大高级研究 institute)

专题命中 具身导航 :embodied AI(abstract);分类 cs.AI

AI总结 本研究提出了一种无需领域知识的AI科学家系统,能自主完成科学流程,通过实验展示了其在心理学研究中的能力,推动了人工智能在科学发现中的应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12795 2026-01-30 cs.DL 50%

A Visual Approach for Health Information Exploration: Adaptive Levels of Visual Granularity and Interaction Analysis

一种健康信息探索的视觉方法:自适应的视觉粒度层次和交互分析

Stefan Lengauer, Lin Shao, Hossein Miri, Michael Bedek, Cordula Kupfer, Maria Zangl, Bettina Kubicek, Barbara Dienstbier, Klaus Jeitler, Cornelia Krenn, Thomas Semlitsch, Carolin Zipp, Dietrich Albert, Andrea Siebenhofer, Tobias Schreck

专题命中 具身导航 :navigation(abstract)

AI总结 本文提出了一种基于自适应文档可视化的健康信息探索系统,通过不同粒度的视觉展示和交互溯源分析,提升用户对健康信息的理解与利用。

Comments 19 pages, 7 figures, journal preprint

Journal ref JUCS - Journal of Universal Computer Science 32 (2026) 4-32

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 具身推理 5 篇

2506.06725 2026-01-30 cs.AI cs.LG 81%

WorldLLM: Improving LLMs' world modeling using curiosity-driven theory-making

WorldLLM: 通过好奇心驱动的理论构建改进LLM的世界建模

Guillaume Levy, Cedric Colas, Pierre-Yves Oudeyer, Thomas Carta, Clement Romac

机构 * Inria(法国国家信息与自动化技术研究院) Univ. of Bordeaux(波尔多大学) MIT(麻省理工学院) Hugging Face

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.LG

AI总结 WorldLLM通过结合贝叶斯推断和好奇心驱动的强化学习,改进LLM在结构化环境中的世界建模能力,提升预测精度并生成可解释的环境理论。

Comments This project's code can be found at https://github.com/flowersteam/WorldLLM. This project was presented at RLDM 2025 (https://rldm.org/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22128 2026-01-30 cs.AI cs.CE q-bio.QM 79%

The Patient is not a Moving Document: A World Model Training Paradigm for Longitudinal EHR

患者并非静态文档:一种用于纵向电子健康记录的world model训练范式

Irsyad Adam, Zekai Chen, David Laprade, Shaun Porwal, David Laub, Erik Reinertsen, Arda Pekis, Kevin Brown

专题命中 具身推理 :world model(title,abstract);分类 cs.AI

AI总结 本文提出SMB-Structure模型,通过结合联合嵌入预测架构和next-token预测,学习捕捉疾病动态的嵌入,实现对高异质性复杂任务的竞争力表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22086 2026-01-30 physics.flu-dyn cs.CV 79%

Learning Transient Convective Heat Transfer with Geometry Aware World Models

利用几何感知的世界模型学习瞬态对流热传递

Onur T. Doganay, Alexander Klawonn, Martin Eigel, Hanno Gottschalk

机构 * Institute of Mathematics, TU Berlin(柏林技术大学数学研究所) Siemens Energy AG(西门子能源有限公司) Weierstrass Institute for Applied Analysis and Stochastics(魏尔斯特拉斯应用分析与概率研究所)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 本文提出一种几何感知的世界模型架构,用于学习瞬态对流热传递,通过双重条件机制和架构适应性提升物理模拟的可控性和精度。

Comments 36 pages, 18 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17323 2026-01-30 cs.CV 57%

SkyReels-V3 Technique Report

SkyReels-V3 技术报告

Debang Li, Zhengcong Fei, Tuanhui Li, Yikun Dou, Zheng Chen, Jiangping Yang, Mingyuan Fan, Jingtao Xu, Jiahua Wang, Baoxuan Gu, Mingshan Chang, Wenjing Cai, Yuqiang Xie, Binjie Mao, Youqiang Zhang, Nuo Pang, Hao Zhang, Yuzhe Jin, Zhiheng Xu, Dixuan Lin, Guibin Chen, Yahui Zhou

机构 * SkyReels Team(SkyReels 团队)

专题命中 具身推理 :world model(abstract);分类 cs.CV

AI总结 SkyReels-V3 提出三种视频生成范式,通过多模态上下文学习框架实现高质量视频生成,涵盖图像到视频、视频扩展和音频引导生成,达到接近闭源系统的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.03762 2026-01-30 cs.CL cs.AI 57%

Brain in a Vat: On Missing Pieces Towards Artificial General Intelligence in Large Language Models

虚拟大脑:迈向大型语言模型中人工通用智能的缺失部件

Yuxi Ma, Chi Zhang, Song-Chun Zhu

机构 * Beijing Insitute for General Artificial Intelligence (BIGAI)(北京通用人工智能研究院)

专题命中 具身推理 :world model(abstract);分类 cs.AI

AI总结 本文探讨了大型语言模型中人工通用智能的缺失部件,提出智能代理应具备无限任务执行、任务生成、价值系统和现实世界模型等特征,并强调知与行的统一对智能发展的重要性。

Journal ref Proceedings of the Annual Meeting of the Cognitive Science Society, 47 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 机器人基础模型 2 篇

2511.19859 2026-01-30 cs.RO 83%

Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought for Robotic Action Generation

统一感知与行动:一种融合模态的管道,具有隐式视觉推理链的机器人行动生成

Xiangkai Ma, Lekai Xing, Han Zhang, Wenzhong Li, Sanglu Lu

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学)

专题命中 机器人基础模型 :robotic(title,abstract);manipulation(abstract);分类 cs.RO

AI总结 VITA通过融合视觉和动作的共享潜在空间,实现机器人行动生成的统一感知与行动,提升了多个任务的成功率和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20966 2026-01-30 cs.RO cs.AI 73%

Parallels Between VLA Model Post-Training and Human Motor Learning: Progress, Challenges, and Trends

VLA模型后训练与人类运动学习的类比:进展、挑战与趋势

Tian-Yu Xiang, Ao-Qun Jin, Xiao-Hu Zhou, Mei-Jiang Gui, Xiao-Liang Xie, Shi-Qi Liu, Shuang-Yi Wang, Sheng-Bin Duan, Fu-Chao Xie, Wen-Kai Wang, Si-Cheng Wang, Ling-Yun Li, Tian Tu, Zeng-Guang Hou

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) The Grainger College of Engineering, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校格拉inger工程学院) CAS Center for Excellence in Brain Science and Intelligence Technology(中国科学院脑科学与智能技术卓越创新中心) Joint Laboratory of Intelligence Science and Technology, Institute of Systems Engineering, Macau University of Science and Technology(澳门科技大学系统工程学院智能科学与技术联合实验室)

专题命中 机器人基础模型 :manipulation(abstract);robotic(abstract);分类 cs.RO、cs.AI

AI总结 本文从人类运动学习角度综述了VLA模型后训练的进展、挑战与趋势,提出四类后训练方法并探讨了其在机器人操作中的应用与未来方向。

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 模仿学习与强化学习 4 篇

2601.21548 2026-01-30 cs.RO cs.AI cs.ET 62%

Training slow silicon neurons to control extremely fast robots with spiking reinforcement learning

训练慢硅神经元以通过脉冲强化学习控制极快的机器人

Irene Ambrosini, Ingo Blakowski, Dmitrii Zendrikov, Cristiano Capone, Luna Gava, Giacomo Indiveri, Chiara De Luca, Chiara Bartolozzi

机构 * Institute of Neuroinformatics, UZH and ETH Zurich(神经信息研究所,苏黎世联邦理工学院和苏黎世联邦理工学院) Istituto Italiano di Tecnologia(意大利技术研究所) Technical University of Munich(慕尼黑技术大学) Natl. Center for Radiation Protection and Computational Physics, Istituto Superiore di Sanità(辐射防护与计算物理国家中心,意大利卫生超级研究所) Digital Society Initiative, University of Zurich(数字社会倡议,苏黎世大学)

专题命中 模仿学习与强化学习 :robotic(abstract);分类 cs.RO、cs.AI

AI总结 通过脉冲强化学习训练慢硅神经元,实现高速机器人控制,展示脑启发方法在快速互动任务中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00564 2026-01-30 cs.LG cs.AI 62%

Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining

通过联合优化的世界-动作模型扩展离线模型基于的强化学习

Jie Cheng, Ruixi Qiao, Yingwei Ma, Binhua Li, Gang Xiong, Qinghai Miao, Yongbin Li, Yisheng Lv

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Alibaba Group(阿里巴巴集团)

专题命中 模仿学习与强化学习 :world model(abstract);分类 cs.AI、cs.LG

AI总结 JOWA通过联合优化的世界-动作模型扩展离线RL,实现高效泛化和高性能

Comments Accepted by ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09797 2026-01-30 cs.RO 57%

FLARE: Agile Flights for Quadrotor Cable-Suspended Payload System via Reinforcement Learning

FLARE:通过强化学习实现四旋翼缆悬载具系统的敏捷飞行

Dongcheng Cao, Jin Zhou, Xian Wang, Shuo Li

机构 * College of Control Science and Engineering, Zhejiang University(控制科学与工程学院,浙江大学)

专题命中 模仿学习与强化学习 :navigation(abstract);分类 cs.RO

AI总结 FLARE通过强化学习实现四旋翼缆悬载具系统的敏捷飞行,相比传统方法提升了3倍的执行速度,并实现了高效的仿真到现实转移。

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.06284 2026-01-30 cs.AI 57%

CURIOUS: Intrinsically Motivated Modular Multi-Goal Reinforcement Learning

CURIOUS:内在动机的模块化多目标强化学习

Cédric Colas, Pierre Fournier, Olivier Sigaud, Mohamed Chetouani, Pierre-Yves Oudeyer

机构 * Flowers Team, Inria and Ensta ParisTech(Inria和Ensta巴黎科技大学)

专题命中 模仿学习与强化学习 :robotic(abstract);分类 cs.AI

AI总结 CURIOUS通过内在动机探索和自动课程学习机制,实现模块化多目标强化学习的自组织发展。

Comments Accepted at ICML 2019 https://github.com/flowersteam/curious

Journal ref Proceedings of the 36th International Conference on Machine Learning 2019

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 机器人数据与评测 8 篇

2601.21713 2026-01-30 cs.RO cs.AI 84%

Disentangling perception and reasoning for improving data efficiency in learning cloth manipulation without demonstrations

解构感知与推理以提高无演示学习布料操控的数据效率

Donatien Delehelle, Fei Chen, Darwin Caldwell

机构 * Advanced Robotics, Istituto Italiano di Tecnologia (IIT)(先进机器人技术研究所(IIT)) epartment of Mechanical and Automation Engineering, T-Stone Robotics Institute, The Chinese University of Hong Kong(机械与自动化工程系,腾讯机器人研究院,香港中文大学) University of Genova(热那亚大学)

专题命中 机器人数据与评测 :manipulation(title,abstract);robotics(abstract);分类 cs.RO、cs.AI

AI总结 本文提出了一种高效且模块化的强化学习方法,通过减少模型大小和训练时间来提高布料操控学习的数据效率,并在现实世界中实现有效迁移。

Comments 6 pages, 4 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21199 2026-01-30 cs.CV cs.AI 73%

Thinker: A vision-language foundation model for embodied intelligence

Thinker:一个用于具身智能的视觉-语言基础模型

Baiyu Pan, Daqin Luo, Junpeng Yang, Jiyuan Wang, Yixuan Zhang, Hailin Shi, Jichao Jiao

专题命中 机器人数据与评测 :robotics(abstract);robotic(abstract);分类 cs.AI、cs.CV

AI总结 Thinker通过构建大规模数据集和改进输入方式,在机器人感知与推理任务中实现了最先进的性能。

Comments IROS 2025, 4 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21222 2026-01-30 cs.AR 67%

FireFly-P: FPGA-Accelerated Spiking Neural Network Plasticity for Robust Adaptive Control

FireFly-P:FPGA加速的脉冲神经网络可塑性用于鲁棒自适应控制

Tenglong Li, Jindong Li, Guobin Shen, Dongcheng Zhao, Qian Zhang, Yi Zeng

专题命中 机器人数据与评测 :robotics(abstract);robotic(abstract)

AI总结 FireFly-P是一种基于FPGA的脉冲神经网络可塑性加速器,通过高效硬件设计实现低功耗、低延迟的自适应控制

Comments 5 pages, 4 figures. Accepted for lecture presentation at the 2026 IEEE International Symposium on Circuits and Systems (ISCAS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02507 2026-01-30 cs.CV cs.RO 62%

Keeping it Local, Tiny and Real: Automated Report Generation on Edge Computing Devices for Mechatronic-Based Cognitive Systems

保持本地化、小巧和现实:面向机电认知系统的边缘计算设备自动化报告生成

Nicolas Schuler, Lea Dewald, Jürgen Graf

专题命中 机器人数据与评测 :robotics(abstract);分类 cs.RO、cs.CV

AI总结 本文提出了一种基于边缘计算设备的自动化报告生成方法,利用本地模型实现多模态传感器数据处理,以提升机电认知系统在不同环境中的评估效率与隐私保护。

Comments 6 pages, 4 figures, 1 table; accepted for MECATRONICS-REM 2025 International Conference, PARIS, FRANCE December 3-5 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21360 2026-01-30 cs.CL cs.AI cs.ET cs.LG cs.SE 62%

The Compliance Paradox: Semantic-Instruction Decoupling in Automated Academic Code Evaluation

合规悖论:自动学术代码评估中的语义指令解耦

Devanshu Sahoo, Manish Prasad, Vasudev Majhi, Arjun Neekhra, Yash Sinha, Murari Mandal, Vinay Chamola, Dhruv Kumar

机构 * KIIT University(KIIT大学)

专题命中 机器人数据与评测 :manipulation(abstract);分类 cs.AI、cs.LG

AI总结 本文揭示了自动学术代码评估中模型因优先考虑隐藏格式约束而无法正确评估代码的问题,提出SPACI和AST-ASIP框架以检测和对抗此类解耦现象。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22027 2026-01-30 cs.AI 57%

CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty

CAR-bench: 评估在现实世界不确定性下LLM代理的一致性和限意识

Johannes Kirmayr, Lukas Stappen, Elisabeth André

机构 * BMW Group Research and Technology(宝马集团研究与技术) Augsburg University(艾希施泰特大学)

专题命中 机器人数据与评测 :navigation(abstract);分类 cs.AI

AI总结 CAR-bench旨在评估LLM代理在现实世界不确定性下的一致性、不确定性处理和能力意识,通过多轮对话和工具使用测试其可靠性和自我意识。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25731 2026-01-30 cs.CV 57%

LaTo: Landmark-tokenized Diffusion Transformer for Fine-grained Human Face Editing

LaTo:基于地标token化的扩散变换器用于精细的人脸编辑

Zhenghao Zhang, Ziying Zhang, Junchao Liao, Xiangyu Meng, Qiang Hu, Siyu Zhu, Xiaoyun Zhang, Long Qin, Weizhi Wang

机构 * Alibaba Cloud Computing(阿里云计算) Shanghai Jiao Tong University(上海交通大学) Fudan University(复旦大学)

专题命中 机器人数据与评测 :manipulation(abstract);分类 cs.CV

AI总结 LaTo通过地标token化扩散变换器实现精细的人脸编辑,提升身份保持与语义一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02250 2026-01-30 cs.HC cs.RO 57%

Designing Effective Human-Swarm Interaction Interfaces: Insights from a User Study on Task Performance

设计高效的人-群体交互界面:基于任务表现的用户研究洞察

Wasura D. Wattearachchi, Erandi Lakshika, Kathryn Kasmarik, Michael Barlow

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.RO

AI总结 本文提出了一种系统方法设计人-群体交互界面,通过用户研究验证了其在目标搜索任务中对机器人群的引导效果,尤其在移动危险情况下表现优异。

Comments 8 pages, 4 figures, 5 tables

Journal ref 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Vienna, Austria, 2025, pp. 3386-3393

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 其他机器人 2 篇

2601.21188 2026-01-30 cs.RO 79%

Disturbance-Aware Flight Control of Robotic Gliding Blimp via Moving Mass Actuation

基于移动质量作动的扰动感知机器人滑翔飞艇飞行控制

Hao Cheng, Feitian Zhang

机构 * Robotics and Control Laboratory, School of Advanced Manufacturing and Robotics, and the State Key Laboratory of Turbulence and Complex Systems, Peking University(机器人与控制实验室,先进制造与机器人学院,湍流与复杂系统国家重点实验室,北京大学)

专题命中 其他机器人 :robotic(title,abstract);分类 cs.RO

AI总结 本文提出基于移动质量作动的扰动感知飞行控制方法,通过MHE-MPC框架提升飞艇在风扰动下的稳定性和控制性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21011 2026-01-30 cs.RO cs.MA cs.OS cs.SE 79%

Meta-ROS: A Next-Generation Middleware Architecture for Adaptive and Scalable Robotic Systems

Meta-ROS:面向适应性和可扩展性机器人系统的下一代中间件架构

Anshul Ranjan, Anoosh Damodar, Neha Chougule, Dhruva S Nayak, Anantharaman P. N, Shylaja S S

机构 * PES University(PES大学)

专题命中 其他机器人 :robotic(title);robotics(abstract);分类 cs.RO

AI总结 Meta-ROS通过高效通信协议和跨平台兼容性,提升机器人系统性能与开发效率,适用于实时AI应用。

Comments Checkout the Python Library - https://pypi.org/project/metaros/ To be Submitted in ACM Transactions on Autonomous and Adaptive Systems (TAAS) Journal

详情

展开后加载摘要…

URL PDF HTML 收藏