arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

2026-01-06 至 2026-01-06 共收录 65 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 机器人操作 20 篇

2601.01618 2026-01-06 cs.RO 88%

Action-Sketcher: From Reasoning to Action via Visual Sketches for Long-Horizon Robotic Manipulation

动作草图生成器:通过视觉草图实现从推理到动作的长周期机器人操作

Huajie Tan, Peterson Co, Yijie Xu, Shanyu Rong, Yuheng Ji, Cheng Chi, Xiansheng Chen, Qiongyu Zhang, Zhongxia Zhao, Pengwei Wang, Zhongyuan Wang, Shanghang Zhang

机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(1 多媒体信息处理国家重点实验室,计算机科学学院,北京大学) Beijing Academy of Artificial Intelligence(2 北京人工智能研究院) University of Sydney(3 新南威尔士大学) Institute of Automation, Chinese Academy of Sciences(4 中国科学院自动化研究所)

专题命中 机器人操作 :manipulation(title,abstract);robotic(title,abstract);分类 cs.RO

AI总结 动作草图生成器通过视觉草图实现从推理到动作的长周期机器人操作,结合视觉-语言-动作框架提升任务执行的鲁棒性和可解释性。

Comments 26 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01948 2026-01-06 cs.RO 85%

Learning Diffusion Policy from Primitive Skills for Robot Manipulation

从基本技能学习扩散策略用于机器人操作

Zhihao Gu, Ming Yang, Difan Zou, Dong Xu

机构 * Dong Xu(东旭教授)

专题命中 机器人操作 :manipulation(title,abstract);robot learning(abstract);robotic(abstract);分类 cs.RO

AI总结 本文提出SDP,一种基于技能的扩散策略,通过整合可解释的技能学习与条件动作规划,提升机器人操作中技能一致性与性能。

Comments Accepted to AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01438 2026-01-06 cs.RO cs.AI 84%

Online Estimation and Manipulation of Articulated Objects

在线估计和操控关节化物体

Russell Buchanan, Adrian Röfer, João Moura, Abhinav Valada, Sethu Vijayakumar

机构 * RIPL-Lab, University of Waterloo, Canada(滑铁卢大学) Robot Learning Lab, University of Freiburg, Germany(弗赖堡大学) School of Informatics, University of Edinburgh, UK(爱丁堡大学)

专题命中 机器人操作 :manipulation(title,abstract);robotic(abstract);分类 cs.RO、cs.AI

AI总结 本文提出了一种结合视觉先验和本体感觉传感的在线估计方法,用于机器人自主操控未知的关节化物体。

Comments This preprint has not undergone peer review or any post-submission improvements or corrections. The Version of Record of this article is published in Autonomous Robots, and is available online at [Link will be updated when available]

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01651 2026-01-06 cs.RO 83%

DemoBot: Efficient Learning of Bimanual Manipulation with Dexterous Hands From Third-Person Human Videos

DemoBot: 从第三人称人类视频高效学习双臂操作与灵巧手

Yucheng Xu, Xiaofeng Mao, Elle Miller, Xinyu Yi, Yang Li, Zhibin Li, Robert B. Fisher

机构 * ByteDance Seed(字节跳动种子)

专题命中 机器人操作 :manipulation(title,abstract);robotic(abstract);分类 cs.RO

AI总结 DemoBot通过强化学习框架从人类视频中高效学习双臂操作技能,解决长周期任务的挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06240 2026-01-06 cs.RO cs.AI 81%

Affordance-Guided Coarse-to-Fine Exploration for Base Placement in Open-Vocabulary Mobile Manipulation

基于 affordance 的粗到细探索用于开放词汇移动操控中的基座放置

Tzu-Jung Lin, Jia-Fong Yeh, Hung-Ting Su, Chung-Yi Lin, Yi-Ting Chen, Winston H. Hsu

专题命中 机器人操作 :manipulation(title,abstract);分类 cs.RO、cs.AI

AI总结 本文提出了一种基于 affordance 的粗到细探索方法,通过结合视觉语言模型的语义理解和几何可行性,提升开放词汇移动操控中基座放置的成功率。

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01869 2026-01-06 cs.DS 78%

Exact Clique Number Manipulation via Edge Interdiction

通过边拦截精确操纵 clique 数量

Yi Zhou, Haoyu Jiang, Chenghao Zhu, André Rossi

专题命中 机器人操作 :manipulation(title,abstract)

AI总结 本文提出了一种两阶段精确算法 RLCM,通过将边拦截 clique 问题转化为参数化边阻塞 clique 问题,有效解决了大规模图中最大 clique 的计算问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01675 2026-01-06 cs.RO 65%

VisuoTactile 6D Pose Estimation of an In-Hand Object using Vision and Tactile Sensor Data

利用视觉和触觉传感器数据进行手持物体的6D位姿估计

Snehal s. Dikhale, Karankumar Patel, Daksh Dhingra, Itoshi Naramura, Akinobu Hayashi, Soshi Iba, Nawid Jamali

机构 * Honda Research Institute USA, Inc.(本田美国研究院) Honda R&D Co., Ltd.(本田研发公司) Department of Mechanical Engineering, University of Washington(华盛顿大学机械工程系)

专题命中 机器人操作 :manipulation(abstract);robotics(comments,journal_ref);分类 cs.RO

AI总结 本文提出利用视觉和触觉数据融合方法,提高机器人在手物体的6D位姿估计精度。

Comments Accepted for publication in IEEE Robotics and Automation Letters (RA-L), January 2022. Presented at ICRA 2022. This is the author's version of the manuscript

Journal ref IEEE Robotics and Automation Letters, vol. 7, no. 2, pp. 2228-2235, April 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01915 2026-01-06 cs.CV 57%

TalkPhoto: A Versatile Training-Free Conversational Assistant for Intelligent Image Editing

TalkPhoto: 一种多功能的无需训练的对话式智能图像编辑助手

Yujie Hu, Zecheng Tang, Xu Jiang, Weiqi Li, Jian Zhang

机构 * School of Electronic and Computer Engineering, Peking University(电子与计算机工程学院,北京大学)

专题命中 机器人操作 :manipulation(abstract);分类 cs.CV

AI总结 TalkPhoto通过对话交互实现无需训练的多功能图像编辑,利用提示模板分层调用现有高级编辑方法,提升编辑精度与质量。

Comments a Conversational Assistant for Intelligent Image Editing

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15991 2026-01-06 quant-ph cs.LG 57%

Quantum Enhanced Anomaly Detection for ADS-B Data using Hybrid Deep Learning

量子增强的ADS-B数据异常检测:基于混合深度学习

Rani Naaman, Felipe Gohring de Magalhaes, Jean-Yves Ouattara, Gabriela Nicolescu

机构 * Department of computer and software engineering, Polytechnique Montreal(计算机与软件工程系,蒙特利尔大学)

专题命中 机器人操作 :manipulation(abstract);分类 cs.LG

AI总结 本文提出一种结合量子和经典机器学习的混合深度学习方法,用于ADS-B数据异常检测,实验表明其在检测准确率上与传统方法相当。

Comments This is the author's version of the work accepted for publication in the IEEE-AIAA Digital Avionics Systems Conference (DASC) 2025. The final version version is available via IEEE Xplore

Journal ref Proc. 44th IEEE/AIAA Digital Avionics Systems Conference (DASC), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01774 2026-01-06 cs.AI cs.CE cs.NA math.NA 57%

Can Large Language Models Solve Engineering Equations? A Systematic Comparison of Direct Prediction and Solver-Assisted Approaches

大型语言模型能解决工程方程吗?直接预测与求解器辅助方法的系统比较

Sai Varun Kodathala, Rakesh Vunnam

机构 * Research and Development Sports Vision, Inc.(Sports Vision公司研发部) Research and Development Vizworld, Inc.(Vizworld公司研发部)

专题命中 机器人操作 :manipulation(abstract);分类 cs.AI

AI总结 本文研究了大型语言模型在解决工程方程中的表现,发现求解器辅助方法在精度上显著优于直接预测,尤其在电子领域效果突出。

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04888 2026-01-06 cs.CV 57%

Training-Free Video Editing via Optical Flow-Enhanced Score Distillation

无需训练的视频编辑 via 光流增强的分数蒸馏

Lianghan Zhu, Yanqi Bao, Jing Huo, Jing Wu, Yu-Kun Lai, Wenbin Li, Yang Gao

专题命中 机器人操作 :manipulation(abstract);分类 cs.CV

AI总结 本研究提出基于预训练文本到视频模型的分数蒸馏方法,通过迭代优化和光流增强技术,解决无需训练视频编辑中的内容保留与时间连续性问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00867 2026-01-06 cs.CR cs.AI cs.CY cs.HC 57%

The Silicon Psyche: Anthropomorphic Vulnerabilities in Large Language Models

硅之心:大型语言模型中的人格化脆弱性

Giuseppe Canale, Kashyap Thimmaraju

机构 * CPF3.org(CPF3组织) Flowguard Institute(Flowguard研究所)

专题命中 机器人操作 :manipulation(abstract);分类 cs.AI

AI总结 本文提出LLMs继承人类心理脆弱性,通过心理测量框架揭示其对权威操纵等攻击的易受性,呼吁开发心理防火墙以保护AI安全。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24366 2026-01-06 quant-ph 50%

Spin vs. position conjugation in quantum simulations with atoms: application to quantum chemistry

自旋与位置在原子量子模拟中的结合:应用于量子化学

N. A. Moroz, K. S. Tikhonov, L. V. Gerasimov, A. D. Manukhova, I. B. Bobrov, S. S. Straupe, D. V. Kupriyanov

专题命中 机器人操作 :manipulation(abstract)

AI总结 本文通过原子量子模拟研究自旋与位置结合对量子统计的影响,实现分子中价电子壳层与核中心的模拟,用于量子化学中单价和双价键的模拟。

Comments This version includes the published Erratum appended at the end of the manuscript

Journal ref Phys. Rev. A 111, 062823 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03440 2026-01-06 cs.HC 50%

manvr3d: A Platform for Human-in-the-loop Cell Tracking in Virtual Reality

manvr3d:一种用于虚拟现实中的人机协作细胞追踪平台

Samuel Pantze, Jean-Yves Tinevez, Matthew McGinity, Ulrik Günther

专题命中 机器人操作 :navigation(abstract)

AI总结 manvr3d通过结合虚拟现实与深度学习,实现人机协作的3D细胞追踪,提升空间理解和导航能力。

Comments 7 pages, 6 figures, published at IEEE VIS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01713 2026-01-06 physics.optics 50%

Controllable Lateral Optical Forces on Janus Particles in Fluid Media

在流体介质中可控的Janus粒子横向光学力

Ziheng Xiu, Yue Chai, Pengyu Wen, Chun Meng, Min-Cheng Zhong, Yu-Xuan Ren, Daohong Song, Hrvoje Buljan, Liqin Tang, Zhigang Chen

专题命中 机器人操作 :manipulation(abstract)

AI总结 本研究通过Janus粒子在流体中实现可控横向光学力,利用偏振角调节实现粒子可控推进,为光学操控提供了新方法。

Comments 16 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01436 2026-01-06 cs.CR cs.PL 50%

Bithoven: Formal Safety for Expressive Bitcoin Smart Contracts

Bithoven:比特币表达性智能合约的正式安全性

Hyunhum Cho, Ik Rae Jeong

专题命中 机器人操作 :manipulation(abstract)

AI总结 Bithoven通过整合类型检查器和资源活跃性分析器,提供比特币智能合约的正式安全性,同时保持高效编译性能。

Comments 15 pages, 3 figures, 4 tables. Submitted to IEEE Transactions on Dependable and Secure Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01046 2026-01-06 cs.CL 50%

KV-Embedding: Training-free Text Embedding via Internal KV Re-routing in Decoder-only LLMs

KV-Embedding: 通过解码器-only LLMs中的内部KV重路由实现无训练文本嵌入

Yixuan Tang, Yi Yang

机构 * The Hong Kong University of Science and Technology(香港科技大学)

专题命中 机器人操作 :manipulation(abstract)

AI总结 KV-Embedding通过内部KV重路由实现解码器-only LLMs的无训练文本嵌入,提升性能达10%并保持长序列鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01020 2026-01-06 cond-mat.quant-gas cond-mat.stat-mech cond-mat.str-el 50%

Tunable chiral spiral phases in a non-Hermitian Ising-Gamma spin chain

可调谐的非厄密安伊辛-伽马自旋链中的手性螺旋相

Run-Dong Huang, Wei-Lin Li, Zhi Li

专题命中 机器人操作 :manipulation(abstract)

AI总结 该研究探讨了非厄密安伊辛-伽马自旋链中耗散对螺旋相手性的影响,揭示了两种反向手性螺旋相的共存机制及其转变依赖于非对角线伽马相互作用的强度。

Journal ref Phys. Rev. A 113, 012203 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24416 2026-01-06 cs.CR 50%

GateChain: A Blockchain Based Application for Country Entry-Exit Registry Management

GateChain: 一种基于区块链的国家入境出境登记管理应用

Mohamad Akkad, Hüseyin Bodur

专题命中 机器人操作 :manipulation(abstract)

AI总结 GateChain是一种基于区块链的国家入境出境登记管理应用,通过分布式账本提高数据完整性与透明度,解决传统边境控制系统存在的数据篡改和互操作性问题。

Comments 9 pages, 4 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19412 2026-01-06 physics.optics 50%

Sculpting ultrafast mid-infrared light for solid-state high harmonic generation

雕刻超快中红外光用于固态高阶谐波生成

Camilo Granados, Bálint Kiss, Eric Cormier, Bikash Kumar Das, Debobrata Rajak, Carmelo Rosales-Guzman, Rajaram Shrestha, Qiwen Zhan, Wenlong Gao

专题命中 机器人操作 :manipulation(abstract)

AI总结 本研究通过生成中红外贝塞尔-高斯涡旋和完美光学涡旋,实现了固态中高阶谐波生成的超快结构光合成,验证了OAM守恒及拓扑特性继承。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 具身导航 22 篇

2512.21887 2026-01-06 cs.RO cs.AI 90%

Aerial World Model for Long-horizon Visual Generation and Navigation in 3D Space

空域世界模型用于3D空间中的长视距视觉生成与导航

Weichen Zhang, Peizhi Tang, Xin Zeng, Fanhang Man, Shiquan Yu, Zichao Dai, Baining Zhao, Hongjin Chen, Yu Shang, Wei Wu, Chen Gao, Xinlei Chen, Xin Wang, Yong Li, Wenwu Zhu

机构 * Tsinghua University(清华大学)

专题命中 具身导航 :navigation(title,abstract);world model(title,abstract);embodied agent(abstract);分类 cs.RO、cs.AI

AI总结 ANWM是一种空域导航世界模型,通过预测未来视觉观察来提升无人机在大规模3D空间中的长视距导航性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02294 2026-01-06 cs.RO 86%

RNBF: Real-Time RGB-D Based Neural Barrier Functions for Safe Robotic Navigation

RNBF: 基于实时RGB-D的神经障碍函数用于安全机器人导航

Satyajeet Das, Yifan Xue, Haoming Li, Nadia Figueroa

机构 * University of Southern California(南加州大学) University of Pennsylvania(宾夕法尼亚大学)

专题命中 具身导航 :navigation(title,abstract);robotic(title);分类 cs.RO

AI总结 本文提出了一种实时RGB-D神经障碍函数框架,用于在未知环境中实现安全机器人导航,通过在线构建可微的SDFs并处理传感器噪声,提升导航安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01872 2026-01-06 cs.RO 84%

CausalNav: A Long-term Embodied Navigation System for Autonomous Mobile Robots in Dynamic Outdoor Scenarios

CausalNav: 一种面向动态户外场景的自主移动机器人长期具身导航系统

Hongbo Duan, Shangyi Luo, Zhiyuan Deng, Yanbo Chen, Yuanhao Chiang, Yi Liu, Fangming Liu, Xueqian Wang

机构 * Center for Artificial Intelligence and Robotics, Shenzhen International Graduate School, Tsinghua University(人工智能与机器人中心,深圳国际研究生院,清华大学) Peng Cheng Laboratory(鹏城实验室)

专题命中 具身导航 :navigation(title,abstract);robotics(abstract,comments);分类 cs.RO

AI总结 CausalNav是一种基于场景图的语义导航框架,通过构建多级语义场景图实现动态户外环境中的稳健导航与远距离规划。

Comments Accepted by IEEE Robotics and Automation Letters (RA-L)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02125 2026-01-06 cs.RO cs.AI 81%

SingingBot: An Avatar-Driven System for Robotic Face Singing Performance

SingingBot: 一种由虚拟角色驱动的机器人面部歌唱表演系统

Zhuoxiong Xu, Xuanchen Li, Yuhao Cheng, Fei Xu, Yichao Yan, Xiaokang Yang

专题命中 具身导航 :robotic(title,abstract);分类 cs.RO、cs.AI

AI总结 SingingBot通过虚拟角色驱动框架实现机器人面部歌唱的丰富情感表达,提出情感动态范围指标评估情感广度,实验表明其在情感表现和同步性方面优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23019 2026-01-06 cs.RO 80%

Stairway to Success: An Online Floor-Aware Zero-Shot Object-Goal Navigation Framework via LLM-Driven Coarse-to-Fine Exploration

通往成功的阶梯:一种通过LLM驱动的粗到细探索的在线楼层感知零样本物体-目标导航框架

Zeying Gong, Rong Li, Tianshuai Hu, Ronghe Qiu, Lingdong Kong, Lingfeng Zhang, Guoyang Zhao, Yiyi Ding, Junwei Liang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The Hong Kong University of Science and Technology(香港科学与技术大学) National University of Singapore(新加坡国立大学) Tsinghua University(清华大学)

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO;robotics(comments)

AI总结 本文提出ASCENT框架,通过LLM驱动的粗到细探索实现在线楼层感知零样本物体-目标导航,无需预建地图或重新训练,适用于多楼层环境。

Comments Accepted to IEEE Robotics and Automation Letters (RAL). Project Page at https://zeying-gong.github.io/projects/ascent

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12752 2026-01-06 cs.RO 79%

MOON: Multi-Objective Optimization-Driven Object-Goal Navigation Using a Variable-Horizon Set-Orienteering Planner

MOON:基于多目标优化的物体-目标导航变量视距集导向规划器

Daigo Nakajima, Kanji Tanaka, Daiki Iwata, Kouki Terashima

机构 * University of Fukui(福井大学)

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO

AI总结 MOON通过多目标优化解决大规模复杂环境中的导航问题,结合可变视距集导向规划和高效神经规划器,提升全局规划与实时决策效率。

Comments 9 pages, 7 figures, technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01546 2026-01-06 cs.AI 79%

Improving Behavioral Alignment in LLM Social Simulations via Context Formation and Navigation

通过情境形成与导航提升LLM社交模拟中的行为一致性

Letian Kong, Qianran, Jin, Renyu Zhang

专题命中 具身导航 :navigation(title,abstract);分类 cs.AI

AI总结 本文提出通过情境形成与导航提升LLM在复杂决策环境中的行为一致性,验证了两阶段框架在不同任务中的有效性。

Comments 39 pages, 2 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01067 2026-01-06 cs.RO 79%

Topological Mapping and Navigation using a Monocular Camera based on AnyLoc

基于AnyLoc的单目相机拓扑映射与导航

Wenzheng Zhang, Yoshitaka Hara, Sousuke Nakamura

机构 * Graduate School of Science and Engineering, Hosei University(Hosei大学工学研究院) Future Robotics Technology Center (fuRo), Chiba Institute of Technology(Chiba技术大学未来机器人技术中心) Faculty of Science and Engineering, Hosei University(Hosei大学工学学院)

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO

AI总结 基于AnyLoc的单目相机实现高效拓扑映射与导航,通过关键节点简化路径规划,提升导航成功率并降低计算成本。

Comments Published in Proc. IEEE CASE 2025. 7 pages, 11 figures

Journal ref Proc. IEEE International Conference on Automation Science and Engineering (CASE), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02208 2026-01-06 cs.RO 79%

ADMM-MCBF-LCA: A Layered Control Architecture for Safe Real-Time Navigation

ADMM-MCBF-LCA:一种用于安全实时导航的分层控制架构

Anusha Srikanthan, Yifan Xue, Vijay Kumar, Nikolai Matni, Nadia Figueroa

机构 * School of Engineering and Applied Science, University of Pennsylvania(工程与应用科学学院,宾夕法尼亚大学)

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO

AI总结 ADMM-MCBF-LCA通过分层控制架构实现安全实时导航,结合离线路径库和在线路径选择,确保在动态环境中安全完成任务。

Journal ref Proc. IEEE Int. Conf. Robot. Autom. (ICRA), 2025, pp. 9243-9250

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22368 2026-01-06 eess.SY cs.SY 78%

Distributed Koopman Operator Learning for Perception and Safe Navigation

分布式Koopman算子学习用于感知与安全导航

Ali Azarbahram, Shenyu Liu, Gian Paolo Incremona

专题命中 具身导航 :navigation(title,abstract)

AI总结 本文提出一种结合模型预测控制与分布式Koopman算子学习的框架,用于在动态交通环境中实现安全高效的自主导航。

详情

展开后加载摘要…

URL PDF HTML 收藏