arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

2026-01-27 至 2026-01-27 共收录 15 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 具身导航 15 篇

2512.10046 2026-01-27 cs.AI 89%

SimWorld-Robotics: Synthesizing Photorealistic and Dynamic Urban Environments for Multimodal Robot Navigation and Collaboration

SimWorld-Robotics: 为多模态机器人导航与协作合成逼真动态城市环境

Yan Zhuang, Jiawei Ren, Xiaokang Ye, Jianzhi Shen, Ruixuan Zhang, Tianai Yue, Muhammad Faayez, Xuhong He, Ziqiao Ma, Lianhui Qin, Zhiting Hu, Tianmin Shu

机构 * University of Virginia(弗吉尼亚大学) UC San Diego(加州大学圣地亚哥分校) Johns Hopkins University(约翰霍普金斯大学) Carnegie Mellon University(卡内基梅隆大学) University of Michigan(密歇根大学)

专题命中 具身导航 :robotics(title,abstract);navigation(title,abstract);embodied AI(abstract);分类 cs.AI

AI总结 SimWorld-Robotics通过合成逼真动态城市环境,提出两个多模态机器人基准测试,评估机器人在复杂场景中的导航、协作与通信能力。

Comments Conference: NeurIPS 2025 (main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18492 2026-01-27 cs.RO 83%

DV-VLN: Dual Verification for Reliable LLM-Based Vision-and-Language Navigation

DV-VLN:基于可靠大语言模型的视觉-语言导航的双重验证

Zijun Li, Shijie Li, Zhenxi Zhang, Bin Li, Shoujun Zhou

机构 * Robotics Engineering Program, College of Engineering, Zhejiang Normal University(浙江师范大学工程学院机器人工程专业) Shenzhen Institutes of Advanced Technology (SIAT), Chinese Academy of Sciences(中国科学院深圳先进技术研究院) Department of Health Technology and Informatics, The Hong Kong Polytechnic University(香港理工大学健康科技与信息技术系)

专题命中 具身导航 :navigation(title,abstract);embodied agent(abstract);分类 cs.RO

AI总结 DV-VLN通过生成-验证范式提升大语言模型在视觉-语言导航中的可靠性,通过双重验证机制提高导航决策的准确性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16078 2026-01-27 cs.RO cs.SY eess.SY 79%

Improve the autonomy of the SE2(3) group based Extended Kalman Filter for Integrated Navigation: Application

提升基于SE2(3)群的扩展卡尔曼滤波器自主性以实现集成导航:应用

Maosong Wang, Jiarui Cui, Wenqi Wu, Peiqi Li, Xianfei Pan

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO

AI总结 本文提出改进基于SE2(3)群的扩展卡尔曼滤波器自主性,通过实验和模拟验证其在高精度导航中的应用效果。

Comments arXiv admin note: substantial text overlap with arXiv:2601.16062. substantial text overlap with arXiv:2601.16062. substantial text overlap with arXiv:2601.16062. substantial text overlap with arXiv:2601.16062

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17428 2026-01-27 cs.RO 70%

Scaling Rough Terrain Locomotion with Automatic Curriculum Reinforcement Learning

通过自动课程强化学习扩展粗糙地形运动

Ziming Li, Chenhao Li, Marco Hutter

机构 * Robotic Systems Lab, ETH Zurich, Switzerland(苏黎世联邦理工学院机器人系统实验室) ETH AI Center, ETH Zurich, Switzerland(苏黎世联邦理工学院人工智能中心)

专题命中 具身导航 :robot learning(abstract);robotic(abstract);分类 cs.RO

AI总结 本文提出基于学习进度的自动课程强化学习框架,使四足机器人在复杂地形上实现稳定高速运动,突破传统方法在地形适应上的限制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20671 2026-01-27 cs.CV 70%

IPFormer: Visual 3D Panoptic Scene Completion with Context-Adaptive Instance Proposals

IPFormer: 基于上下文自适应实例提案的视觉3D全景场景补全

Markus Gross, Aya Fahmy, Danit Niwattananan, Dominik Muhle, Rui Song, Daniel Cremers, Henri Meeß

机构 * Fraunhofer Institute IVI(弗劳恩霍夫研究所IVI) Technical University of Munich(慕尼黑技术大学) Munich Center for Machine Learning(慕尼黑机器学习中心) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 具身导航 :robotics(abstract);navigation(abstract);分类 cs.CV

AI总结 IPFormer通过上下文自适应实例提案实现视觉3D全景场景补全,提升场景理解与泛化能力。

Journal ref Advances in Neural Information Processing Systems (NeurIPS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17383 2026-01-27 cs.CV cs.AI 62%

Physical Prompt Injection Attacks on Large Vision-Language Models

针对大视觉-语言模型的物理提示注入攻击

Chen Ling, Kai Hu, Hangcheng Liu, Xingshuo Han, Tianwei Zhang, Changhai Ou

机构 * School of Cyber Science and Engineering, Wuhan University(武汉大学计算机科学与工程学院) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院) College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院)

专题命中 具身导航 :navigation(abstract);分类 cs.AI、cs.CV

AI总结 本研究提出了一种无需访问模型或输入的物理提示注入攻击方法,通过物理物体注入恶意指令,成功攻击多种大视觉-语言模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02328 2026-01-27 cs.RO cs.LG 62%

Path Planning using a One-shot-sampling Skeleton Map

基于一次性采样骨架图的路径规划

Gabriel O. Flores-Aquino, Octavio Gutierrez-Frias, Juan Irving Vasquez

机构 * Centro de Investigación en Matematicas (CIMAT), A.C.(数学研究所(CIMAT)) Unidad Profesional Interdisciplinaria en Ingeniería y Tecnologías Avanzadas (UPIITA), Instituto Politécnico Nacional (IPN)(跨学科工程与先进技术单位(UPIITA),墨西哥理工学院(IPN)) Centro de Innovación y Desarrollo Tecnológico en Cómputo (CIDETEC), Instituto Politécnico Nacional (IPN)(计算创新与技术发展中心(CIDETEC),墨西哥理工学院(IPN))

专题命中 具身导航 :navigation(abstract);分类 cs.RO、cs.LG

AI总结 本文提出基于SkelUnet的高效路径规划方法,通过一次性采样骨架图快速生成安全路径,提升导航效率和安全性。

Comments Submitted to IEEE Latin America Transactions

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13170 2026-01-27 cs.CV cs.AI 62%

Unified-EGformer: Exposure Guided Lightweight Transformer for Mixed-Exposure Image Enhancement

统一曝光引导轻量变换器:用于混合曝光图像增强

Eashan Adhikarla, Kai Zhang, Rosaura G. VidalMata, Manjushree Aithal, Nikhil Ambha Madhusudhana, John Nicholson, Lichao Sun, Brian D. Davison

机构 * Lehigh University(莱维大学) Lenovo Research(联想研究)

专题命中 具身导航 :navigation(abstract);分类 cs.AI、cs.CV

AI总结 Unified-EGformer通过轻量级变换器架构实现混合曝光图像增强,具备高效推理和跨任务泛化能力。

Comments Under submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18168 2026-01-27 cs.CV 57%

TempDiffReg: Temporal Diffusion Model for Non-Rigid 2D-3D Vascular Registration

TempDiffReg:用于非刚性2D-3D血管配准的时序扩散模型

Zehua Liu, Shihao Zou, Jincai Huang, Yanfang Zhang, Chao Tong, Weixin Si

机构 * School of Computer Science and Engineering, Beihang University, Beijing, China(北京航空航天大学计算机科学与工程学院) State Key Laboratory of Virtual Reality Technology and Systems, Beihang University, Beijing, China(北京航空航天大学虚拟现实技术与系统国家重点实验室) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen, China(中国科学院深圳先进技术研究所) School of Computer Science and Control Engineering, Shenzhen University of Advanced Technology, Shenzhen, China(深圳先进技术大学计算机科学与控制工程学院) Department of Interventional Radiology, Shenzhen People’s Hospital, Shenzhen, China(深圳人民医院介入放射科)

专题命中 具身导航 :navigation(abstract);分类 cs.CV

AI总结 TempDiffReg通过时序扩散模型和结构感知视角n点模块,实现高精度的2D-3D血管配准,提升TACE手术的准确性和安全性。

Comments Accepted by IEEE BIBM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17231 2026-01-27 cs.RO cs.AR 57%

Real-Time, Energy-Efficient, Sampling-Based Optimal Control via FPGA Acceleration

基于FPGA加速的实时、节能的采样最优控制

Tanmay Desai, Brian Plancher, R. Iris Bahar

机构 * Colorado School of Mines(科罗拉多矿业学院) Barnard College(巴纳德学院) Columbia University(哥伦比亚大学) Dartmouth College(达特茅斯学院)

专题命中 具身导航 :robotics(abstract);分类 cs.RO

AI总结 本文提出一种基于FPGA加速的MPPI控制方法,通过深度流水线和并行性优化,实现嵌入式平台上的高效能和低能耗控制。

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01054 2026-01-27 cs.CR cs.AI cs.CY 57%

Autonomous Penetration Testing: Solving Capture-the-Flag Challenges with LLMs

自主渗透测试:利用大语言模型解决夺旗挑战

Isabelle Bakker, John Hastings

机构 * The Beacom College of Computer and Cyber Sciences(计算机与网络安全科学学院) Dakota State University(达科他州立大学)

专题命中 具身导航 :navigation(abstract);分类 cs.AI

AI总结 利用大语言模型自主解决初级渗透测试任务,展示其在自动化简单攻击流程中的潜力,同时揭示安全环境对LLM攻击的挑战。

Comments 6 pages, 2 figures, 3 tables

Journal ref 2025 IEEE Cyber Awareness and Research Symposium (CARS'25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18583 2026-01-27 physics.optics eess.IV physics.app-ph 50%

Uncooled Poisson Bolometer for High-Speed Event-Based Long-wave Thermal Imaging

无冷却的泊松 bolometer 用于高速基于事件的长波热成像

Mohamed A. Mousa, Leif Bauer, Utkarsh Singh, Ziyi Yang, Angshuman Deka, Zubin Jacob

专题命中 具身导航 :navigation(abstract)

AI总结 本研究提出了一种无冷却的泊松 bolometer,实现了高速、低功耗的长波热成像,其事件率高达 1,250 Hz,超越传统无冷却微bolometers 的时间分辨率 4 倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18328 2026-01-27 cs.HC 50%

MarioChart: Autonomous Tangibles as Active Proxy Interfaces for Embodied Casual Data Exploration

MarioChart:自主实体作为具身偶然数据探索的主动代理接口

Shaozhang Dai, Kadek Ananta Satriadi, Jim Smiley, Barrett Ens, Lonni Besançon, Tim Dwyer

专题命中 具身导航 :manipulation(abstract)

AI总结 MarioChart通过自主实体作为主动代理接口,提升偶然数据探索中的短期空间记忆和数据分析效率,同时保持与传统触屏在长期记忆等指标上的相似性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18179 2026-01-27 cs.HC 50%

Exploring Customizable Interactive Tools for Therapeutic Homework Support in Mental Health Counseling

探索用于心理健康辅导中治疗作业支持的可定制交互工具

Yimeng Wang, Liabette Escamilla, Yinzhou Wang, Bianca R. Augustine, Yixuan Zhang

专题命中 具身导航 :navigation(abstract)

AI总结 TheraTrack是一款可定制的治疗师工具,利用大型语言模型整合多维数据,优化治疗作业跟踪,减少认知负荷并提高数据验证效率。

Comments 22 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17769 2026-01-27 cs.HC 50%

Reflexa: Uncovering How LLM-Supported Reflection Scaffolding Reshapes Creativity in Creative Coding

Reflexa:揭示LLM支持的反思支架如何重塑创意编码中的创造力

Anqi Wang, Zhengyi Li, Lan Luo, Xin Tong, Pan Hui

专题命中 具身导航 :navigation(abstract)

AI总结 Reflexa通过系统化的反思支架提升创意编码中的创造力,通过结构化反思模式增强可控性、探索广度和原创性。

详情

展开后加载摘要…

URL PDF HTML 收藏