arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

2025-12-19 至 2025-12-19 共收录 51 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 具身推理 3 篇

2512.16924 2025-12-19 cs.CV 57%

The World is Your Canvas: Painting Promptable Events with Reference Images, Trajectories, and Text

世界是你的画布:通过参考图像、轨迹和文本绘画可提示事件

Hanlin Wang, Hao Ouyang, Qiuyu Wang, Yue Yu, Yihao Meng, Wen Wang, Ka Leong Cheng, Shuailei Ma, Qingyan Bai, Yixuan Li, Cheng Chen, Yanhong Zeng, Xing Zhu, Yujun Shen, Qifeng Chen

机构 * HKUST(香港科技大学) Ant Group(蚂蚁集团) ZJU(浙江大学) NEU(南京大学) CUHK(香港中文大学) NTU(南洋理工大学)

专题命中 具身推理 :world model(abstract);分类 cs.CV

AI总结 WorldCanvas通过结合文本、轨迹和参考图像,实现可提示的多代理交互事件生成,提升世界模型的交互性和可控性。

Comments Project page and code: https://worldcanvas.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 模仿学习与强化学习 5 篇

2508.16807 2025-12-19 cs.RO cs.AI cs.LG cs.SY eess.SY 83%

Autonomous UAV Flight Navigation in Confined Spaces: A Reinforcement Learning Approach

在受限空间中自主无人机飞行导航:一种强化学习方法

Marco S. Tayar, Lucas K. de Oliveira, Felipe Andrade G. Tommaselli, Juliano D. Negri, Thiago H. Segreto, Ricardo V. Godoy, Marcelo Becker

机构 * Department of Mechanical Engineering, University of São Paulo(机械工程系,圣保罗大学)

专题命中 模仿学习与强化学习 :navigation(title,abstract);分类 cs.RO、cs.AI、cs.LG;robotics(journal_ref)

AI总结 本文通过比较PPO和SAC在高保真模拟器中实现精确飞行,发现PPO在高精度安全导航任务中表现更优。

Journal ref 2025 Latin American Robotics Symposium (LARS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16911 2025-12-19 cs.LG cs.AI cs.RO 80%

Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL Finetuning

后验行为克隆:为高效强化学习微调预训练BC策略

Andrew Wagenmaker, Perry Dong, Raymond Tsao, Chelsea Finn, Sergey Levine

机构 * UC Berkeley(伯克利大学) Stanford(斯坦福大学)

专题命中 模仿学习与强化学习 :robotics(abstract);manipulation(abstract);robotic(abstract);分类 cs.RO、cs.AI、cs.LG

AI总结 本文提出后验行为克隆策略,通过建模示范者行为的后验分布来提升强化学习微调效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16813 2025-12-19 cs.NI cs.AI cs.DC cs.LG eess.SP 62%

Coordinated Anti-Jamming Resilience in Swarm Networks via Multi-Agent Reinforcement Learning

通过多智能体强化学习实现蜂群网络中的协同抗干扰韧性

Bahman Abolhassani, Tugba Erpek, Kemal Davaslioglu, Yalin E. Sagduyu, Sastry Kompella

专题命中 模仿学习与强化学习 :robotic(abstract);分类 cs.AI、cs.LG

AI总结 本文提出基于QMIX算法的多智能体强化学习框架,用于提升蜂群网络在反应式干扰下的通信韧性,通过仿真验证其在对抗环境中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15705 2025-12-19 cs.CV 57%

GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization

GeoVista: 基于网络增强的代理视觉推理用于地理定位

Yikun Wang, Zuyan Liu, Ziyi Wang, Han Hu, Pengfei Liu, Yongming Rao

机构 * Fudan University(复旦大学) Tencent Hunyuan(腾讯文元) Tsinghua University(清华大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 模仿学习与强化学习 :manipulation(abstract);分类 cs.CV

AI总结 GeoVista通过整合图像缩放和网络搜索工具,提升地理定位任务的代理模型性能,实现优于开源模型的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16787 2025-12-19 cs.AI 57%

Enter the Void - Planning to Seek Entropy When Reward is Scarce

进入虚无 - 在奖励稀缺时规划以寻求熵

Ashish Sundar, Chunbo Luo, Xiaoyang Wang

机构 * Department of Computer Science University of Exeter(计算机科学系埃克塞特大学)

专题命中 模仿学习与强化学习 :world model(abstract);分类 cs.AI

AI总结 本文提出了一种基于世界模型的分层规划方法,在奖励稀缺时主动寻找信息丰富的状态,提升样本效率,应用于Dreamer时在迷宫任务中效率提升50%。

Comments 10 pages without appendix, 15 Figures, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 人机交互与遥操作 2 篇

2512.15776 2025-12-19 cs.AI cs.MA cs.RO 76%

Emergence: Overcoming Privileged Information Bias in Asymmetric Embodied Agents via Active Querying

涌现:通过主动查询克服不对称具身智能体中的特权信息偏差

Shaun Baek, Sam Liu, Joseph Ukpong

机构 * Emory University(埃默里大学)

专题命中 人机交互与遥操作 :embodied agent(title);分类 cs.RO、cs.AI

AI总结 本文通过主动查询机制解决不对称具身智能体中的特权信息偏差问题,揭示了沟通接地错误对协作成功率的影响。

Comments 12 pages, 9 pages of content, 6 tables, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06189 2025-12-19 cs.RO cs.HC 57%

Accessible and Pedagogically-Grounded Explainability for Human-Robot Interaction: A Framework Based on UDL and Symbolic Interfaces

可及且以教学为导向的机器人可解释性:基于UDL和符号接口的框架

Francisco J. Rodríguez Lera, Raquel Fernández Hernández, Sonia Lopez González, Miguel Angel González-Santamarta, Francisco Jesús Rodríguez Sedano, Camino Fernandez Llamas

机构 * Grupo de Robótica , EIIIA Campus de Vegazana, Universidad de León(机器人组,EIIIA校区,莱昂大学) C.E.E. Ntra. Sra. Del Sagrado Corazón(圣心宗女会)

专题命中 人机交互与遥操作 :robotics(abstract);分类 cs.RO

AI总结 本文提出基于UDL和符号接口的框架,旨在通过多模态解释板提升人机交互中不同需求用户的可解释性与教学导向性。

Comments 6 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 机器人数据与评测 10 篇

2508.17468 2025-12-19 cs.CV cs.AI cs.LG cs.RO 80%

A Synthetic Dataset for Manometry Recognition in Robotic Applications

用于机器人应用的测压识别合成数据集

Pedro Antonio Rabelo Saraiva, Enzo Ferreira de Souza, Joao Manoel Herrera Pinheiro, Thiago H. Segreto, Ricardo V. Godoy, Marcelo Becker

机构 * University of São Paulo(圣保罗大学)

专题命中 机器人数据与评测 :robotic(title);分类 cs.RO、cs.AI、cs.CV;robotics(journal_ref)

AI总结 本文提出了一种混合数据合成方法,通过结合过程渲染和AI驱动视频生成,提升机器人应用中测压识别的训练效果和可靠性。

Journal ref 2025 Latin American Robotics Symposium (LARS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01194 2025-12-19 cs.RO 70%

RoboLoc: A Benchmark Dataset for Point Place Recognition and Localization in Indoor-Outdoor Integrated Environments

RoboLoc:用于室内-室外集成环境中的点位识别和定位的基准数据集

Jaejin Jeon, Seonghoon Ryoo, Sang-Duck Lee, Soomok Lee, Seungwoo Jeong

机构 * Department of Data, Network and AI, Ajou University(数据、网络与人工智能系,亚运会大学) Korea Railroad Research Institute(韩国铁路研究院) Department of Mobility Engineering, Ajou University(移动工程系,亚运会大学)

专题命中 机器人数据与评测 :robotics(abstract);navigation(abstract);分类 cs.RO

AI总结 RoboLoc是一个用于室内外集成环境的点位识别和定位基准数据集,通过真实机器人轨迹和多领域转换,评估多种先进模型的泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16295 2025-12-19 cs.AI 70%

OS-Oracle: A Comprehensive Framework for Cross-Platform GUI Critic Models

OS-Oracle: 一个跨平台GUI批评模型的综合框架

Zhenyu Wu, Jingjing Xie, Zehao Li, Bowen Yang, Qiushi Sun, Zhaoyang Liu, Zhoumianze Liu, Yu Qiao, Xiangyu Yue, Zun Wang, Zichen Ding

机构 * Shanghai Jiaotong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室) CUHK MMLab(香港大学MMLab) The University of Hong Kong(香港大学) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 机器人数据与评测 :manipulation(abstract);navigation(abstract);分类 cs.AI

AI总结 OS-Oracle提出了一种跨平台GUI批评模型框架,通过合成数据和两阶段训练方法,实现了在移动、网页和桌面平台上的高效批评评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13107 2025-12-19 cs.CV cs.AI 62%

Diffusion-Based Restoration for Multi-Modal 3D Object Detection in Adverse Weather

基于扩散的多模态3D物体检测在恶劣天气中的修复

Zhijian He, Feifei Liu, Yuwei Li, Zhanpeng Luo, Jintao Cheng, Xieyuanli Chen, Xiaoyu Tang

机构 * School of Xingzhi College, South China Normal University(星智学院,华南师范大学) College of Big Data and Internet, Shenzhen Technology University(大数据与互联网学院,深圳科技大学) School of Data Science and Engineering, Xingzhi College, South China Normal University(数据科学与工程学院,星智学院,华南师范大学) Department of Electronic and Computer Engineering, Hong Kong University of Science and Technology(电子与计算机工程系,香港科技大学) College of Intelligence Science and Technology, National University of Defense Technology(智能科学与技术学院,国防科技大学)

专题命中 机器人数据与评测 :robotics(abstract);分类 cs.AI、cs.CV

AI总结 DiffFusion通过基于扩散的修复和自适应跨模态融合,提升多模态3D物体检测在恶劣天气中的鲁棒性与清洁数据性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15957 2025-12-19 cs.CV cs.AI 62%

Seeing is Believing (and Predicting): Context-Aware Multi-Human Behavior Prediction with Vision Language Models

看见即信仰(并预测):基于视觉语言模型的上下文感知多人类行为预测

Utsav Panchal, Yuchen Liu, Luigi Palmieri, Ilche Georgievski, Marco Aiello

机构 * Institute of Architecture of Application Systems, University of Stuttgart, Germany(应用系统建筑研究所,斯图加特大学,德国) Bosch Research, Germany(博世研究,德国)

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.AI、cs.CV

AI总结 CAMP-VLM通过结合视觉语言模型与上下文特征,提升了多人类行为预测的准确性,其在预测精度上比基线模型高66.9%。

Comments Accepted at IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16493 2025-12-19 cs.CV 57%

YOLO11-4K: An Efficient Architecture for Real-Time Small Object Detection in 4K Panoramic Images

YOLO11-4K: 一种高效的实时小目标检测架构用于4K全景图像

Huma Hafeez, Matthew Garratt, Jo Plested, Sankaran Iyer, Arcot Sowmya

机构 * School of Engineering \& Technology, University of New South Wales, Canberra, Australia School of Systems \& Computing, University of New South Wales, Canberra, Australia School of Computer Science \& Engineering, University of New South Wales, Sydney, Australia

专题命中 机器人数据与评测 :navigation(abstract);分类 cs.CV

AI总结 YOLO11-4K通过高效架构实现4K全景图像中小目标的实时高精度检测,相比YOLO11在精度和速度上均有显著提升。

Comments Conference paper just submitted

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16213 2025-12-19 cs.CV math.DG 57%

Enhanced 3D Shape Analysis via Information Geometry

通过信息几何增强的3D形状分析

Amit Vishwakarma, K. S. Subrahamanian Moosath

机构 * Indian Institute of Space Science and Technology(印度空间科学与技术研究所)

专题命中 机器人数据与评测 :robotics(abstract);分类 cs.CV

AI总结 本文提出基于信息几何的3D点云分析方法,通过高斯混合模型和改进的KL散度实现稳定且准确的形状比较。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12793 2025-12-19 cs.RO 57%

VLG-Loc: Vision-Language Global Localization from Labeled Footprint Maps

VLG-Loc: 从标注足迹地图实现视觉-语言全局定位

Mizuho Aoki, Kohei Honda, Yasuhiro Yoshimura, Takeshi Ishita, Ryo Yonetani

机构 * The Department of Mechanical Systems Engineering, Nagoya University(名古屋大学机械系统工程系) CyberAgent AI Lab(CyberAgent AI实验室)

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.RO

AI总结 VLG-Loc通过视觉-语言模型实现从标注足迹地图的全局定位,结合蒙特卡洛定位框架和概率融合提升环境变化下的鲁棒性。

Comments v2: Updated the citation of SparseLoc from an arXiv preprint to its published version in the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07392 2025-12-19 cs.CL cs.AI 57%

Voice-Interactive Surgical Agent for Multimodal Patient Data Control

多模态患者数据控制的语音交互手术代理

Hyeryun Park, Byung Mo Gu, Jun Hee Lee, Byeong Hyeon Choi, Sekeun Kim, Hyun Koo Kim, Kyungsang Kim

机构 * Image Guided Precision Cancer Surgery Institute, College of Medicine, Korea University(影像引导精准癌症手术研究所,韩国大学医学院) Department of Radiology, Massachusetts General Hospital(放射科,麻省总医院) Department of Thoracic and Cardiovascular Surgery, Korea University Guro Hospital, College of Medicine, Korea University(胸外科和心血管外科,韩国大学Guro医院,韩国大学医学院) Department of Biomedical Sciences, College of Medicine, Korea University(生物医学科学系,韩国大学医学院)

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.AI

AI总结 本文提出一种基于大型语言模型的语音交互手术代理,用于多模态患者数据的高效控制与处理,通过分层多代理框架提升手术流程中的数据操作效率与鲁棒性。

Comments 14 pages, 13 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16826 2025-12-19 cs.CV cs.AI 54%

Next-Generation License Plate Detection and Recognition System using YOLOv8

基于YOLOv8的下一代车牌检测与识别系统

Arslan Amin, Rafia Mumtaz, Muhammad Jawad Bashir, Syed Mohammad Hassan Zaidi

机构 * School of Electrical Engineering Computer Science (SEECS), \ University of Sciences Ghulam Ishaq Khan Institute of Engineering Sciences \ Technology (GIKI), Topi, District, Swabi, \ Pakhtunkhwa 23460, Pakistan, Email:prorector\

专题命中 机器人数据与评测 :robotics(comments,journal_ref);分类 cs.AI、cs.CV

AI总结 本文提出基于YOLOv8的车牌检测与识别系统,通过优化模型配置提升实时准确率,为智能交通系统提供高效解决方案。

Comments 6 pages, 5 figures. Accepted and published in the 2023 IEEE 20th International Conference on Smart Communities: Improving Quality of Life using AI, Robotics and IoT (HONET)

Journal ref 2023 IEEE 20th International Conference on Smart Communities: Improving Quality of Life using AI, Robotics and IoT (HONET)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 其他机器人 3 篇

2512.15993 2025-12-19 cs.CV 79%

Eyes on the Grass: Biodiversity-Increasing Robotic Mowing Using Deep Visual Embeddings

草地之眼:利用深度视觉嵌入提升生物多样性的机器人修剪

Lars Beckers, Arno Waes, Aaron Van Campenhout, Toon Goedemé

机构 * EAVISE-PSI research group, Department of Electrical Engineering ESAT, KU Leuven, Belgium(EAVISE-PSI研究组,电气工程系ESAT,卢森堡大学)

专题命中 其他机器人 :robotic(title,abstract);分类 cs.CV

AI总结 本文提出利用深度视觉嵌入提升生物多样性的机器人修剪方法,通过选择性修剪与保护行为交替,实现草坪生态价值的提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13036 2025-12-19 cs.RO eess.SP 57%

Dual-Channel Tomographic Tactile Skin with Pneumatic Pressure Sensing for Improved Force Estimation

双通道触觉皮肤与气压传感用于改进的力估计

Haofeng Chen, Jiri Kubik, Bedrich Himmel, Matej Hoffmann, Hyosang Lee

机构 * Department of Cybernetics, Faculty of Electrical Engineering, Czech Technical University in Prague(电子工程学院自动化系,捷克技术大学)

专题命中 其他机器人 :robotic(abstract);分类 cs.RO

AI总结 本文提出双通道触觉皮肤结合EIT与气压传感,通过位置感知校正实现多接触力估计,提升力估计精度并简化校准流程。

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16178 2025-12-19 cs.CV cs.AI cs.RO 56%

Towards Closing the Domain Gap with Event Cameras

通过事件相机缩小领域差距

M. Oltan Sevinc, Liao Wu, Francisco Cruz

机构 * School of Computer Science and Engineering, University of New South Wales(新南威尔士大学计算机科学与工程学院) School of Mechanical and Manufacturing Engineering, University of New South Wales(新南威尔士大学机械与制造工程学院) Escuela de Ingeniería, Universidad Central de Chile(智利中央大学工程学院)

专题命中 其他机器人 :分类 cs.RO、cs.AI、cs.CV;robotics(comments)

AI总结 本文提出使用事件相机以解决昼夜光照差异带来的领域差距问题,展示其在不同光照条件下更稳定的性能和更优的跨领域表现。

Comments Accepted to Australasian Conference on Robotics and Automation (ACRA), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏