arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

2025-12-19 至 2025-12-19 共收录 10 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 机器人数据与评测 10 篇

2508.17468 2025-12-19 cs.CV cs.AI cs.LG cs.RO 80%

A Synthetic Dataset for Manometry Recognition in Robotic Applications

用于机器人应用的测压识别合成数据集

Pedro Antonio Rabelo Saraiva, Enzo Ferreira de Souza, Joao Manoel Herrera Pinheiro, Thiago H. Segreto, Ricardo V. Godoy, Marcelo Becker

机构 * University of São Paulo(圣保罗大学)

专题命中 机器人数据与评测 :robotic(title);分类 cs.RO、cs.AI、cs.CV;robotics(journal_ref)

AI总结 本文提出了一种混合数据合成方法,通过结合过程渲染和AI驱动视频生成,提升机器人应用中测压识别的训练效果和可靠性。

Journal ref 2025 Latin American Robotics Symposium (LARS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01194 2025-12-19 cs.RO 70%

RoboLoc: A Benchmark Dataset for Point Place Recognition and Localization in Indoor-Outdoor Integrated Environments

RoboLoc:用于室内-室外集成环境中的点位识别和定位的基准数据集

Jaejin Jeon, Seonghoon Ryoo, Sang-Duck Lee, Soomok Lee, Seungwoo Jeong

机构 * Department of Data, Network and AI, Ajou University(数据、网络与人工智能系,亚运会大学) Korea Railroad Research Institute(韩国铁路研究院) Department of Mobility Engineering, Ajou University(移动工程系,亚运会大学)

专题命中 机器人数据与评测 :robotics(abstract);navigation(abstract);分类 cs.RO

AI总结 RoboLoc是一个用于室内外集成环境的点位识别和定位基准数据集,通过真实机器人轨迹和多领域转换,评估多种先进模型的泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16295 2025-12-19 cs.AI 70%

OS-Oracle: A Comprehensive Framework for Cross-Platform GUI Critic Models

OS-Oracle: 一个跨平台GUI批评模型的综合框架

Zhenyu Wu, Jingjing Xie, Zehao Li, Bowen Yang, Qiushi Sun, Zhaoyang Liu, Zhoumianze Liu, Yu Qiao, Xiangyu Yue, Zun Wang, Zichen Ding

机构 * Shanghai Jiaotong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室) CUHK MMLab(香港大学MMLab) The University of Hong Kong(香港大学) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 机器人数据与评测 :manipulation(abstract);navigation(abstract);分类 cs.AI

AI总结 OS-Oracle提出了一种跨平台GUI批评模型框架,通过合成数据和两阶段训练方法,实现了在移动、网页和桌面平台上的高效批评评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13107 2025-12-19 cs.CV cs.AI 62%

Diffusion-Based Restoration for Multi-Modal 3D Object Detection in Adverse Weather

基于扩散的多模态3D物体检测在恶劣天气中的修复

Zhijian He, Feifei Liu, Yuwei Li, Zhanpeng Luo, Jintao Cheng, Xieyuanli Chen, Xiaoyu Tang

机构 * School of Xingzhi College, South China Normal University(星智学院,华南师范大学) College of Big Data and Internet, Shenzhen Technology University(大数据与互联网学院,深圳科技大学) School of Data Science and Engineering, Xingzhi College, South China Normal University(数据科学与工程学院,星智学院,华南师范大学) Department of Electronic and Computer Engineering, Hong Kong University of Science and Technology(电子与计算机工程系,香港科技大学) College of Intelligence Science and Technology, National University of Defense Technology(智能科学与技术学院,国防科技大学)

专题命中 机器人数据与评测 :robotics(abstract);分类 cs.AI、cs.CV

AI总结 DiffFusion通过基于扩散的修复和自适应跨模态融合,提升多模态3D物体检测在恶劣天气中的鲁棒性与清洁数据性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15957 2025-12-19 cs.CV cs.AI 62%

Seeing is Believing (and Predicting): Context-Aware Multi-Human Behavior Prediction with Vision Language Models

看见即信仰(并预测):基于视觉语言模型的上下文感知多人类行为预测

Utsav Panchal, Yuchen Liu, Luigi Palmieri, Ilche Georgievski, Marco Aiello

机构 * Institute of Architecture of Application Systems, University of Stuttgart, Germany(应用系统建筑研究所,斯图加特大学,德国) Bosch Research, Germany(博世研究,德国)

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.AI、cs.CV

AI总结 CAMP-VLM通过结合视觉语言模型与上下文特征,提升了多人类行为预测的准确性,其在预测精度上比基线模型高66.9%。

Comments Accepted at IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16493 2025-12-19 cs.CV 57%

YOLO11-4K: An Efficient Architecture for Real-Time Small Object Detection in 4K Panoramic Images

YOLO11-4K: 一种高效的实时小目标检测架构用于4K全景图像

Huma Hafeez, Matthew Garratt, Jo Plested, Sankaran Iyer, Arcot Sowmya

机构 * School of Engineering \& Technology, University of New South Wales, Canberra, Australia School of Systems \& Computing, University of New South Wales, Canberra, Australia School of Computer Science \& Engineering, University of New South Wales, Sydney, Australia

专题命中 机器人数据与评测 :navigation(abstract);分类 cs.CV

AI总结 YOLO11-4K通过高效架构实现4K全景图像中小目标的实时高精度检测,相比YOLO11在精度和速度上均有显著提升。

Comments Conference paper just submitted

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16213 2025-12-19 cs.CV math.DG 57%

Enhanced 3D Shape Analysis via Information Geometry

通过信息几何增强的3D形状分析

Amit Vishwakarma, K. S. Subrahamanian Moosath

机构 * Indian Institute of Space Science and Technology(印度空间科学与技术研究所)

专题命中 机器人数据与评测 :robotics(abstract);分类 cs.CV

AI总结 本文提出基于信息几何的3D点云分析方法,通过高斯混合模型和改进的KL散度实现稳定且准确的形状比较。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12793 2025-12-19 cs.RO 57%

VLG-Loc: Vision-Language Global Localization from Labeled Footprint Maps

VLG-Loc: 从标注足迹地图实现视觉-语言全局定位

Mizuho Aoki, Kohei Honda, Yasuhiro Yoshimura, Takeshi Ishita, Ryo Yonetani

机构 * The Department of Mechanical Systems Engineering, Nagoya University(名古屋大学机械系统工程系) CyberAgent AI Lab(CyberAgent AI实验室)

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.RO

AI总结 VLG-Loc通过视觉-语言模型实现从标注足迹地图的全局定位,结合蒙特卡洛定位框架和概率融合提升环境变化下的鲁棒性。

Comments v2: Updated the citation of SparseLoc from an arXiv preprint to its published version in the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07392 2025-12-19 cs.CL cs.AI 57%

Voice-Interactive Surgical Agent for Multimodal Patient Data Control

多模态患者数据控制的语音交互手术代理

Hyeryun Park, Byung Mo Gu, Jun Hee Lee, Byeong Hyeon Choi, Sekeun Kim, Hyun Koo Kim, Kyungsang Kim

机构 * Image Guided Precision Cancer Surgery Institute, College of Medicine, Korea University(影像引导精准癌症手术研究所,韩国大学医学院) Department of Radiology, Massachusetts General Hospital(放射科,麻省总医院) Department of Thoracic and Cardiovascular Surgery, Korea University Guro Hospital, College of Medicine, Korea University(胸外科和心血管外科,韩国大学Guro医院,韩国大学医学院) Department of Biomedical Sciences, College of Medicine, Korea University(生物医学科学系,韩国大学医学院)

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.AI

AI总结 本文提出一种基于大型语言模型的语音交互手术代理,用于多模态患者数据的高效控制与处理,通过分层多代理框架提升手术流程中的数据操作效率与鲁棒性。

Comments 14 pages, 13 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16826 2025-12-19 cs.CV cs.AI 54%

Next-Generation License Plate Detection and Recognition System using YOLOv8

基于YOLOv8的下一代车牌检测与识别系统

Arslan Amin, Rafia Mumtaz, Muhammad Jawad Bashir, Syed Mohammad Hassan Zaidi

机构 * School of Electrical Engineering Computer Science (SEECS), \ University of Sciences Ghulam Ishaq Khan Institute of Engineering Sciences \ Technology (GIKI), Topi, District, Swabi, \ Pakhtunkhwa 23460, Pakistan, Email:prorector\

专题命中 机器人数据与评测 :robotics(comments,journal_ref);分类 cs.AI、cs.CV

AI总结 本文提出基于YOLOv8的车牌检测与识别系统,通过优化模型配置提升实时准确率,为智能交通系统提供高效解决方案。

Comments 6 pages, 5 figures. Accepted and published in the 2023 IEEE 20th International Conference on Smart Communities: Improving Quality of Life using AI, Robotics and IoT (HONET)

Journal ref 2023 IEEE 20th International Conference on Smart Communities: Improving Quality of Life using AI, Robotics and IoT (HONET)

详情

展开后加载摘要…

URL PDF HTML 收藏