arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

International Conference on Robotics and Automation · 会议 · Robotics

2026-07-14 至 2026-07-14 共收录 2
2607.10925 2026-07-14 cs.RO 新提交

Mapping Pamir: Multi-Session Visual-Inertial SLAM and 3D Reconstruction of an Underwater Shipwreck

绘制帕米尔:水下沉船的多会话视觉惯性SLAM与3D重建

Michalis Chatzispyrou, Luke Horgan, Hyunkil Hwang, Harish Sathishchandra, Chinmay Burgul, Monika Roznere, Alberto Quattrini Li, Philippos Mordohai, Ioannis Rekleitis

机构 * University of Delaware(特拉华大学) Stevens Institute of Technology(史蒂文斯理工学院) Binghamton University(宾汉姆顿大学) Dartmouth College(达特茅斯学院)

AI总结 该研究利用经济相机与潜水计算机数据,基于SVIn2和COLMAP框架,提出水下环境多会话映射框架,实现巴巴多斯海岸沉船的多会话映射,首次对沉船内外进行映射,第三个会话采用双相机拓宽视野。

Comments 8 pages, 12 figures. Accepted to ICRA 2026, Vienna, Austria

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.10797 2026-07-14 cs.CV 新提交

Compositional Context Fine-Tuning Vision-Language Model for Complex Assembly Action Understanding from Videos

用于从视频理解复杂装配动作的组合上下文微调视觉语言模型

Hao Zheng, Jinyi Huang, Tiantian Zheng, Xun Xu, Tuka Alhanai

机构 * New York University Abu Dhabi(纽约大学阿布扎比分校) The University of Auckland(奥克兰大学)

AI总结 针对装配动作理解难题,提出组合上下文微调方法及层划分交替训练方法,将动作分解为语义元素并微调视觉语言模型,创建相关数据集,实验证明该方法优于基线且能提供可解释预测,助力人机协作装配。

Comments Accepted by ICRA 2026. Video Understanding; Vision Language Model; Multi-modal LLM; Action Recognition; Assembly

详情

展开后加载摘要…

URL PDF HTML 收藏