arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

2025-08-07 至 2025-08-07 共收录 1 信号源:cs.RO, cs.CV, cs.AI, cs.LG

1. 数据集与评测 1 篇

2508.04681 2025-08-07 cs.CV 70%

Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions

Liang Xu, Chengqun Yang, Zili Lin, Fei Xu, Yifan Liu, Congsheng Xu, Yiyi Zhang, Jie Qin, Xingdong Sheng, Yunhui Liu, Xin Jin, Yichao Yan, Wenjun Zeng, Xiaokang Yang

机构 * MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University(人工智能教育部重点实验室,上海交通大学AI研究院) Ningbo Institute of Digital Twin, Eastern Institute of Technology(宁波数字孪生研究所,东部技术研究所) Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生关键实验室) MoE Key Lab of AI, School of Computer Science, Shanghai Jiao Tong University(人工智能教育部重点实验室,上海交通大学计算机学院) Nanjing University of Aeronautics and Astronautics(南京航空航天大学) Lenovo(联想公司)

专题命中 数据集与评测 :vision-language-action(abstract);action model(abstract);分类 cs.CV

Comments Accepted to ICCV 2025. Project Page: https://liangxuy.github.io/InterVLA/

详情

展开后加载摘要…

URL PDF HTML 收藏