arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

2025-09-09 至 2025-09-09 共收录 8 信号源:cs.RO, cs.CV, cs.AI, cs.LG

1. VLA模型 6 篇

2509.05578 2025-09-09 cs.AI cs.RO 85%

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision

Ruixun Liu, Lingyu Kong, Derun Li, Hang Zhao

机构 * Shanghai Qi Zhi Institute(上海启智研究所) Xi’an Jiaotong University(西安交通大学) Fudan University(复旦大学) Shanghai Jiao Tong University(上海交通大学) Tsinghua University(清华大学)

专题命中 VLA模型 :vision-language-action(title);action model(title);分类 cs.RO、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19757 2025-09-09 cs.RO cs.CV 84%

Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Zhi Hou, Tianyi Zhang, Yuwen Xiong, Haonan Duan, Hengjun Pu, Ronglei Tong, Chengyang Zhao, Xizhou Zhu, Yu Qiao, Jifeng Dai, Yuntao Chen

机构 * Shanghai AI Lab(上海人工智能实验室) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) MMLab, The Chinese University of Hong Kong(香港中文大学MMLab) Peking University(北京大学) SenseTime Research(商汤科技研究院) Tsinghua University(清华大学) HKISI, CAS(中国科学院香港中文大学研究所)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(abstract);分类 cs.RO、cs.CV

Comments Preprint; https://robodita.github.io; To appear in ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06268 2025-09-09 q-bio.MN 50%

Fast-Slow Analysis of a Model For the Stimulation of Enzymatic Activity by a Competitive Inhibitor

Garrett Young, Mitchell Riley, Colleen Mitchell

专题命中 VLA模型 :action model(abstract)

Comments 26 pages including references, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06074 2025-09-09 cs.CL 50%

Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis

Zhenqi Jia, Rui Liu, Berrak Sisman, Haizhou Li

机构 * Inner Mongolia University(内蒙古大学) Center for Language and Speech Processing (CLSP)(语言与语音处理中心) School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)人工智能学院)

专题命中 VLA模型 :action model(abstract)

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14388 2025-09-09 hep-ex nucl-ex 50%

A simulation model investigation of neutron-oxygen inelastic scattering and subsequent nucleus deexcitation based on experimental data

Y. Hino, Y. Ashida, T. Tano, Y. Koshio

专题命中 VLA模型 :action model(abstract)

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05675 2025-09-09 math.NA cs.DC cs.NA 50%

Workflow for High-Fidelity Dynamic Analysis of Structures with Pile Foundation

Amin Pakzad, Pedro Arduino, Wenyang Zhang, Ertugrul Tacirouglu

专题命中 VLA模型 :action model(abstract)

Comments 8 pages, 20 figures, conference paper, Proceedings of the XVII PCSMGE

Journal ref Proceedings of the 17th Pan-American Conference on Soil Mechanics and Geotechnical Engineering, 2024, 1733-1740

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 动作表示与策略 1 篇

2508.17230 2025-09-09 cs.CV 57%

4D Visual Pre-training for Robot Learning

Chengkai Hou, Yanjie Ze, Yankai Fu, Zeyu Gao, Songbo Hu, Yue Yu, Shanghang Zhang, Huazhe Xu

机构 * Peking University(北京大学) Tsinghua University(清华大学) Shanghai Qizhi Institute(上海启智研究院) CASIA(中国科学院自动化研究所) Shanghai AI Lab(上海人工智能实验室)

专题命中 动作表示与策略 :vision-language-action(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 数据集与评测 1 篇

2509.05513 2025-09-09 cs.CV cs.AI cs.RO 67%

OpenEgo: A Large-Scale Multimodal Egocentric Dataset for Dexterous Manipulation

Ahad Jawaid, Yu Xiang

机构 * Department of Computer Science, The University of Texas at Dallas(德克萨斯大学达拉斯分校计算机科学系) Physical Automation, Inc.(Physical Automation 公司)

专题命中 数据集与评测 :vision-language-action(abstract);分类 cs.RO、cs.CV、cs.AI

Comments 4 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏