arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

共收录 9157 信号源:cs.RO, cs.CV, cs.AI, cs.LG

1. VLA模型 9157 篇

2511.19236 2025-11-25 cs.RO cs.AI 81%

SENTINEL: A Fully End-to-End Language-Action Model for Humanoid Whole Body Control

SENTINEL:一种用于人形机器人全身控制的端到端语言-动作模型

Yuxuan Wang, Haobin Jiang, Shiqing Yao, Ziluo Ding, Zongqing Lu

机构 * Peking University(北京大学) BeingBeyond

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO、cs.AI

AI总结 SENTINEL是一种端到端语言-动作模型,通过直接映射语言指令和本体感觉输入到低层动作,实现人形机器人全身控制,并支持多模态扩展。

Comments 23 pages, 8 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04055 2025-11-19 q-bio.QM cs.AI cs.LG 81%

Benchmark on Drug Target Interaction Modeling from a Drug Structure Perspective

Xinnan Zhang, Jialin Wu, Junyi Xie, Tianlong Chen, Kaixiong Zhou

机构 * University of Minnesota(明尼苏达大学) University of California San Diego(加州大学圣地亚哥分校) UNC Chapel Hill(北卡罗来纳大学教堂山分校) North Carolina State University(北卡罗来纳州立大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15691 2025-11-13 cs.LG cs.AI 81%

What Do Latent Action Models Actually Learn?

Chuheng Zhang, Tim Pearce, Pushi Zhang, Kaixin Wang, Xiaoyu Chen, Wei Shen, Li Zhao, Jiang Bian

机构 * Microsoft Research(微软研究院) Tsinghua University(清华大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

Comments Accepted by NeurIPS-25

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18865 2025-09-24 cs.RO cs.LG 81%

Bi-VLA: Bilateral Control-Based Imitation Learning via Vision-Language Fusion for Action Generation

Masato Kobayashi, Thanpimon Buamanee

机构 * D3 Center, The University of Osaka(大阪大学D3中心) Graduate School of Information Science and Technology, The University of Osaka(大阪大学信息科学与技术研究生院) Graduate School of Maritime Sciences, Kobe University(兵库大学海运研究生院)

专题命中 VLA模型 :VLA(title,abstract);分类 cs.RO、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14067 2025-09-22 cs.CV cs.AI 81%

VLA-Mark: A cross modal watermark for large vision-language alignment model

Shuliang Liu, Qi Zheng, Jesse Jiaxi Xu, Yibo Yan, Junyan Zhang, He Geng, Aiwei Liu, Peijie Jiang, Jia Liu, Yik-Cheung Tam, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) University of Toronto(多伦多大学) Ant Group, Alibaba(蚂蚁集团,阿里巴巴) New York University Shanghai(纽约大学上海分校)

专题命中 VLA模型 :VLA(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by the main conference, EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16365 2025-09-15 cs.CV cs.AI 81%

JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse

Muyao Li, Zihao Wang, Kaichen He, Xiaojian Ma, Yitao Liang

机构 * Peking University(北京大学) BIGAI(大疆创新)

专题命中 VLA模型 :VLA(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16806 2025-07-28 cs.RO cs.AI 81%

DyWA: Dynamics-adaptive World Action Model for Generalizable Non-prehensile Manipulation

Jiangran Lyu, Ziming Li, Xuesong Shi, Chaoyi Xu, Yizhou Wang, He Wang

机构 * Center on Frontiers of Computing Studies, School of Computer Science, Peking University(前沿计算研究教育部重点实验室,计算机学院,北京大学) Galbot Inst. for Artificial Intelligence, Peking University(人工智能研究所,北京大学) State Key Laboratory of General Artificial Intelligence, Peking University(通用人工智能国家重点实验室,北京大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO、cs.AI

Comments Project Page:https://pku-epic.github.io/DyWA/

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.10705 2025-07-16 cs.LG cs.AI 81%

Learning Safe Numeric Planning Action Models

Argaman Mordoch, Shahaf S. Shperberg, Roni Stern, Berndan Juba

机构 * Software and Information Systems Engineering, Ben-Gurion University of the Negev(本·古里安大学软件与信息系统工程系) Department of Computer Science and Engineering, Washington University(华盛顿大学计算机科学与工程系)

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13152 2025-07-01 cs.CV cs.AI 81%

Interpretable Interaction Modeling for Trajectory Prediction via Agent Selection and Physical Coefficient

Shiji Huang, Lei Ye, Min Chen, Wenhai Luo, Dihong Wang, Chenqi Xu, Deyuan Liang

机构 * College of Computer Science and Technology, Zhejiang University of Technology(计算机科学与技术学院,浙江工业大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by International Conference on Intelligent Robots and Systems (IROS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22391 2025-06-06 cs.LG cs.AI 81%

A Large Recurrent Action Model: xLSTM enables Fast Inference for Robotics Tasks

Thomas Schmied, Thomas Adler, Vihang Patil, Maximilian Beck, Korbinian Pöppel, Johannes Brandstetter, Günter Klambauer, Razvan Pascanu, Sepp Hochreiter

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02298 2025-06-04 cs.CL cs.AI cs.LG 81%

LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback

Thai Hoang, Kung-Hsiang Huang, Shirley Kokane, Jianguo Zhang, Zuxin Liu, Ming Zhu, Jake Grigsby, Tian Lan, Michael S Ryoo, Chien-Sheng Wu, Shelby Heinecke, Huan Wang, Silvio Savarese, Caiming Xiong, Juan Carlos Niebles

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

Comments LAM Simulator framework for agentic data generation

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08613 2025-05-27 cs.CV cs.AI 81%

Cross-Modal Bidirectional Interaction Model for Referring Remote Sensing Image Segmentation

Zhe Dong, Yuzhe Sun, Tianzhu Liu, Wangmeng Zuo, Yanfeng Gu

机构 * School of Electronics and Information Engineering, Harbin Institute of Technology(电子信息工程学院,哈尔滨工业大学) Faculty of Computing, Harbin Institute of Technology(计算机学院,哈尔滨工业大学) Peng Cheng Laboratory, Shenzhen(鹏城实验室,深圳)

专题命中 VLA模型 :action model(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12272 2025-05-20 cs.AI cs.LG 81%

Enhancing Knowledge Graph Completion with GNN Distillation and Probabilistic Interaction Modeling

Lingzhi Wang, Pengcheng Huang, Haotian Li, Yuliang Wei, Guodong Xin, Rui Zhang, Donglin Zhang, Zhenzhou Ji, Wei Wang

机构 * Shandong Key Laboratory of Industrial Network Security(山东工业网络安全重点实验室) School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院)

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00200 2025-04-28 cs.RO cs.CV 81%

Unified Video Action Model

Shuang Li, Yihuai Gao, Dorsa Sadigh, Shuran Song

机构 * Stanford University(斯坦福大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO、cs.CV

Comments Project website: https://unified-video-action-model.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02572 2025-03-05 cs.RO cs.AI 81%

RaceVLA: VLA-based Racing Drone Navigation with Human-like Behaviour

Valerii Serpiva, Artem Lykov, Artyom Myshlyaev, Muhammad Haris Khan, Ali Alridha Abdulkarim, Oleg Sautenkov, Dzmitry Tsetserukou

专题命中 VLA模型 :VLA(title,abstract);分类 cs.RO、cs.AI

Comments 6 pages, 6 figures. Submitted to IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14795 2025-02-24 cs.RO cs.CV 81%

Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration

Pengxiang Ding, Jianfei Ma, Xinyang Tong, Binghong Zou, Xinxin Luo, Yiguo Fan, Ting Wang, Hongchao Lu, Panzhong Mo, Jinxin Liu, Yuefan Wang, Huaicheng Zhou, Wenshuo Feng, Jiacheng Liu, Siteng Huang, Donglin Wang

专题命中 VLA模型 :VLA(title,abstract);分类 cs.RO、cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14891 2025-02-13 cs.RO cs.CV 81%

Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation

Guokang Wang, Hang Li, Shuyuan Zhang, Di Guo, Yanhong Liu, Huaping Liu

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO、cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04408 2025-02-10 cs.LG cs.AI 81%

Transforming Multimodal Models into Action Models for Radiotherapy

Matteo Ferrante, Alessandra Carosi, Rolando Maria D Angelillo, Nicola Toschi

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06067 2024-10-10 cs.CV cs.LG 81%

Contrastive Learning to Fine-Tune Feature Extraction Models for the Visual Cortex

Alex Mulrooney, Austin J. Brockmeier

专题命中 VLA模型 :action model(title,abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.03215 2024-09-06 cs.CL cs.AI cs.LG 81%

xLAM: A Family of Large Action Models to Empower AI Agent Systems

Jianguo Zhang, Tian Lan, Ming Zhu, Zuxin Liu, Thai Hoang, Shirley Kokane, Weiran Yao, Juntao Tan, Akshara Prabhakar, Haolin Chen, Zhiwei Liu, Yihao Feng, Tulika Awalgaonkar, Rithesh Murthy, Eric Hu, Zeyuan Chen, Ran Xu, Juan Carlos Niebles, Shelby Heinecke, Huan Wang, Silvio Savarese, Caiming Xiong

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

Comments Technical report for the Salesforce xLAM model series

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.13627 2024-07-16 cs.CV cs.AI 81%

Vamos: Versatile Action Models for Video Understanding

Shijie Wang, Qi Zhao, Minh Quan Do, Nakul Agarwal, Kwonjoon Lee, Chen Sun

专题命中 VLA模型 :action model(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to ECCV 2024 (European Conference on Computer Vision). Code and models are released at https://brown-palm.github.io/Vamos/

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17802 2024-05-29 cs.LG cs.AI q-bio.BM 81%

Multi-level Interaction Modeling for Protein Mutational Effect Prediction

Yuanle Mo, Xin Hong, Bowen Gao, Yinjun Jia, Yanyan Lan

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.14055 2024-04-08 cs.RO cs.AI 81%

Transforming a Quadruped into a Guide Robot for the Visually Impaired: Formalizing Wayfinding, Interaction Modeling, and Safety Mechanism

J. Taery Kim, Wenhao Yu, Yash Kothari, Jie Tan, Greg Turk, Sehoon Ha

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO、cs.AI

Comments 16 pages, 8 figures

Journal ref Proceedings of The 7th Conference on Robot Learning, PMLR 229:2288-2303, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.16189 2024-01-30 cs.CV cs.RO 81%

FIMP: Future Interaction Modeling for Multi-Agent Motion Prediction

Sungmin Woo, Minjung Kim, Donghyeong Kim, Sungjun Jang, Sangyoun Lee

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO、cs.CV

Comments Accepted by ICRA 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.10660 2023-01-26 cs.LG cs.AI 81%

Interaction Modeling with Multiplex Attention

Fan-Yun Sun, Isaac Kauvar, Ruohan Zhang, Jiachen Li, Mykel Kochenderfer, Jiajun Wu, Nick Haber

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2022, project website: https://cs.stanford.edu/~sunfanyun/imma/

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.08234 2022-11-16 cs.LG cs.AI 81%

Build generally reusable agent-environment interaction models

Jun Jin, Hongming Zhang, Jun Luo

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

Comments Accepted in Foundation Models for Decision Making Workshop at Neural Information Processing Systems, 2022. Slides: https://docs.google.com/presentation/d/1PMS2xwTcztP2pPk1bsjqkQscI39Wy5tpmoE5-_ZC7Fo/edit?usp=sharing

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.13014 2022-09-28 q-bio.BM cs.AI cs.LG q-bio.MN 81%

Predicting Protein-Ligand Binding Affinity via Joint Global-Local Interaction Modeling

Yang Zhang, Gengmo Zhou, Zhewei Wei, Hongteng Xu

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.05492 2019-12-12 cs.AI cs.LG 81%

Neural-Symbolic Descriptive Action Model from Images: The Search for STRIPS

Masataro Asai

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

Comments Technical Report; not going to be submitted to the conference

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.01992 2018-10-05 cs.AI cs.LG 81%

Action Model Acquisition using LSTM

Ankuj Arora, Humbert Fiorino, Damien Pellier, Sylvie Pesty

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1507.04285 2015-07-16 cs.LG cs.AI cs.LO 81%

Learning Action Models: Qualitative Approach

Thomas Bolander, Nina Gierasimczuk

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

Comments 18 pages, accepted for LORI-V: The Fifth International Conference on Logic, Rationality and Interaction, October 28-31, 2015, National Taiwan University, Taipei, Taiwan

详情

展开后加载摘要…

URL PDF HTML 收藏