arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

共收录 9832 信号源:cs.RO, cs.CV, cs.AI, cs.LG

1. VLA模型 9157 篇

2504.09298 2025-04-15 cs.CV 57%

A Lightweight Moment Retrieval System with Global Re-Ranking and Robust Adaptive Bidirectional Temporal Search

Tinh-Anh Nguyen-Nhu, Huu-Loc Tran, Nguyen-Khang Le, Minh-Nhat Nguyen, Tien-Huy Nguyen, Hoang-Long Nguyen-Huu, Huu-Phong Phan-Nguyen, Huy-Thach Pham, Quan Nguyen, Hoang M. Le, Quang-Vinh Dinh

专题命中 VLA模型 :action model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20211 2025-04-15 cs.LG cs.IR 57%

Generative Regression Based Watch Time Prediction for Short-Video Recommendation

Hongxu Ma, Kai Tian, Tao Zhang, Xuefeng Zhang, Han Zhou, Chunjie Chen, Han Li, Jihong Guan, Shuigeng Zhou

专题命中 VLA模型 :action model(abstract);分类 cs.LG

Comments 10 pages, 5 figures, conference or other essential info

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06771 2025-04-10 cs.HC cs.AI 57%

AI, Help Me Think$\unicode{x2014}$but for Myself: Assisting People in Complex Decision-Making by Providing Different Kinds of Cognitive Support

Leon Reicherts, Zelun Tony Zhang, Elisabeth von Oswald, Yuanting Liu, Yvonne Rogers, Mariam Hassib

专题命中 VLA模型 :action model(abstract);分类 cs.AI

Comments To be published at ACM CHI 2025 Conference on Human Factors in Computing Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05303 2025-04-08 cs.CV 57%

InteractVLM: 3D Interaction Reasoning from 2D Foundational Models

Sai Kumar Dwivedi, Dimitrije Antić, Shashank Tripathi, Omid Taheri, Cordelia Schmid, Michael J. Black, Dimitrios Tzionas

专题命中 VLA模型 :action model(abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05076 2025-04-08 cs.CV 57%

Content-Distortion High-Order Interaction for Blind Image Quality Assessment

Shuai Liu, Qingyu Mao, Chao Li, Jiacong Chen, Fanyang Meng, Yonghong Tian, Yongsheng Liang

专题命中 VLA模型 :action model(abstract);分类 cs.CV

Comments 19 pages (main text: 14 pages + appendix: 5 pages), 9 figures, 23 tables. In submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04780 2025-04-08 cs.CV 57%

Bottom-Up Scattering Information Perception Network for SAR target recognition

Chenxi Zhao, Daochang Wang, Siqian Zhang, Gangyao Kuang

专题命中 VLA模型 :action model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08646 2025-04-01 cs.CV 57%

StreamChat: Chatting with Streaming Video

Jihao Liu, Zhiding Yu, Shiyi Lan, Shihao Wang, Rongyao Fang, Jan Kautz, Hongsheng Li, Jose M. Alvare

专题命中 VLA模型 :action model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11935 2025-03-24 cs.CV 57%

Design of an Expression Recognition Solution Based on the Global Channel-Spatial Attention Mechanism and Proportional Criterion Fusion

Jun Yu, Yang Zheng, Lei Wang, Yongqi Wang, Shengfan Xu

专题命中 VLA模型 :action model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15432 2025-03-20 cond-mat.mtrl-sci cs.LG 57%

Accurate, transferable, and verifiable machine-learned interatomic potentials for layered materials

Johnathan D. Georgaras, Akash Ramdas, Chung Hsuan Shan, Elena Halsted, Berwyn, Tianshu Li, Felipe H. da Jornada

专题命中 VLA模型 :action model(abstract);分类 cs.LG

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13389 2025-03-18 cs.LG physics.geo-ph 57%

Investigating the effect of CPT in lateral spreading prediction using Explainable AI

Cheng-Hsi Hsiao, Ellen Rathje, Krishna Kumar

专题命中 VLA模型 :action model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06861 2025-03-11 cs.CL cs.AI 57%

Enhanced Multi-Tuple Extraction for Alloys: Integrating Pointer Networks and Augmented Attention

Mengzhe Hei, Zhouran Zhang, Qingbao Liu, Yan Pan, Xiang Zhao, Yongqian Peng, Yicong Ye, Xin Zhang, Shuxin Bai

专题命中 VLA模型 :action model(abstract);分类 cs.AI

Comments 17 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06565 2025-03-11 cs.CV 57%

Future-Aware Interaction Network For Motion Forecasting

Shijie Li, Xun Xu, Si Yong Yeo, Xulei Yang

专题命中 VLA模型 :action model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.10726 2025-03-10 cs.AI 57%

Planning Domain Model Acquisition from State Traces without Action Parameters

Tomáš Balyo, Martin Suda, Lukáš Chrpa, Dominik Šafránek, Stephan Gocht, Filip Dvořák, Roman Barták, G. Michael Youngblood

专题命中 VLA模型 :action model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04170 2025-03-07 cs.ET cs.AI 57%

Towards Intelligent Transportation with Pedestrians and Vehicles In-the-Loop: A Surveillance Video-Assisted Federated Digital Twin Framework

Xiaolong Li, Jianhao Wei, Haidong Wang, Li Dong, Ruoyang Chen, Changyan Yi, Jun Cai, Dusit Niyato, Xuemin, Shen

专题命中 VLA模型 :action model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03512 2025-03-06 cs.CL cs.LG 57%

An Aspect Extraction Framework using Different Embedding Types, Learning Models, and Dependency Structure

Ali Erkan, Tunga Güngör

专题命中 VLA模型 :action model(abstract);分类 cs.LG

Comments Aspect-based Sentiment Analysis, Aspect Extraction, Natural Language Processing, Machine Learning, Deep Neural Networks, Turkish

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03204 2025-03-06 cs.CV 57%

Find Matching Faces Based On Face Parameters

Setu A. Bhatt, Harshadkumar B. Prajapati, Vipul K. Dabhi, Ankush Tyagi

专题命中 VLA模型 :action model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.06666 2025-03-04 cs.CL cs.AI cs.SD eess.AS 57%

LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Qingkai Fang, Shoutao Guo, Yan Zhou, Zhengrui Ma, Shaolei Zhang, Yang Feng

专题命中 VLA模型 :action model(abstract);分类 cs.AI

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17443 2025-02-26 cs.SE cs.AI 57%

AI Agentic workflows and Enterprise APIs: Adapting API architectures for the age of AI agents

Vaibhav Tupe, Shrinath Thube

专题命中 VLA模型 :action model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15956 2025-02-25 cs.CV 57%

Human Motion Prediction, Reconstruction, and Generation

Canxuan Gang, Yiran Wang

专题命中 VLA模型 :action model(abstract);分类 cs.CV

Comments Tech report

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11490 2025-02-18 cs.LG cs.DC cs.IR 57%

GPU-accelerated Multi-relational Parallel Graph Retrieval for Web-scale Recommendations

Zhuoning Guo, Guangxing Chen, Qian Gao, Xiaochao Liao, Jianjia Zheng, Lu Shen, Hao Liu

专题命中 VLA模型 :action model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10620 2025-02-18 cs.AI 57%

ProMRVL-CAD: Proactive Dialogue System with Multi-Round Vision-Language Interactions for Computer-Aided Diagnosis

Xueshen Li, Xinlong Hou, Ziyi Huang, Yu Gan

专题命中 VLA模型 :action model(abstract);分类 cs.AI

Comments 17 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02945 2025-02-06 cs.CL cs.AI 57%

LLM-KT: Aligning Large Language Models with Knowledge Tracing using a Plug-and-Play Instruction

Ziwei Wang, Jie Zhou, Qin Chen, Min Zhang, Bo Jiang, Aimin Zhou, Qinchun Bai, Liang He

专题命中 VLA模型 :action model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13295 2025-01-24 cs.AI 57%

Parallel Belief Contraction via Order Aggregation

Jake Chandler, Richard Booth

专题命中 VLA模型 :action model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.03819 2025-01-09 cs.CV 57%

Cross-Skeleton Interaction Graph Aggregation Network for Representation Learning of Mouse Social Behaviour

Feixiang Zhou, Xinyu Yang, Fang Chen, Long Chen, Zheheng Jiang, Hui Zhu, Reiko Heckel, Haikuan Wang, Minrui Fei, Huiyu Zhou

专题命中 VLA模型 :action model(abstract);分类 cs.CV

Comments Accepted to IEEE Transactions on Image Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08827 2024-12-31 cs.CV 57%

RGBT Tracking via All-layer Multimodal Interactions with Progressive Fusion Mamba

Andong Lu, Wanyu Wang, Chenglong Li, Jin Tang, Bin Luo

专题命中 VLA模型 :action model(abstract);分类 cs.CV

Comments Accepted by AAAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.09261 2024-12-24 cs.LG 57%

MSHyper: Multi-Scale Hypergraph Transformer for Long-Range Time Series Forecasting

Zongjiang Shang, Ling Chen, Binqing Wu, Dongliang Cui

专题命中 VLA模型 :action model(abstract);分类 cs.LG

Comments 12 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.02291 2024-12-24 cs.CV 57%

Focusing on what to decode and what to train: SOV Decoding with Specific Target Guided DeNoising and Vision Language Advisor

Junwen Chen, Yingcheng Wang, Keiji Yanai

专题命中 VLA模型 :VLA(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10846 2024-12-17 cs.CV cs.HC 57%

Detecting Activities of Daily Living in Egocentric Video to Contextualize Hand Use at Home in Outpatient Neurorehabilitation Settings

Adesh Kadambi, José Zariffa

专题命中 VLA模型 :action model(abstract);分类 cs.CV

Comments To be submitted to IEEE Transactions on Neural Systems and Rehabilitation Engineering. 11 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.11777 2024-12-12 cs.AI 57%

What you get is what you see: Decomposing Epistemic Planning using Functional STRIPS

Guang Hu, Tim Miller, Nir Lipovetzky

专题命中 VLA模型 :action model(abstract);分类 cs.AI

Comments 20 pages, 3 figures, 4 experiments, journal paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09914 2024-11-18 cs.CR cs.CV 57%

mmSpyVR: Exploiting mmWave Radar for Penetrating Obstacles to Uncover Privacy Vulnerability of Virtual Reality

Luoyu Mei, Ruofeng Liu, Zhimeng Yin, Qingchuan Zhao, Wenchao Jiang, Shuai Wang, Kangjie Lu, Tian He

专题命中 VLA模型 :action model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏