arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

共收录 9866 信号源:cs.RO, cs.CV, cs.AI, cs.LG

1. VLA模型 9186 篇

2412.11337 2024-12-17 cs.RO cs.AI cs.CV 75%

Modality-Driven Design for Multi-Step Dexterous Manipulation: Insights from Neuroscience

Naoki Wake, Atsushi Kanehira, Daichi Saito, Jun Takamatsu, Kazuhiro Sasabuchi, Hideki Koike, Katsushi Ikeuchi

专题命中 VLA模型 :vision-language-action(abstract);action model(abstract);分类 cs.RO、cs.CV、cs.AI

Comments 8 pages, 5 figures, 2 tables. Last updated on December 14th, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.12871 2024-05-10 cs.CV cs.AI cs.CL cs.LG 75%

An Embodied Generalist Agent in 3D World

Jiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu, Puhao Li, Yan Wang, Qing Li, Song-Chun Zhu, Baoxiong Jia, Siyuan Huang

专题命中 VLA模型 :vision-language-action(abstract);VLA(abstract);分类 cs.CV、cs.AI、cs.LG

Comments ICML 2024. The first four authors contribute equally. Project page: https://embodied-generalist.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12208 2026-04-15 cs.RO cs.AI 74%

Unveiling the Surprising Efficacy of Navigation Understanding in End-to-End Autonomous Driving

揭示端到端自动驾驶中导航理解的惊人效能

Zhihua Hua, Junli Wang, Pengfei LI, Qihao Jin, Bo Zhang, Kehua Sheng, Yilun Chen, Zhongxue Gan, Wenchao Ding

机构 * College of Intelligent Robotics and Advanced Manufacturing, Fudan University(复旦大学智能机器人与先进制造学院) Didi Chuxing(滴滴出行) Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 VLA模型 :VLA(abstract,abstract_cn);分类 cs.RO、cs.AI

AI总结 本文提出SNG框架,通过真实导航模式高效表示全局导航信息,结合导航路径与分步信息,提升自动驾驶的全局与局部规划能力,实现无需辅助损失函数的高精度导航建模。

Comments 8 pages, 6 figures. ICRA 2026. Code available at https://fudan-magic-lab.github.io/SNG-VLA-web

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27933 2026-08-07 cs.AI 版本更新 74%

The Geometry of Flow-Matching Uncertainty: A Cost-free Uncertainty Proxy and Its Application in Flow-based VLA Failure Detection

流匹配不确定性的几何本质与自由代理

Ziyang Rao, Yiren Zhao, Weiyu Guo, Ben Fei, Yandong Guo, Hui Xiong

专题命中 VLA模型 :VLA(title);分类 cs.AI

AI总结 该研究针对流匹配(FM)不确定性估计方法的缺陷,提出去噪加速度(accel)作为高泛化无成本的不确定性代理,可提前识别FM生成动作的失败,性能优于相关基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28762 2026-08-03 cs.LG 新提交 74%

Feature Interaction Modeling for Physics-Informed Neural Networks and Neural Operators

面向物理信息神经网络与神经算子的特征交互建模

Quan Gu, Hongxia Liu

专题命中 VLA模型 :action model(title);分类 cs.LG

AI总结 该研究将因子分解机衍生的特征交互模块嵌入物理信息神经网络与神经算子,提出FM-PINN、FM-Operator等模型,提升了激波主导等问题的PDE解近似精度,为相关物理建模提供了新方向。

Comments 37 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17568 2026-05-21 cs.LG 74%

Structured Neural Marked Point Processes for Interpretable Event Interaction Modeling

结构化神经标记点过程用于可解释的事件交互建模

Zhitong Xu, Qiwei Yuan, Yinghao Chen, Shandian Zhe, Bin Shen

机构 * Kahlert School of Computing, University of Utah(犹他大学计算学院) Celonis AI

专题命中 VLA模型 :action model(title);分类 cs.LG

AI总结 本文提出了一种结构化神经标记点过程(SNMPP),通过显式发现事件级和类别级的关系,实现高灵活性的建模,同时在合成和现实数据集上验证了其揭示结构关系和强预测性能的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18627 2026-05-19 cs.AI 74%

Learning Lifted Action Models from Traces with Minimal Information About Actions and States

从动作轨迹中学习提升的动作模型:最少关于动作和状态的信息

Jonas Gösgens, Niklas Jansen, Hector Geffner

机构 * RWTH Aachen University(亚琛工业大学)

专题命中 VLA模型 :action model(title);分类 cs.AI

AI总结 本文研究了在不完全信息下从动作轨迹中学习STRIPS+动作域的问题,提出了三种通用情况下的算法和完备性结果,假设选定的动作参数完全可观察,从而在不同可观察性假设下确定等效域的学习条件。

Comments accepted at KR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24188 2026-04-28 cs.RO cs.GR 74%

Generalizable Friction Coefficient Estimation via Material Embedding and Proxy Interaction Modeling

通过材料嵌入和代理交互建模实现通用摩擦系数估计

Zhendong Wang, Huamin Wang

机构 * Style3D Research(Style3D研究院)

专题命中 VLA模型 :action model(title);分类 cs.RO

AI总结 本文提出基于代理材料的摩擦系数估计框架,通过材料嵌入和融合函数实现高效预测,减少实验成本并提升鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17058 2026-03-19 cs.GT cs.MA cs.RO cs.SY eess.SY math.OC 74%

Asymmetric Nash Seeking via Best Response Maps: Global Linear Convergence and Robustness to Inexact Reaction Models

不对称纳什寻求 via 最佳反应映射:全局线性收敛性与对不精确反应模型的鲁棒性

Mahdis Rabbani, Navid Mojahed, Shima Nazari

机构 * Department of Mechanical and Aerospace Engineering, University of California, Davis(加州大学戴维斯分校机械与航空航天工程系)

专题命中 VLA模型 :action model(title);分类 cs.RO

AI总结 本文研究了不对称信息双玩家约束博弈的全局线性收敛性和对不精确反应模型的鲁棒性,提出了一种不对称投影梯度下降-最佳反应迭代方法,证明了在最佳反应映射精确时的全局线性收敛性,并分析了不精确情况下的误差界。

Comments 6 Pages, 2 Figures, Preprint submitted to IEEE L-CSS and CDC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00110 2026-03-03 cs.RO 74%

Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation

从预训练视频模型中学习物理:一种多模态连续和序列世界交互模型用于机器人操作

Zijian Song, Qichang Li, Sihan Qin, Yuhao Chen, Tianshui Chen, Liang Lin, Guangrun Wang

机构 * Sun Yat-sen University(中山大学) Guangdong Key Laboratory of Big Data Analysis(广东大数据分析与处理重点实验室) X-Era AI Lab(X-Era人工智能实验室) Guangdong University of Technology(广东工业大学)

专题命中 VLA模型 :action model(title);分类 cs.RO

AI总结 PhysGen通过预训练视频模型学习物理知识,实现机器人操作的连续和序列世界交互,优于现有基线方法。

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15771 2026-01-23 cs.LG q-bio.BM 74%

Rethinking Drug-Drug Interaction Modeling as Generalizable Relation Learning

重新思考药物-药物相互作用建模作为可推广的关系学习

Dong Xu, Jiantao Wu, Qihua Pan, Sisi Yuan, Zexuan Zhu, Junkai Ji

机构 * School of Artificial Intelligence, Shenzhen University(人工智能学院,深圳大学)

专题命中 VLA模型 :action model(title);分类 cs.LG

AI总结 本文提出GenRel-DDI框架,通过关系学习方法提升药物-药物相互作用预测的泛化能力,显著优于现有方法。

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18239 2026-01-21 cs.IR cs.LG 74%

LIME: Link-based user-item Interaction Modeling with decoupled xor attention for Efficient test time scaling

基于解耦异或注意力的链接式用户-物品交互建模

Yunjiang Jiang, Ayush Agarwal, Yang Liu, Bi Xue

机构 * Meta

专题命中 VLA模型 :action model(title);分类 cs.LG

AI总结 LIME通过解耦异或注意力机制,实现了高效推荐系统中计算复杂度的降低,提升了推理速度和用户参与度。

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07003 2026-01-13 stat.ME cs.LG stat.CO 74%

Unity Forests: Improving Interaction Modelling and Interpretability in Random Forests

统一森林:改进随机森林中的交互建模和可解释性

Roman Hornung, Alexander Hapfelmeier

机构 * Institute for Medical Information Processing, Biometry and Epidemiology, Faculty of Medicine, Ludwig Maximilian University of Munich (LMU)(慕尼黑路德维希-马克西米利安大学医学信息处理与流行病学研究所) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) Institute of General Practice and Health Services Research, Department Clinical Medicine, TUM School of Medicine and Health, Technical University of Munich (TUM)(慕尼黑技术大学医学与健康学院一般医学与健康服务研究研究所) Institute of AI and Informatics in Medicine, TUM School of Medicine and Health, Technical University of Munich (TUM)(慕尼黑技术大学医学与健康学院医学人工智能与信息学研究所)

专题命中 VLA模型 :action model(title);分类 cs.LG

AI总结 统一森林通过优化树根分割提升交互建模和可解释性,改进了随机森林的变量重要性度量和预测性能。

Comments 33 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03828 2025-12-04 cs.RO 74%

IM HERE: Interaction Model for Human Effort Based Robot Engagement

IM HERE: 以人为中心的机器人互动模型

Dominykas Strazdas, Magnus Jung, Jan Marquenie, Ingo Siegert, Ayoub Al-Hamadi

机构 * Neuro-Information Technology Otto von Guericke University(神经信息技术奥托·冯·格里克大学) Mobile Dialog Systems Otto von Guericke University(移动对话系统奥托·冯·格里克大学)

专题命中 VLA模型 :action model(title);分类 cs.RO

AI总结 IM HERE提出了一种基于人类努力的机器人互动模型,旨在通过建模社交行为来实现自主系统与社会规范的协调。

Comments 8 pages, 5 figures

Journal ref 2025 IEEE Conference on Cognitive and Computational Aspects of Situation Management (CogSIMA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09553 2025-11-20 cs.CV 74%

GLD-Road:A global-local decoding road network extraction model for remote sensing images

Ligao Deng, Yupeng Deng, Yu Meng, Jingbo Chen, Zhihao Xi, Diyou Liu, Qifeng Chu

机构 * Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航空信息研究所) School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences(中国科学院大学电子电气与通信工程学院) Heilongjiang Geographic Information Engineering Institute(黑龙江地理信息工程研究所)

专题命中 VLA模型 :action model(title);分类 cs.CV

Journal ref ISPRS J. Photogramm. Remote Sens. 228, 741-755 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02728 2025-11-07 cs.RO 74%

Team Xiaomi EV-AD VLA: Caption-Guided Retrieval System for Cross-Modal Drone Navigation -- Technical Report for IROS 2025 RoboSense Challenge Track 4

Lingfeng Zhang, Erjia Xiao, Yuchen Zhang, Haoxiang Fu, Ruibin Hu, Yanbiao Ma, Wenbo Ding, Long Chen, Hangjun Ye, Xiaoshuai Hao

机构 * Tsinghua University(清华大学) Xiaomi EV(小米电动车) Georgia Institute of Technology(佐治亚理工学院) National University of Singapore(新加坡国立大学) The Chinese University of Hong Kong(香港中文大学) Renmin University of China(中国人民大学)

专题命中 VLA模型 :VLA(title);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.03655 2025-11-04 stat.ML cs.LG 74%

Bayesian Additive Main Effects and Multiplicative Interaction Models using Tensor Regression for Multi-environmental Trials

Antonia A. L. Dos Santos, Danilo A. Sarti, Rafael A. Moral, Andrew C. Parnell

机构 * Hamilton Institute, Department of Mathematics(哈里顿研究所,数学系) Statistics, Maynooth University, Ireland(统计学,梅诺特大学,爱尔兰) School of Mathematics(数学学院) Statistics, Insight Centre for Data Analytics, University College Dublin, Ireland(统计学,洞察数据分析师中心,都柏林大学学院,爱尔兰)

专题命中 VLA模型 :action model(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10458 2025-10-02 cs.CV cs.CL cs.HC 74%

GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents

Run Luo, Lu Wang, Wanwei He, Longze Chen, Jiaming Li, Xiaobo Xia

机构 * Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所) University of Chinese Academy of Sciences(中国科学院大学) National University of Singapore(新加坡国立大学)

专题命中 VLA模型 :action model(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21449 2025-09-01 cs.AI 74%

Learning Lifted Action Models From Traces of Incomplete Actions and States

Niklas Jansen, Jonas Gösgens, Hector Geffner

机构 * RWTH Aachen University(亚琛工业大学)

专题命中 VLA模型 :action model(title);分类 cs.AI

Comments To be presented at KR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.00248 2025-06-04 cs.IR cs.AI 74%

TransAct: Transformer-based Realtime User Action Model for Recommendation at Pinterest

Xue Xia, Pong Eksombatchai, Nikil Pancha, Dhruvil Deven Badani, Po-Wei Wang, Neng Gu, Saurabh Vishwas Joshi, Nazanin Farahpour, Zhiyuan Zhang, Andrew Zhai

专题命中 VLA模型 :action model(title);分类 cs.AI

Comments \c{opyright} {ACM} {2023}. This is the author's version of the work. It is posted here for your personal use. Not for redistribution. The definitive Version of Record was published in KDD'23, http://dx.doi.org/10.1145/3580305.3599918

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07352 2025-05-12 cs.LG stat.ML 74%

Generating Origin-Destination Matrices in Neural Spatial Interaction Models

Ioannis Zachos, Mark Girolami, Theodoros Damoulas

机构 * Department of Engineering, Cambridge University(剑桥大学工程系) The Alan Turing Institute(艾伦·图灵研究所) Departments of Statistics & Computer Science, University of Warwick(沃里克大学统计学与计算机科学系)

专题命中 VLA模型 :action model(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17453 2025-03-25 cs.CV 74%

Feature-Based Dual Visual Feature Extraction Model for Compound Multimodal Emotion Recognition

Ran Liu, Fengyu Zhang, Cong Yu, Longjiang Yang, Zhuofan Wen, Siyuan Zhang, Hailiang Yao, Shun Chen, Zheng Lian, Bin Liu

专题命中 VLA模型 :action model(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09247 2025-02-14 cs.CL cs.AI 74%

The Joint Entity-Relation Extraction Model Based on Span and Interactive Fusion Representation for Chinese Medical Texts with Complex Semantics

Danni Feng, Runzhi Li, Jing Wang, Siyu Yan, Lihong Ma, Yunli Xing

专题命中 VLA模型 :action model(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.17740 2024-12-09 cs.IR cs.AI 74%

All-in-One: Heterogeneous Interaction Modeling for Cold-Start Rating Prediction

Shuheng Fang, Kangfei Zhao, Yu Rong, Zhixun Li, Jeffrey Xu Yu

专题命中 VLA模型 :action model(title);分类 cs.AI

Comments 14 pages, 9 figures

Journal ref ICDE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23452 2024-11-01 cs.CL cs.AI 74%

Graph-Augmented Relation Extraction Model with LLMs-Generated Support Document

Vicky Dong, Hao Yu, Yao Chen

专题命中 VLA模型 :action model(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.12610 2024-11-01 cs.RO 74%

A Joint Modeling of Vision-Language-Action for Target-oriented Grasping in Clutter

Kechun Xu, Shuqi Zhao, Zhongxiang Zhou, Zizhang Li, Huaijin Pi, Yue Wang, Rong Xiong

专题命中 VLA模型 :vision-language-action(title);分类 cs.RO

Comments Accepted by ICRA 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01854 2024-10-04 eess.IV cs.CV 74%

A Novel Feature Extraction Model for the Detection of Plant Disease from Leaf Images in Low Computational Devices

Rikathi Pal, Anik Basu Bhaumik, Arpan Murmu, Sanoar Hossain, Biswajit Maity, Soumya Sen

专题命中 VLA模型 :action model(title);分类 cs.CV

Comments 10 Pages, 8 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11406 2024-09-20 cs.LG cs.CR q-fin.ST 74%

Designing an attack-defense game: how to increase robustness of financial transaction models via a competition

Alexey Zaytsev, Maria Kovaleva, Alex Natekin, Evgeni Vorsin, Valerii Smirnov, Georgii Smirnov, Oleg Sidorshin, Alexander Senin, Alexander Dudin, Dmitry Berestnev

专题命中 VLA模型 :action model(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10848 2024-06-18 cs.RO cs.HC 74%

A transparency-based action model implemented in a robotic physical trainer for improved HRI

Aharony Naama, Krakovski Maya, Edan Yael

专题命中 VLA模型 :action model(title);分类 cs.RO

Comments 22 pages, 15 figures, 2 tables and 2 APPENDICES

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.17626 2024-05-08 cs.LG q-bio.QM stat.AP stat.CO 74%

Using Pre-training and Interaction Modeling for ancestry-specific disease prediction in UK Biobank

Thomas Le Menestrel, Erin Craig, Robert Tibshirani, Trevor Hastie, Manuel Rivas

专题命中 VLA模型 :action model(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏