arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

2026-02-06 至 2026-02-06 共收录 7 信号源:cs.RO, cs.CV, cs.AI, cs.LG

1. VLA模型 6 篇

2602.05049 2026-02-06 cs.CV cs.AI cs.LG cs.RO 90%

VISTA: Enhancing Visual Conditioning via Track-Following Preference Optimization in Vision-Language-Action Models

VISTA: 通过视觉跟踪偏好优化增强视觉条件化在视觉-语言-动作模型中

Yiye Chen, Yanan Jian, Xiaoyi Dong, Shuxin Cao, Jing Wu, Patricio Vela, Benjamin E. Lundell, Dongdong Chen

机构 * Nvidia(Nvidia公司) Microsoft(微软公司) University of Oxford(牛津大学)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract,comments);分类 cs.RO、cs.CV、cs.AI

AI总结 VISTA通过视觉跟踪偏好优化提升视觉条件化,增强VLA模型在视觉-动作对齐和任务性能上的表现。

Comments In submission. Project website: https://vista-vla.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05966 2026-02-06 cs.CV cs.AI 62%

LSA: Localized Semantic Alignment for Enhancing Temporal Consistency in Traffic Video Generation

LSA:局部语义对齐用于增强交通视频生成中的时间一致性

Mirlan Karimov, Teodora Spasojevic, Markus Braun, Julian Wiederer, Vasileios Belagiannis, Marc Pollefeys

机构 * Mercedes-Benz AG(梅赛德斯-奔驰集团) ETH Zurich(苏黎世联邦理工学院) Friedrich-Alexander University Erlangen-Nuremberg(埃朗根-纽伦堡弗里德里希-亚历山大大学)

专题命中 VLA模型 :action model(abstract);分类 cs.CV、cs.AI

AI总结 LSA通过局部语义对齐提升交通视频生成的时间一致性,无需外部控制信号。

Comments Accepted to IEEE IV 2026. 8 pages, 3 figures. Code available at https://github.com/mirlanium/LSA

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05560 2026-02-06 eess.SP physics.ao-ph 50%

Depth estimation of a monoharmonic source using a vertical linear array at fixed distance

利用垂直线阵在固定距离上估计单谐源深度

Yangjin Xu, Wei Gao, Xiaolei Li, Qinghang Zeng

专题命中 VLA模型 :VLA(abstract)

AI总结 本文提出了一种基于正交约束模态搜索的深度估计方法,用于在未知海床参数条件下利用垂直线阵估计单谐源深度,并通过实验验证了其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05329 2026-02-06 cs.CR 50%

SynAT: Enhancing Security Knowledge Bases via Automatic Synthesizing Attack Tree from Crowd Discussions

SynAT: 通过自动合成 crowds 的攻击树来增强安全知识库

Ziyou Jiang, Lin Shi, Guowei Yang, Xuyan Ma, Fenglong Li, Qing Wang

专题命中 VLA模型 :action model(abstract)

AI总结 SynAT通过自动合成攻击树提升安全知识库,利用LLM和提示学习提取安全讨论中的事件和关系,有效增强安全防护能力。

Comments 28 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03655 2026-02-06 physics.bio-ph 50%

V-Reactor Dynamics: Dual Chaotic Systems and Synchronizing Human Defenses with Viral Evolution

V-Reactor Dynamics: 双重混沌系统与通过病毒进化同步人类防御

Yong-Shou Chen

专题命中 VLA模型 :action model(abstract)

AI总结 V-Reactor Dynamics通过构建双重混沌系统,利用反应参数ρ预测病毒传播动态,实现提前预警和主动防御大流行病。

Comments 6 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08845 2026-02-06 hep-ph hep-ex hep-lat nucl-ex nucl-th 50%

Exclusive photoproduction of light and heavy vector mesons: thresholds to very high energies

轻重向量介子的独占光电产生:从阈值到极高能级

Lin Tang, Hui-Yu Xing, Minghui Ding, Craig D. Roberts

专题命中 VLA模型 :action model(abstract)

AI总结 本文提出了一种模型,用于描述γ + p → V + p反应,统一了不同介子的截面数据,并探讨了其在高能下的行为及与胶子分布等物理量的联系。

Comments 22 pages, 19 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 部署与泛化 1 篇

2602.05441 2026-02-06 cs.RO cs.AI 73%

Benchmarking Affordance Generalization with BusyBox

通过BusyBox评估具身泛化能力

Dean Fortier, Timothy Adamson, Tess Hellebrekers, Teresa LaScala, Kofi Ennin, Michael Murray, Andrey Kolobov, Galen Mullins

机构 * Microsoft Research(微软研究院) Mississippi State University(密苏里州立大学)

专题命中 部署与泛化 :vision-language-action(abstract);VLA(abstract);分类 cs.RO、cs.AI

AI总结 本文提出BusyBox基准,用于评估VLA模型在具身泛化能力上的表现,通过模块化设计挑战模型对新物体的操作能力。

详情

展开后加载摘要…

URL PDF HTML 收藏