arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

共收录 594 信号源:cs.RO, cs.CV, cs.AI, cs.LG

1. VLA模型 594 篇

2607.25703 2026-07-29 astro-ph.CO 新提交 67%

Simulation-based tension quantification of the cosmic dipole

基于模拟的宇宙偶极张力量化

Mali Land-Strykowski, Harry T. J. Bevins, Oliver T. Oayda, Geraint F. Lewis

专题命中 VLA模型 :VLA(abstract,abstract_cn)

AI总结 研究针对宇宙偶极张力超预期挑战,提出基于模拟的推理架构,用神经比率估计器测量张力,经验证可准确恢复真实值,应用于多数据集测得了不同张力,能在新观测时代实现稳健张力量化。

Comments 14 pages, 9 figures, accepted for publication in MNRAS

Journal ref Mon Not R Astron Soc (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22966 2026-07-28 astro-ph.GA 新提交 67%

A quasar hatching from a buried red phase at z = 3.7

一个在 z = 3.7 从深埋红相孵化出的类星体

Zheng Ma, Yongda Zhu, Zhiyuan Ji, Eiichi Egami, Marcia J. Rieke, Xiaohui Fan, Jianwei Lyu, George H. Rieke, Fengwu Sun, Yang Sun, George D. Becker, Andrew J. Bunker, Francesco D'Eugenio, Xiangyu Jin, Ignas Juodžbalis, Weizhe Liu, Roberto Maiolino, Pierluigi Rinaldi, Feige Wang, Christopher N. A. Willmer, Yunjing Wu, Jinyi Yang, Junyu Zhang, Peixin Zhu

专题命中 VLA模型 :VLA(abstract,abstract_cn)

AI总结 研究 z = 3.7 的红类星体“幼体”,通过多波长数据如 JWST 等观测,发现其活动核部分暴露,有宽发射线等特征,核附近有致密气体,多相气体被加速,连续谱呈红色且下降,为探索类星体过渡阶段提供独特机会。

Comments 26 pages, 5 main figures, and 10 Extended Data figures. Submitted, comments are welcome

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21841 2026-07-27 astro-ph.GA 新提交 67%

A Multiwavelength Inventory for the Local Group L-band Survey I: Atlas and Radial Profiles of Local Group Galaxies

本星系群L波段巡天I的多波长星表:本星系群星系的星图与径向分布

Cosima Eibensteiner, Adam K. Leroy, Jiayi Sun, Erik Roslowsky, Eric W. Koch, Nickolas Pingel, Chang-Goo Kim, Laura B. Chomiuk, Ryan Chown, Alberto D. Bolatto, Michael P. Busch, Julianne J. Dalcanton, Amanda A. Kepley, Christina W. Lindberg, Eve C. Ostriker, Jürgen Ott, Sumit K. Sarbadhicary, Adam Smercina, Snežana Stanimirović, Elizabeth Tarantino, Vicente Villanueva, Tobin M. Wainer, Fabian Walter, Thomas G. Williams

专题命中 VLA模型 :VLA(abstract,abstract_cn)

AI总结 该研究结合LGLBS的HI数据与多波段星图,分析六个本星系星系气体、恒星等径向分布,发现原子气体盘最延展,$\Sigma_{\rm HI}$分布特点及局部结构,给出气体消耗时间等,公开多波长数据为相关研究提供参考。

Comments Accepted for publication in ApJS; 21 pages, 12 figures, 6 tables Public avaible data under: https://dx.doi.org/10.11570/26.0020

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18440 2026-07-22 astro-ph.GA 新提交 67%

Vz-GAL Dusty Star-Forming Galaxies: Revisiting the CO-H2 Conversion Factor Tension

Vz-GAL尘埃星形成星系:重新审视CO-H2转换因子的张力

Prachi Prajapati, Axel Weiss, Dominik Riechers, Tom J. L. C. Bakx, Leindert A. Boogaard, Diana Ismail, Pierre Cox, Andrew J. Baker, Roberto Neri, Matthew Lehnert, Chentao Yang, Emilio Romano-Diaz, Hiddo S. B. Algera, Stefano Berta, Edoardo Borsato, Kirsty M. Butler, Asantha Cooray, Bethany Jones, Amelie Saintonge, Paul van der Werf

专题命中 VLA模型 :VLA(abstract,abstract_cn)

AI总结 研究高红移尘埃星形成星系中CO-H2转换因子的“张力”,通过对21个星系样本用多种方法推导分子气体质量,发现当前数据不支持αCO = 0.8,中间到接近银河系的值可行,校准需解析分子气体观测、物理建模和对尘埃性质的约束。

Comments Submitted to ApJ

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.17776 2026-07-21 astro-ph.GA 新提交 67%

Structural and dynamical properties of Tidal dwarf galaxies in the tails and bridge of the Guitar galaxy Arp 105

吉他星系Arp 105尾巴和桥中潮汐矮星系的结构与动力学性质

Jyoti Prakash, Kanak Saha

专题命中 VLA模型 :VLA(abstract,abstract_cn)

AI总结 研究阿贝尔1185星系团中Arp 105系统内潮汐矮星系及潮汐桥,通过多波长分析其恒星形成、金属度等性质,利用光谱能量分布建模得恒星质量,还估计动力学质量比,发现暗物质不足,后续观测或增进对潮汐矮星系形成的理解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14428 2026-07-17 astro-ph.GA 新提交 67%

Accurate proper motions of the protostellar system VLA1623-2417

原恒星系统VLA1623 - 2417的精确自行

Ricardo Hernández Garnica, Laurent Loinard, Carlos Carrasco-González, Jazmín Ordóñez-Toro, Johanan Ramírez-Arellano, María José Maureira, Isaac C. Radley, Eleonora Bianchi, Claire J. Chandler, Luis F. Rodríguez, Rosa M. Torres, Aina Palau

专题命中 VLA模型 :VLA(abstract,abstract_cn)

AI总结 对原恒星系统VLA1623 - 2417进行天体测量分析,结合多阵列观测数据得出各分量自行。Aa/Ab有轨道运动但难估质量,W与A、B投影距离减小,其运动不符弹出情景,推测W与A/B关系及轨道情况。

Comments Submitted to MNRAS

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04443 2026-07-09 cs.CV cs.AI cs.GR cs.LG 新提交 67%

Wan-Streamer v0.2: Higher Resolution, Same Latency

万流播器v0.2:更高分辨率,相同延迟

Lianghua Huang, Zhi-Fan Wu, Yupeng Shi, Wei Wang, Mengyang Feng, Junjie He, Chen-Wei Xie, Yu Liu, Jingren Zhou, Ang Wang, Bang Zhang, Baole Ai, Chen Liang, Cheng Yu, Chongyang Zhong, Jinwei Qi, Kai Zhu, Pandeng Li, Peng Zhang, Wenyuan Zhang, Xinhua Cheng, Yitong Huang, Yun Zheng, Yuxiang Bao, Yuzheng Wang, Zoubin Bi

机构 * Alibaba Group(阿里巴巴集团)

专题命中 VLA模型 :action model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 介绍万流播器v0.2,它保持v0.1建模公式,将交互输出流分辨率提高,在保持低延迟下支持场景中景智能体。核心方法是优化架构,主要贡献是提升分辨率同时保持低延迟,集中硬件于视觉生成。

Comments Website: https://wan-streamer.com/

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03592 2026-07-07 astro-ph.GA 新提交 67%

From Atomic Gas to Star Formation

从原子气体到恒星形成

Erik Rosolowsky, Eric Koch, Sambit Roychowdhury, Luca Cortese, Filippo M. Maccagni, Barbara Catinella, Nushkia Chamba, Timothy A. Davis, Cosima Eibensteiner, Eric Murphy, Mamta Pandey-Pommier, Amidou Sorgho

专题命中 VLA模型 :VLA(abstract,abstract_cn)

AI总结 研究星系中气体循环,此前因对占质量主导的冷原子相了解少而受限。利用SKA及其他设备高分辨率观测,可解决恒星形成星际介质演化的开放性问题,介绍了相关观测的近期结果。

Comments Published in Advancing Astrophysics with the SKA II (AASKAII), 2026 (arXiv:2606.20366). Report-no:AASKAII/Rosolowsky01

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03368 2026-07-07 astro-ph.CO 新提交 67%

The giant radio source 0917+75: Origin and properties

巨型射电源0917+75:起源与性质

G. Giovannini, N. Biava, M. Girardi, W. Boschin, R. Barrena, A. Bonafede, L. Feretti, C. Ferrari, F. Govoni, M. Iacobelli, M. Murgia, E. Orru', R. Pizzo, V. Vacca

专题命中 VLA模型 :VLA(abstract,abstract_cn)

AI总结 研究巨型射电源GRS0917 + 75,通过光学观测、低频射电成像及多频段射电数据分析,将其分类为特定类型星系,探讨其性质、起源及与环境的联系。

Comments 12 pages, 13 figures, accepted for publication in A&A

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02417 2026-07-03 cs.RO cs.CV cs.LG 新提交 67%

LIME: Learning Intent-aware Camera Motion from Egocentric Video

LIME: 从自我中心视频学习意图感知的相机运动

Boyang Sun, Jiajie Li, Yung-Hsu Yang, Chenyangguang Zhang, Tim Engelbracht, Sunghwan Hong, Cesar Cadena, Marc Pollefeys, Hermann Blum

机构 * ETH Zurich(苏黎世联邦理工学院) Microsoft(微软) University of Bonn(波恩大学)

专题命中 VLA模型 :vision-language-action(abstract);分类 cs.RO、cs.CV、cs.LG

AI总结 提出LIME模型,从自我中心视频中挖掘多意图相机运动监督,结合自回归观察增益输出与连续流匹配位姿头,实现语言条件相机运动生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26044 2026-06-25 astro-ph.GA astro-ph.SR 新提交 67%

Jets and Outflows in Young Stellar Objects with the SKAO

年轻恒星天体中的喷流与外流与SKAO

Giovanni Sabatini, Gemma Busquet, Carlos Carrasco-González, Adriana Rodríguez-Kamenetzky, Codella Claudio, Linda Podio, Antonio Martínez-Henares, Josep Miquel Girart, Marta De Simone, Luca Cacciapuoti, Guillem Anglada, Lukasz Tychoniec, Lisa Giani, Manoj Puravankara, Francesca Bacciotti, Rafael Bachiller, Eleonora Bianchi, Guillermo Blázquez-Calero, Tyler L. Bourke, Stefano Bovino, Paola Caselli, Francesco Cavallaro, Cecilia Ceccarelli, Elena Diaz-Marquez, Stefano Facchini, Antonio Garufi, Greta Guidi, Tomoya Hirota, John D. Ilee, Adriano Ingallinera, Izaskun Jiménez-Serra, Valerio Lattanzi, Chin-Fei Lee, Manuela Lippi, Alessandro Lupi, Liton Majumdar, Mayank Narang, Mayra Osorio, Marco Padovani, Jaime Pineda, Isaac Radley, Basmah Riaz, Luis Felipe Rodríguez, Alvaro Sánchez-Monge, Alberto Sanna, Silvia Spezzano, Leonardo Testi, Claudia Toci, Alessio Traficante, Himanshu Tyagi, Grazia Maria Umana

专题命中 VLA模型 :VLA(abstract,abstract_cn)

AI总结 本文综述了SKAO如何通过高分辨率厘米波观测、射电复合线和同步辐射,解决年轻恒星天体喷流/外流的加速、准直及化学影响等关键问题。

Comments Published in Advancing Astrophysics with the SKA II (AASKAII), 2026 (arXiv:2606.20366). Report-no:AASKAII/Sabatini01. Advancing Astrophysics with the SKA II (AASKAII) outlines the transformative scientific advances that will be enabled by the SKA telescopes

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20980 2026-06-23 cs.CV cs.AI cs.RO 新提交 67%

Robusto-2: Benchmarking Humans & VLMs for Autonomous Driving in Lima & New York City

Robusto-2: 在利马与纽约市对自动驾驶中人类与VLM的基准测试

Adrian Cespedes, Marcelo Chincha, Dunant Cusipuma, Victor Flores-Benites, David Ortega, Arturo Deza

机构 * Artificio Lima, Peru

专题命中 VLA模型 :action model(abstract);分类 cs.RO、cs.CV、cs.AI

AI总结 本研究通过视觉问答范式,比较人类驾驶员(来自利马和纽约)与视觉语言模型在利马和纽约的驾驶场景中的表现,发现人类与模型在回答事实、评级、反事实和推理四类问题时存在差异,但地理因素影响不显著。

Comments 11 pages main body. 42 pages total. Data publicly available online

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22845 2026-06-23 astro-ph.GA 新提交 67%

Observations of a Possible Transient Magnetically Arrested Accretion State in a Nearby Quasar: OQ208

邻近类星体OQ208中可能的瞬态磁捕获吸积态的观测

Brian Punsly, Cormac Reynolds, Paola Marziani, Gary J. Hill, Alexander B. Pushkarev, Carlo Stanghellini, Frank Schinzel, Christopher O'Dea, Gregory R. Zeimann, Andrew Biggs, Jian-Min Wang, Pu Du, Alberto Floris, Mauro D'Onofrio, Levi Malmstrom

专题命中 VLA模型 :VLA(abstract,abstract_cn)

AI总结 基于39年多波段数据,发现OQ208在1997-2001年间Hα发射线减弱与射电耀斑同步,支持瞬态磁捕获吸积态模型。

Comments To appear in ApJ

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09814 2026-06-18 astro-ph.SR astro-ph.GA 新提交 67%

ALMA measurements of mass loss and wind clumping in the massive stars of the Arches cluster

ALMA对Arches星团中大质量恒星的质量损失和风团块结构的测量

James P. Perry, Raman K. Prinja, Danielle M. Fenech, Francisco Najarro

专题命中 VLA模型 :VLA(summary_cn)

AI总结 利用ALMA和VLA数据,测量Arches星团中23颗大质量恒星的电离风质量损失率和团块结构,发现Wolf-Rayet星以热自由-自由辐射为主,而O型星存在非热同步辐射,并揭示风团块随半径减小。

Comments 16 pages, 5 figures, 7 tables, 2 appendices. Accepted for publication in MNRAS

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17210 2026-06-17 astro-ph.GA astro-ph.CO 新提交 67%

Testing masking effectiveness using multi-line image cubes based on COSMOS2020 for [CII] line intensity mapping at $z_{[CII]} > 3.5$

基于COSMOS2020的多线图像立方体测试掩膜有效性用于$z_{[CII]} > 3.5$的[CII]谱线强度映射

J. Clarke, C. Karoumpis, A. Dev, D. Riechers, T. Oak, Y. Okada, K. Narita, F. Bertoldi

专题命中 VLA模型 :VLA(abstract,abstract_cn)

AI总结 利用COSMOS2020星系目录构建CO和[CII]发射线的强度映射立方体,通过掩膜技术恢复高红移[CII]信号,并评估噪声影响。

Comments 23 pages, 22 figures, submitted to A&A

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13642 2026-06-12 gr-qc astro-ph.IM hep-ph 新提交 67%

Search for High-Frequency Gravitational Waves via Geomagnetic Conversion with Radio Telescopes

通过射电望远镜利用地磁转换搜索高频引力波

Hongliang Tian, Lei Wu, Xiaolong Yang, Qiang Yuan, Bin Zhu

专题命中 VLA模型 :VLA(abstract,abstract_cn)

AI总结 利用甚大阵列和阿塔卡马大型毫米波/亚毫米波阵列,通过逆Gertsenshtein效应搜索高频引力波,未发现信号,将特征应变上限提高至三个数量级。

Comments 6 pages, 3 figures + Supplemental Materials(8 pages, 5 figures)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19613 2026-08-21 cs.RO cs.CV 新提交 62%

What Matters for Latent Actions in Robot Learning

机器人学习中隐式动作的关键影响因素

Xizhou Bu, Qingda Hu, Lei Zhou, Lingfeng Zhang, Yingbo Tang, Zihao Liu, Xinyi Tao, Zhiqiang Ma, Qingqiu Huang, Chufeng Tang, Hongbo Wang, Jing Zhang, Jiayi Ma, Hangjun Ye, Wei Li, Xiaoshuai Hao

机构 * College of Intelligent Robotics and Advanced Manufacturing, Fudan University(复旦大学智能机器人与先进制造学院) Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Guangdong Provincial Key Laboratory of Computility Microelectronics, Faculty of Computility Microelectronics, Shenzhen University of Advanced Technology(深圳先进技术研究院计算机微科学与技术学院、广东省计算微电子重点实验室) School of Aeronautics and Astronautics, Sichuan University(四川大学航空航天学院) Suzhou Evans Intelligent Technology Co., Ltd.(苏州伊文斯智能科技有限公司) Morphi Intelligence Technology Co., Ltd.(墨飞智能科技有限公司) School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院) Xiaomi EV(小米汽车)

专题命中 VLA模型 :action model(abstract);分类 cs.RO、cs.CV

AI总结 本研究通过统一框架整合代表性隐式动作模型,系统探究41种设计选择,发现用隐式动作微调视觉语言模型主干可为下游机器人操纵策略学习提供更强初始化。

Comments Project page: https://carldegio.github.io/latent_action.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18555 2026-08-20 cs.LG cs.AI 新提交 62%

Performance Drift Detection in Machine Learning as a Service (MLaaS) for IoT Environments

物联网环境下机器学习即服务(MLaaS)的性能漂移检测

Deepak Kanneganti, Sajib Mistry, Sheik Mohammad Mostakim Fattah, Erik Elmroth, Aneesh Krishna, Monowar Bhuyan

机构 * Curtin University(科廷大学) Umeå University(于默奥大学)

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

AI总结 针对物联网环境下MLaaS客户端为黑盒的漂移检测难题,提出MPDD模型与APDDM机制,实验显示二者可显著提升检测准确率、降低漏检率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14655 2026-08-18 cs.LG cs.CV 新提交 62%

Diagnosing and Mitigating Perception-Decision Misalignment in Omni-LLMs via Modality Subspace Activation

通过模态子空间激活诊断并缓解全模态大语言模型(Omni-LLMs)中的感知-决策失配问题

Hongbo Jiang, Jie Li, Yunhang Shen, Tianyu Xie, Pingyang Dai

专题命中 VLA模型 :action model(abstract);分类 cs.CV、cs.LG

AI总结 针对Omni-LLMs存在的感知-决策失配问题,本文提出因果模态敏感性及对应诊断方法,构建CausalMSBench数据集,提出无需训练的MSA框架以恢复模型的因果模态敏感性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21290 2026-07-24 cs.LG cs.AI 新提交 62%

Multi-Task Learning for Heterogeneous Prediction from Video Game State with Transfer Learning

基于迁移学习的视频游戏状态异构预测多任务学习

Jonas Peché, Aliaksei Tsishurou, Alexander Zap, Günter Wallner

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

AI总结 研究视频游戏状态异构预测,通过适配多模态架构,采用跨任务联合训练共享模型,结合多种信息,经实验比较单多任务训练、评估策略及测试预训练等,还研究游戏内迁移,以降低成本并提升泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04648 2026-07-07 cs.LG cs.AI 新提交 62%

Machine Learning for Depression Screening and Intervention: an Original Circadian Rhythm Score-based Methodology

用于抑郁症筛查和干预的机器学习:一种基于昼夜节律评分的原创方法

Bin Wang, Shuo Lian, Yuanyuan Hou, Dexian Wang, Peilan He, Feng Hong, Yanwei Yu, Tianrui Li

机构 * Ocean University of China(中国海洋大学) The Affiliated Hospital of Qingdao University(青岛大学附属医院) Chengdu University of Traditional Chinese Medicine(成都中医药大学) Southwest Jiaotong University(西南交通大学)

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

AI总结 研究针对大规模行为数据筛查抑郁症的挑战,提出昼夜节律评分(CRS),构建可解释筛查框架及干预推理方法,通过实验验证其有效性,为抑郁症筛查和干预提供框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25434 2026-06-25 cs.LG cs.AI 新提交 62%

Interpretable Concept-Guided Polynomial Tabular Kolmogorov-Arnold Network for EEG-Based Mild Cognitive Impairment Detection

可解释概念引导的多项式表格Kolmogorov-Arnold网络用于基于EEG的轻度认知障碍检测

Yosef Bernardus Wirian, Qiang Cheng

机构 * Computer Science Department, University of Kentucky(肯塔基大学计算机科学系) Institute for Biomedical Informatics, University of Kentucky(肯塔基大学生物医学信息学研究所)

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

AI总结 提出CPTabKAN模型,通过概念编码、二阶多项式交互和傅里叶参数化TabKAN分类器,在睡眠EEG数据上实现MCI检测,F1达0.9038。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25753 2026-06-25 cs.LG cs.AI math.OC physics.comp-ph physics.optics 新提交 62%

Gradient-based inverse lithography for EUV masks via the waveguide method and a physics-informed neural operator

基于梯度方法的EUV掩模逆光刻:波导方法与物理信息神经算子

Vasiliy A. Es'kin, Egor V. Ivanov

机构 * University of Nizhny Novgorod(下诺夫哥罗德大学)

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

AI总结 提出一种基于梯度方法的极紫外掩模逆光刻技术,利用可微波导方法和波导神经算子作为端到端物理引擎,通过自动微分恢复掩模吸收体介电常数,实验验证了在多种材料下实现所需晶圆场。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21623 2026-06-23 cs.CV cs.AI 新提交 62%

A DVDrive Approach for doScenes Instructed Driving Challenge

一种面向 doScenes 指令驾驶挑战的 DVDrive 方法

Zijian Fu, Xiangyang Chu, Mengshi Qi, Huadong Ma, Guanghao Zhang, Wei Li

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Xiaomi EV(小米汽车)

专题命中 VLA模型 :vision-language-action(abstract);分类 cs.CV、cs.AI

AI总结 提出基于 OmniDrive 的指令条件轨迹预测方法,引入 DVPE 分视角感知模块减少跨视角干扰,提升多视角视觉定位与语言指令对齐能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.19408 2026-06-19 cs.LG cs.RO 新提交 62%

FlexLAM: Resolving the Bottleneck Trade-off in Latent Action Learning

FlexLAM: 解决潜在动作学习中的瓶颈权衡

Takanori Yoshimoto, Yang Hu, Naruya Kondo, Tatsuya Matsushima

机构 * University of Tsukuba(筑波大学) The University of Tokyo(东京大学)

专题命中 VLA模型 :action model(abstract);分类 cs.RO、cs.LG

AI总结 针对潜在动作模型中固定容量瓶颈导致的权衡问题,提出FlexLAM,通过嵌套dropout实现变长潜在动作,在不增加架构或损失的情况下,在稀缺标签和低回报任务中优于固定容量模型,并支持推理时调整令牌预算。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17480 2026-06-17 cs.CV cs.RO 新提交 62%

GeneralVLA-2: Geometry-Aware Reconstruction and Governed Memory for Robot Planning

GeneralVLA-2: 几何感知重建与受控记忆用于机器人规划

Haoyu Wang, Guoqing Ma, Zeyu Zhang, Yandong Guo, Boxin Shi, Hao Tang

机构 * School of Computer Science, Peking University(北京大学计算机科学学院) CASIA(中国科学院自动化研究所) AI 2 Robotics

专题命中 VLA模型 :vision-language-action(abstract);分类 cs.RO、cs.CV

AI总结 针对机器人规划中3D物体重建幻觉和记忆质量不可控的问题,提出GeoFuse-MV3D几何先验引导重建分支和受控长期记忆系统,在GSO-30和Terminal-Bench等基准上显著提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14777 2026-06-16 cs.CV cs.AI 新提交 62%

JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence

JoyAI-VL-Interaction: 实时视觉-语言交互智能

Dingyu Yao, Junhao Zhou, Chenxu Yang, Chuanyu Qin, Haowen Hou, Zheming Liang, Congcong Wang, Yuhang Cao, Shenglong Ye, Shuai Xie, Shuhuan Gu, Haoyang Huang, Qingyi Si, Nan Duan, Jiaqi Wang

机构 * JD.com(京东)

专题命中 VLA模型 :action model(abstract);分类 cs.CV、cs.AI

AI总结 提出一种持续观察、自主决定是否回应的视觉-语言交互模型,并开源8B规模模型及完整部署系统,在六个真实场景中优于现有方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18184 2026-08-20 cs.CV 新提交 57%

Human-Centric Intelligence in the Era of Foundation Models: A Survey

基础模型时代的以人为中心的智能:一项综述

Yang Chen, Tianqi Wang, Xiaorui Jiang, Yilei Man, Yihua Shao, Mengyuan Liu, Zhi Chen, Xiaofeng Cao, Qibin Zhao, Chi Harold Liu, Albert Y. Zomaya, Nicu Sebe, Jingren Zhou, Dacheng Tao, Song Guo, Jingcai Guo

机构 * The Hong Kong Polytechnic University(香港理工大学) Peking University(北京大学) University of Southern Queensland(南昆士兰大学) Tongji University(同济大学) RIKEN Center for Advanced Intelligence Project(理化学研究所先进智能项目中心) Beijing Institute of Technology(北京理工大学) The University of Sydney(悉尼大学) University of Trento(特伦托大学) Alibaba Group(阿里巴巴集团) Nanyang Technological University(南洋理工大学) Hong Kong University of Science and Technology(香港科技大学)

专题命中 VLA模型 :action model(abstract);分类 cs.CV

AI总结 本文针对基础模型时代以人为中心的智能进展碎片化问题,提出全谱人类上下文分类法,梳理其方法基础、相关方法与资源,探讨挑战方向,为该领域推进提供参考。

Comments GitHub Repo: https://github.com/cseeyangchen/Human-Centric-AI; Project Page: https://cseeyangchen.github.io/Human-Centric-AI/homepage/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18026 2026-08-19 cs.LG cs.CE 新提交 57%

TabNSM: Neural Sparse Mixer for Tabular Regression

TabNSM:面向表格回归的神经稀疏混合器

Ali Eslamian, Qiang Cheng

机构 * University of Kentucky(肯塔基大学) Institute for Biomedical Informatics(生物医学信息学研究所)

专题命中 VLA模型 :action model(abstract);分类 cs.LG

AI总结 该研究提出TabNSM框架,通过自适应稀疏交互模块、多阶段回归头等组件,在9个真实回归基准上实现高维表格回归的优性能与可扩展性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17496 2026-08-19 cs.RO 新提交 57%

Calibrated Predictive Safety for Heterogeneous Robots: An Action-Conditioned JEPA Framework with Model-Based Safety Shields

异构机器人的校准预测安全性:一种带基于模型安全盾的动作条件联合嵌入预测架构(JEPA)框架

Kaiming Zhong, Tianhua Liu, Yue Wang

机构 * Guangdong Bifang Intelligent Control Technology Co., Ltd.(广东毕方智能控制技术有限公司)

专题命中 VLA模型 :vision-language-action(abstract);分类 cs.RO

AI总结 本研究提出带基于模型安全盾的动作条件JEPA框架,用于异构机器人,可预测候选动作的任务进展与物理风险,在仿真中提升了成功率并降低碰撞漏报率。

Comments 17 pages, 9 figures. Simulation-only empirical results on LIBERO-Long (no real-robot experiments). Source, figure-generation scripts and reproducibility checklist included. Level-3 offline reranking significance test not executed; see Sec. 7 (Scope and honesty statement) for detailed disclosure

详情

展开后加载摘要…

URL PDF HTML 收藏