arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-11-24 至 2025-11-24 共收录 3 信号源:cs.CV, cs.AI, cs.LG

1. GUI与屏幕智能体 3 篇

2511.14004 2025-11-24 cs.RO 67%

Searching in Space and Time: Unified Memory-Action Loops for Open-World Object Retrieval

在空间和时间中搜索:统一的记忆-动作循环用于开放世界物体检索

Taijing Chen, Sateesh Kumar, Junhong Xu, Georgios Pavlakos, Joydeep Biswas, Roberto Martín-Martín

机构 * Department of Computer Science, The University of Texas at Austin(计算机科学系,德克萨斯大学奥斯汀分校)

专题命中 GUI与屏幕智能体 :vision-language model(abstract);grounding(abstract)

AI总结 STAR框架通过统一记忆查询和具身动作,提升开放世界中时空物体检索的效率与准确性。

Comments https://amrl.cs.utexas.edu/STAR/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06182 2025-11-24 cs.RO 67%

OpenVLN: Open-world Aerial Vision-Language Navigation

OpenVLN: 开放世界空中视觉语言导航

Peican Lin, Gan Sun, Chenxi Liu, Fazeng Li, Weihong Ren, Yang Cong

机构 * School of Automation Science and Engineering, South China University of Technology(华南理工大学自动化科学与工程学院) State Key Laboratory of Robotics, Shenyang Institute of Automation, Chinese Academy of Sciences(中国科学院沈阳自动化研究所机器人重点实验室) State Key Laboratory of Robotics and System, School of Mechanical Engineering and Automation, Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)机械工程与自动化学院机器人与系统重点实验室)

专题命中 GUI与屏幕智能体 :vision-language model(abstract);VLM(abstract)

AI总结 OpenVLN通过强化学习和长视距规划器提升无人机在复杂空中环境中的长视距导航能力。

Comments Content: 8 pages 4 figures, conference paper under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17401 2025-11-24 cs.RO cs.HC 50%

Feasibility of Embodied Dynamics Based Bayesian Learning for Continuous Pursuit Motion Control of Assistive Mobile Robots in the Built Environment

基于具身动力学的贝叶斯学习在辅助移动机器人连续追击运动控制中的可行性

Xiaoshan Zhou, Carol C. Menassa, Vineet R. Kamat

专题命中 GUI与屏幕智能体 :grounding(abstract)

AI总结 本文提出基于具身动力学的贝叶斯学习方法,有效提升辅助移动机器人在复杂环境中的连续追击运动控制性能。

Comments 37 pages, 9 figures, and 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏