arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-26 至 2025-09-26 共收录 6 信号源:cs.CV, cs.AI, cs.LG

1. GUI与屏幕智能体 6 篇

2509.21126 2025-09-26 cs.LG cs.AI 81%

Teaching RL Agents to Act Better: VLM as Action Advisor for Online Reinforcement Learning

Xiefeng Wu, Jing Zhao, Shu Zhang, Mingyu Hu

机构 * Wuhan University(武汉大学)

专题命中 GUI与屏幕智能体 :VLM(title);vision-language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21189 2025-09-26 cs.RO cs.AI cs.CV 73%

Human-like Navigation in a World Built for Humans

Bhargav Chandaka, Gloria X. Wang, Haozhe Chen, Henry Che, Albert J. Zhai, Shenlong Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 GUI与屏幕智能体 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments CoRL 2025. Project website: https://reasonnav.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20841 2025-09-26 cs.RO cs.AI cs.LG 62%

ImaginationPolicy: Towards Generalizable, Precise and Reliable End-to-End Policy for Robotic Manipulation

Dekun Lu, Wei Gao, Kui Jia

机构 * Dekun Lu ∗ , Wei Gao ∗ and Kui Jia ∗(作者)

专题命中 GUI与屏幕智能体 :VLM(abstract);分类 cs.AI、cs.LG

Comments First two authors contribute equally. Project page: https://sites.google.com/view/imaginationpolicy

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21259 2025-09-26 cs.NI cs.AI 57%

Semantic Edge-Cloud Communication for Real-Time Urban Traffic Surveillance with ViT and LLMs over Mobile Networks

Murat Arda Onsu, Poonam Lohan, Burak Kantarci, Aisha Syed, Matthew Andrews, Sean Kennedy

专题命中 GUI与屏幕智能体 :multimodal large language model(abstract);分类 cs.AI

Comments 17 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21006 2025-09-26 cs.RO cs.AI 57%

AnywhereVLA: Language-Conditioned Exploration and Mobile Manipulation

Konstantin Gubernatorov, Artem Voronov, Roman Voronov, Sergei Pasynkov, Stepan Perminov, Ziang Guo, Dzmitry Tsetserukou

机构 * Intelligent Space Robotics Laboratory, Center for Digital Engineering, Skolkovo Institute of Science and Technology(智能空间机器人实验室,数字工程中心,斯克尔科夫科学与技术研究所)

专题命中 GUI与屏幕智能体 :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17084 2025-09-26 cs.CV 57%

MoCLIP-Lite: Efficient Video Recognition by Fusing CLIP with Motion Vectors

Binhua Huang, Ni Wang, Arjun Pakrashi, Soumyabrata Dev

机构 * The ADAPT SFI Research Centre(ADAPT SFI研究机构) School of Computer Science, University College Dublin(大学学院计算机科学系) Amazon Development Center Germany GmbH(亚马逊德国开发中心) Beijing-Dublin International College(北京-都柏林国际学院)

专题命中 GUI与屏幕智能体 :vision-language model(abstract);分类 cs.CV

Comments 6 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏