arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-11-14 至 2025-11-14 共收录 2 信号源:cs.CV, cs.AI, cs.LG

1. GUI与屏幕智能体 2 篇

2508.05294 2025-11-14 cs.RO cs.AI cs.LG 81%

Towards Embodied Agentic AI: Review and Classification of LLM- and VLM-Driven Robot Autonomy and Interaction

Sahar Salimpour, Lei Fu, Kajetan Rachwał, Pascal Bertrand, Kevin O'Sullivan, Robert Jakob, Farhad Keramat, Leonardo Militano, Giovanni Toffetti, Harry Edelman, Jorge Peña Queralta

机构 * Department of Computing, University of Turku(图尔库大学计算机系) Institute of Computer Science, Zurich University of Applied Sciences(应用科学大学计算机科学研究所) Centre for Artificial Ingelligence, Zurich University of Applied Sciences(应用科学大学人工智能中心) Agentic Systems Lab, Department of Management, Technology and Economics, ETH Zürich(苏黎世联邦理工学院管理、科技与经济系代理系统实验室) Faculty of Mathematics and Information Science, Warsaw University of Technology(华沙技术大学数学与信息科学学院)

专题命中 GUI与屏幕智能体 :VLM(title);vision-language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10615 2025-11-14 cs.CV cs.CL 57%

Towards Blind and Low-Vision Accessibility of Lightweight VLMs and Custom LLM-Evals

Shruti Singh Baghel, Yash Pratap Singh Rathore, Sushovan Jena, Anurag Pradhan, Amit Shukla, Arnav Bhavsar, Pawan Goyal

机构 * Indian Institute of Technology Mandi(印度理工学院曼迪分校) Vellore Institute of Technology(韦洛雷理工学院) Indian Institute of Technology Kharagpur(印度理工学院哈里科普分校)

专题命中 GUI与屏幕智能体 :vision-language model(abstract);分类 cs.CV

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏