arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-02-04 至 2026-02-04 共收录 3 信号源:cs.CV, cs.AI, cs.LG

1. GUI与屏幕智能体 3 篇

2602.03750 2026-02-04 cs.CV cs.AI 81%

Zero-shot large vision-language model prompting for automated bone identification in paleoradiology x-ray archives

零样本大视觉语言模型提示法用于古放射学x光档案中骨骼自动识别

Owen Dong, Lily Gao, Manish Kota, Bennett A. Landmana, Jelena Bekvalac, Gaynor Western, Katherine D. Van Schaik

专题命中 GUI与屏幕智能体 :vision-language model(title);vision language model(abstract);分类 cs.CV、cs.AI

AI总结 利用大视觉语言模型实现古放射学x光图像中骨骼自动识别,提升内容导航效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02548 2026-02-04 cs.LG cs.AI cs.CV cs.MA 67%

ToolTok: Tool Tokenization for Efficient and Generalizable GUI Agents

ToolTok: 为高效且通用的GUI代理的工具标记化

Xiaoce Wang, Guibin Zhang, Junzhe Li, Jinzhe Tu, Chun Li, Ming Li

机构 * Department of Computer Science, Tsinghua University, Beijing, China(清华大学计算机科学系) Guangming Laboratory, Shenzhen, China(光明实验室) National University of Singapore, Singapore(新加坡国立大学) Peking University, Beijing, China(北京大学) MSU-BIT-SMBU Joint Research Center of Applied Mathematics, Shenzhen MSU–BIT University(MSU-BIT-SMBU应用数学联合研究中心)

专题命中 GUI与屏幕智能体 :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 ToolTok通过多步路径寻找范式,利用语义锚定机制和易到难课程,实现高效且通用的GUI代理,以较少数据获得优异性能。

Comments 8 pages main paper, 18 pages total, 8 figures, 5 tables, code at https://github.com/ZephinueCode/ToolTok

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21602 2026-02-04 cs.RO 50%

AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation

AIR-VLA:面向空中操作的视觉-语言-动作系统

Jianli Sun, Bin Tian, Qiyao Zhang, Chengxiang Li, Zihan Song, Zhiyong Cui, Yisheng Lv, Yonglin Tian

机构 * The Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Automation, Beijing Institute of Technology(北京理工大学自动化学院) School of Information and Intelligent Engineering, University of Sanya(三亚大学信息与智能工程学院) School of Mechanical and Vehicle Engineering, Hunan University(湖南大学机械与车辆工程学院) State Key Lab of Intelligent Transportation Systems, School of Transportation Science and Engineering, Beihang University(北京航空航天大学交通科学与工程学院)

专题命中 GUI与屏幕智能体 :VLM(abstract)

AI总结 AIR-VLA提出首个针对空中操作的视觉-语言-动作系统,通过构建仿真环境和多模态数据集,评估主流模型并揭示其在无人机移动、机械臂控制和高层规划中的能力和限制。

详情

展开后加载摘要…

URL PDF HTML 收藏