arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-08 至 2025-09-08 共收录 3 信号源:cs.CV, cs.AI, cs.LG

1. GUI与屏幕智能体 3 篇

2506.13205 2025-09-08 cs.CR cs.AI 83%

Poison Once, Control Anywhere: Clean-Text Visual Backdoors in VLM-based Mobile Agents

Xuan Wang, Siyuan Liang, Zhe Liu, Yi Yu, Aishan Liu, Yuliang Lu, Xitong Gao, Ee-Chien Chang

专题命中 GUI与屏幕智能体 :VLM(title,abstract);vision-language model(abstract);分类 cs.AI

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00388 2025-09-08 cs.RO 78%

Find Everything: A General Vision Language Model Approach to Multi-Object Search

Daniel Choi, Angus Fung, Haitong Wang, Aaron Hao Tan

专题命中 GUI与屏幕智能体 :vision language model(title,abstract)

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02544 2025-09-08 cs.CL 67%

Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions

Xinbei Ma, Yiting Wang, Yao Yao, Tongxin Yuan, Aston Zhang, Zhuosheng Zhang, Hai Zhao

机构 * School of Computer Science(计算机学院) Key Laboratory of Shanghai Education Commission for Intelligent Interaction and Cognitive Engineering(智能交互与认知工程重点实验室) Shanghai Jiao Tong University(上海交通大学) Shanghai Key Laboratory of Trusted Data Circulation and Governance in Web3(Web3可信数据流通与治理上海市重点实验室) GenAI, Meta(Meta GenAI)

专题命中 GUI与屏幕智能体 :multimodal large language model(abstract);MLLM(abstract)

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏