arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

共收录 2982 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. GUI与网页智能体 2982 篇

2512.19107 2025-12-23 cs.AI 57%

FC-MIR: A Mobile Screen Awareness Framework for Intent-Aware Recommendation based on Frame-Compressed Multimodal Trajectory Reasoning

FC-MIR:一种基于帧压缩多模态轨迹推理的移动屏幕感知框架,用于意图感知推荐

Zhe Yang, Xiaoshuang Sheng, Zhengnan Zhang, Jidong Wu, Zexing Wang, Xin He, Shenghua Xu, Guanjing Xiong

机构 * vivo AI Lab(vivo人工智能实验室) Zhejiang University(浙江大学)

专题命中 GUI与网页智能体 :agent(abstract);分类 cs.AI

AI总结 FC-MIR框架通过帧压缩多模态轨迹推理,提升移动设备上意图感知推荐的效率与实用性,支持轻量级部署并探索生成操作与建议。

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.00162 2025-12-23 cs.LG cs.NA math.NA 57%

Modeling Large-Scale Walking and Cycling Networks: A Machine Learning Approach Using Mobile Phone and Crowdsourced Data

对大规模步行和骑行网络建模:一种利用移动电话和众包数据的机器学习方法

Meead Saberi, Tanapon Lilasathapornkit

专题命中 GUI与网页智能体 :planning(abstract);分类 cs.LG

AI总结 本研究采用机器学习方法,利用移动电话和众包数据对澳大利亚新南威尔士州大规模步行和骑行网络进行建模,以提高主动交通规划的准确性。

Comments 22 pages, 8 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14014 2025-12-17 cs.AI 57%

MobileWorldBench: Towards Semantic World Modeling For Mobile Agents

MobileWorldBench: 向移动智能体的语义世界建模迈进

Shufan Li, Konstantinos Kallidromitis, Akash Gokul, Yusuke Kato, Kazuki Kozuka, Aditya Grover

机构 * UCLA(加州大学洛杉矶分校) Panasonic AI Research(松下人工智能研究) Salesforce AI Research(Salesforce人工智能研究)

专题命中 GUI与网页智能体 :planning(abstract);分类 cs.AI

AI总结 MobileWorldBench通过语义世界模型提升移动智能体任务成功率

Comments 21 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11147 2025-12-15 cs.CR cs.AI 57%

MiniScope: A Least Privilege Framework for Authorizing Tool Calling Agents

MiniScope: 一种基于最小特权原则的工具调用代理授权框架

Jinhao Zhu, Kevin Tseng, Gil Vernik, Xiao Huang, Shishir G. Patil, Vivian Fang, Raluca Ada Popa

机构 * University of California, Berkeley(加州大学伯克利分校) IBM Research(IBM研究院)

专题命中 GUI与网页智能体 :agentic(abstract);分类 cs.AI

AI总结 MiniScope通过重建权限层次结构和移动式权限模型,实现最小特权原则,以平衡安全性和易用性,同时在权限最小化和成本控制方面优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10934 2025-12-12 cs.RO cs.LG 57%

Curriculum-Based Reinforcement Learning for Autonomous UAV Navigation in Unknown Curved Tubular Conduit

基于课程的学习强化学习用于自主无人机在未知弯曲管状通道中的导航

Zamirddine Mari, Jérôme Pasquet, Julien Seinturier

机构 * DGA Techniques Navales - Direction Générale de l’Armement(法国国防总局海军技术部) LIRMM, TETIS, CNRS, Université de Montpellier Paul-Valéry(蒙彼利埃大学保罗-瓦莱里耶大学) LIS, CNRS, Université de Toulon(土伦大学)

专题命中 GUI与网页智能体 :agent(abstract);分类 cs.LG

AI总结 本文提出基于课程学习的强化学习方法,使无人机在未知弯曲管状环境中实现自主导航,通过结合激光雷达和视觉信息,克服部分可观测性挑战,提升导航稳定性与泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09920 2025-12-11 cs.RO cs.AI cs.CV 57%

LISN: Language-Instructed Social Navigation with VLM-based Controller Modulating

LISN: 基于VLM控制器的指令引导社交导航

Junting Chen, Yunchuan Li, Panfeng Jiang, Jiacheng Du, Zixuan Chen, Chenrui Tie, Jiajun Deng, Lin Shao

机构 * RoboScience Co.(RoboScience公司) ShanghaiTech University(上海科技大学) Nanjing University(南京大学) University of Science and Technology of China(中国科学技术大学)

专题命中 GUI与网页智能体 :agent(abstract);分类 cs.AI

AI总结 LISN通过VLM控制器实现指令引导的社交导航,提出首个标准化基准并提升动态避障能力。

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06912 2025-12-10 cs.RO cs.LG 57%

Khalasi: Energy-Efficient Navigation for Surface Vehicles in Vortical Flow Fields

Khalasi:在涡流流场中表面车辆的节能导航

Rushiraj Gadhvi, Sandeep Manjanna

机构 * Plaksha University(普拉克斯大学)

专题命中 GUI与网页智能体 :planning(abstract);分类 cs.LG

AI总结 本文提出了一种基于强化学习的端到端方法,用于在涡流流场中实现表面车辆的节能导航,通过局部速度测量学习流感知的导航策略,显著提升能源效率和泛化能力。

Comments Under Review for International Conference on Robotics and Automation (ICRA 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06490 2025-12-09 cs.LG 57%

Optimizing LLMs Using Quantization for Mobile Execution

用量化优化LLMs在移动执行中的应用

Agatsya Yadav, Renta Chintala Bhargavi

机构 * School of Computer Science(计算机科学学院) Engineering(工程学院) Vellore Institute of Technology(韦洛雷理工学院) Chennai, India(印度钦奈)

专题命中 GUI与网页智能体 :workflow(abstract);分类 cs.LG

AI总结 本文通过4位后训练量化和GGUF格式优化,实现Llama 3.2 3B模型在移动设备上的高效部署。

Comments 11 pages, 1 equation, 2 tables. Author Accepted Manuscript (AAM) of a paper published in Springer LNNS, ICT4SD 2025. DOI: 10.1007/978-3-032-06697-8_33

Journal ref Fong, S., Dey, N., Joshi, A. (eds) ICT Analysis and Applications. ICT4SD 2025. Lecture Notes in Networks and Systems, vol 1654. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04105 2025-12-05 cs.CY cs.AI cs.HC 57%

LegalWebAgent: Empowering Access to Justice via LLM-Based Web Agents

LegalWebAgent:通过基于大语言模型的网络代理赋能正义获取

Jinzhe Tan, Karim Benyekhlef

机构 * Cyberjustice Laboratory, Faculty of Law, University of Montreal, Canada(法律正义实验室,法学院,蒙特利尔大学,加拿大)

专题命中 GUI与网页智能体 :agent(abstract);分类 cs.AI

AI总结 LegalWebAgent通过多模态大语言模型实现网络代理,解决普通公民在法律问题中的获取障碍,实现从查询到行动的全流程自动化,实验结果显示其在复杂场景中的高自主性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03429 2025-12-04 cs.RO cs.AI 57%

World Models for Autonomous Navigation of Terrestrial Robots from LIDAR Observations

基于LIDAR观测的地面机器人自主导航的World Models

Raul Steinmetz, Fabio Demo Rosa, Victor Augusto Kich, Jair Augusto Bottega, Ricardo Bedin Grando, Daniel Fernando Tello Gamarra

机构 * Universidade Federal de Santa Maria, Brazil(巴西联邦大学圣玛利亚分校) University of Tsukuba, Japan(日本筑波大学) Universidade Federal de Rio Grande(巴西联邦大学里约格兰德分校) Universidad Tecnológica del Uruguay, Uruguay(乌拉圭技术大学)

专题命中 GUI与网页智能体 :agent(abstract);分类 cs.AI

AI总结 本文提出基于DreamerV3算法的模型驱动强化学习框架,通过整合MLP-VAE实现高维LIDAR数据的高效编码与潜在表示学习,提升地面机器人自主导航的效率与鲁棒性。

Comments Accepted for publication in the Journal of Intelligent and Fuzzy Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00049 2025-12-02 cs.RO cs.AI 57%

Socially aware navigation for mobile robots: a survey on deep reinforcement learning approaches

面向移动机器人的社会感知导航:深度强化学习方法的综述

Ibrahim Khalil Kabir, Muhammad Faizan Mysorewala

专题命中 GUI与网页智能体 :agent(abstract);分类 cs.AI

AI总结 本文综述了深度强化学习在社会感知导航中的应用,分析了关键技术和挑战,提出未来需结合多种方法并建立平衡的评估标准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17532 2025-11-26 cs.NI cs.AI 57%

Denoising Refinement Diffusion Models for Simultaneous Generation of Multi-scale Mobile Network Traffic

去噪细化扩散模型用于多尺度移动网络流量的同时生成

Xiaoqian Qi, Haoye Chai, Sichang Liu, Lei Yue, Raoyuan Pan, Yue Wang, Yong Li

机构 * Department of Electronic Engineering, BNRist, Tsinghua University(电子工程系,北京理工大学,清华大学) State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications(网络与交换技术国家重点实验室,北京邮电大学) China Mobile Communications Group Guangxi Co., Ltd.(中国移动通信集团广西有限公司)

专题命中 GUI与网页智能体 :planning(abstract);分类 cs.AI

AI总结 ZoomDiff通过多阶段去噪细化扩散模型实现多尺度移动网络流量的同时生成,提升了流量预测的准确性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19477 2025-11-26 cs.SE 57%

Building Browser Agents: Architecture, Security, and Practical Solutions

构建浏览器代理:架构、安全与实用解决方案

Aram Vardanyan

专题命中 GUI与网页智能体 :agent(abstract);分类 cs.SE

AI总结 本文提出通过专用工具和编程约束构建安全浏览器代理,实现85%的成功率。

Comments 30 pages, 22 figures. Production architecture and benchmark evaluation of browser agents

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12937 2025-11-26 cs.AI cs.CV 57%

Yanyun-3: Enabling Cross-Platform Strategy Game Operation with Vision-Language Models

Yanyun-3: 通过视觉-语言模型实现跨平台战略游戏操作

Guoyan Wang, Yanyan Huang, Chunlin Chen, Lifeng Wang, Yuxiang Sun

专题命中 GUI与网页智能体 :agent(abstract);分类 cs.AI

AI总结 Yanyun-3通过整合视觉-语言模型和结构化多模态数据组织,实现了跨平台战略游戏操作的自动化,提升了任务执行效率和泛化能力。

Comments 32 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17225 2025-11-24 cs.RO cs.AI cs.CV 57%

TP-MDDN: Task-Preferenced Multi-Demand-Driven Navigation with Autonomous Decision-Making

TP-MDDN:任务优先的多需求驱动导航与自主决策

Shanshan Li, Da Huang, Yu He, Yanwei Fu, Yu-Gang Jiang, Xiangyang Xue

机构 * Fudan University(复旦大学) Shanghai Jiao Tong University(上海交通大学) Shanghai Innovation Institution(上海创新机构)

专题命中 GUI与网页智能体 :planning(abstract);分类 cs.AI

AI总结 TP-MDDN通过引入任务优先的多需求驱动导航与自主决策系统,提升复杂任务下的导航性能与环境理解能力。

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15567 2025-11-20 cs.CV cs.CL cs.HC 57%

Computer-Use Agents as Judges for Generative User Interface

计算机使用代理作为生成用户界面的法官

Kevin Qinghong Lin, Siyuan Hu, Linjie Li, Zhengyuan Yang, Lijuan Wang, Philip Torr, Mike Zheng Shou

机构 * University of Oxford(牛津大学) Show Lab, National University of Singapore(新加坡国立大学Show实验室) Microsoft(微软公司)

专题命中 GUI与网页智能体 :agent(abstract);分类 cs.CL

AI总结 本研究提出Coder-CUA协作框架,通过代理作为法官与编码模型协作,提升自动GUI设计的效率和可靠性。

Comments Project: https://showlab.github.io/AUI Github: https://github.com/showlab/AUI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23452 2025-11-20 cs.IR cs.SE 57%

What About Emotions? Guiding Fine-Grained Emotion Extraction from Mobile App Reviews

Quim Motger, Marc Oriol, Max Tiessler, Xavier Franch, Jordi Marco

专题命中 GUI与网页智能体 :planning(abstract);分类 cs.SE

Comments Accepted at the 33rd IEEE International Requirements Engineering 2025 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11845 2025-11-18 cs.RO cs.AI cs.AR 57%

Autonomous Underwater Cognitive System for Adaptive Navigation: A SLAM-Integrated Cognitive Architecture

K. A. I. N Jayarathne, R. M. N. M. Rathnayaka, D. P. S. S. Peiris

机构 * Department of Computational Mathematics University of Moratuwa(计算数学系大学莫图瓦)

专题命中 GUI与网页智能体 :planning(abstract);分类 cs.AI

Comments 6 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11323 2025-11-17 cs.AI 57%

RLSLM: A Hybrid Reinforcement Learning Framework Aligning Rule-Based Social Locomotion Model with Human Social Norms

Yitian Kou, Yihe Gu, Chen Zhou, DanDan Zhu, Shuguang Kuai

专题命中 GUI与网页智能体 :agent(abstract);分类 cs.AI

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09309 2025-11-13 cs.HC cs.AI 57%

TaskSense: Cognitive Chain Modeling and Difficulty Estimation for GUI Tasks

Yiwen Yin, Zhian Hu, Xiaoxi Xu, Chun Yu, Xintong Wu, Wenyu Fan, Yuanchun Shi

机构 * Tsinghua University(清华大学) University of Washington(华盛顿大学) Cornell University(康奈尔大学) University of Sydney(悉尼大学)

专题命中 GUI与网页智能体 :agent(abstract);分类 cs.AI

Comments 22 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09157 2025-11-13 cs.AI 57%

ProBench: Benchmarking GUI Agents with Accurate Process Information

Leyang Yang, Ziwei Wang, Xiaoxuan Tang, Sheng Zhou, Dajun Chen, Wei Jiang, Yong Li

专题命中 GUI与网页智能体 :agent(abstract);分类 cs.AI

Comments Paper accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08271 2025-11-13 cs.CV cs.SE 57%

SWAN -- Enabling Fast and Mobile Histopathology Image Annotation through Swipeable Interfaces

Sweta Banerjee, Timo Gosch, Sara Hester, Viktoria Weiss, Thomas Conrad, Taryn A. Donovan, Nils Porsche, Jonas Ammeling, Christoph Stroblberger, Robert Klopfleisch, Christopher Kaltenecker, Christof A. Bertram, Katharina Breininger, Marc Aubreville

机构 * Flensburg University of Applied Sciences(弗劳恩霍夫应用科技大学) University of Veterinary Medicine(兽医大学) Schwarzman Animal Medical Center(施瓦茨曼动物医疗中心) Medical University of Vienna(维也纳医学大学) Technische Hochschule Ingolstadt(英戈尔施塔特技术大学)

专题命中 GUI与网页智能体 :workflow(abstract);分类 cs.SE

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04307 2025-11-11 cs.AI 57%

GUI-360$^\circ$: A Comprehensive Dataset and Benchmark for Computer-Using Agents

Jian Mu, Chaoyun Zhang, Chiming Ni, Lu Wang, Bo Qiao, Kartik Mathur, Qianhui Wu, Yuhang Xie, Xiaojun Ma, Mengyu Zhou, Si Qin, Liqun Li, Yu Kang, Minghua Ma, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang

机构 * Nanjing University(南京大学) Microsoft(微软) ZJU-UIUC(浙大-UIUC) Peking University(北京大学)

专题命中 GUI与网页智能体 :agent(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15922 2025-11-11 cs.LG cs.RO 57%

The Dark Side of Rich Rewards: Understanding and Mitigating Noise in VLM Rewards

Sukai Huang, Shu-Wei Liu, Nir Lipovetzky, Trevor Cohn

机构 * Google DeepMind(谷歌DeepMind)

专题命中 GUI与网页智能体 :agent(abstract);分类 cs.LG

Comments accepted by PRL Workshop Series @ ICAPS 2025. 11 main body pages, 21 appendix pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22940 2025-11-05 cs.LG 57%

Generating Auxiliary Tasks with Reinforcement Learning

Judah Goldfeder, Matthew So, Hod Lipson

机构 * Department of Computer Science(计算机科学系) Columbia University(哥伦比亚大学) Department of Mechanical Engineering(机械工程系)

专题命中 GUI与网页智能体 :agent(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21964 2025-11-04 cs.HC cs.CL 57%

UI-Evol: Automatic Knowledge Evolving for Computer Use Agents

Ziyun Zhang, Xinyi Liu, Xiaoyi Zhang, Jun Wang, Gang Chen, Yan Lu

机构 * School of Software and Microelectronics, Peking University(北京大学软件与微电子学院) Microsoft Research Asia(微软亚洲研究院)

专题命中 GUI与网页智能体 :agent(abstract);分类 cs.CL

Comments Accepted to ICML 2025 Workshop on Computer Use Agents

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00265 2025-11-04 cs.CL cs.CR 57%

AgentBnB: A Browser-Based Cybersecurity Tabletop Exercise with Large Language Model Support and Retrieval-Aligned Scaffolding

Arman Anwar, Zefang Liu

机构 * School of Electrical and Computer Engineering(电气与计算机工程学院) Georgia Institute of Technology(佐治亚理工学院) School of Computational Science and Engineering(计算科学与工程学院)

专题命中 GUI与网页智能体 :agent(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00033 2025-11-04 cs.RO cs.AI 57%

STRIDER: Navigation via Instruction-Aligned Structural Decision Space Optimization

Diqi He, Xuehao Gao, Hao Li, Junwei Han, Dingwen Zhang

机构 * Northwestern Polytechnical University(西北工业大学) Nanyang Technological University(南洋理工大学) Chongqing University of Posts and Telecommunications(重庆邮电大学)

专题命中 GUI与网页智能体 :agent(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26997 2025-11-03 cs.LG 57%

Gradient Descent as Loss Landscape Navigation: a Normative Framework for Deriving Learning Rules

John J. Vastola, Samuel J. Gershman, Kanaka Rajan

机构 * Harvard Medical School(哈佛医学院) Harvard University(哈佛大学)

专题命中 GUI与网页智能体 :planning(abstract);分类 cs.LG

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26040 2025-10-31 cs.RO cs.LG 57%

Accelerating Real-World Overtaking in F1TENTH Racing Employing Reinforcement Learning Methods

Emily Steiner, Daniel van der Spuy, Futian Zhou, Afereti Pama, Minas Liarokapis, Henry Williams

机构 * Centre for Automation and Robotic Engineering Science, The University of Auckland, New Zealand(自动化与机器人工程科学中心,奥克兰大学,新西兰) New Dexterity research group, The University of Auckland, New Zealand(新灵活性研究组,奥克兰大学,新西兰)

专题命中 GUI与网页智能体 :agent(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏