arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

共收录 15756 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. Agent评测 15756 篇

2002.06306 2020-02-18 cs.LG cs.AI cs.MA stat.ML 73%

Jelly Bean World: A Testbed for Never-Ending Learning

Emmanouil Antonios Platanios, Abulhair Saparov, Tom Mitchell

专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG

Comments Published as a conference paper at ICLR 2020

Journal ref International Conference on Learning Representations 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.01562 2019-11-06 cs.LG cs.AI cs.RO 73%

DeepRacer: Educational Autonomous Racing Platform for Experimentation with Sim2Real Reinforcement Learning

Bharathan Balaji, Sunil Mallya, Sahika Genc, Saurabh Gupta, Leo Dirac, Vineet Khare, Gourav Roy, Tao Sun, Yunzhe Tao, Brian Townsend, Eddie Calleja, Sunil Muralidhara, Dhanasekar Karuppasamy

专题命中 Agent评测 :agent(abstract);planning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.01401 2019-09-18 cs.LG cs.AI cs.RO stat.ML 73%

Unsupervised Emergence of Egocentric Spatial Structure from Sensorimotor Prediction

Alban Laflaquière, Michael Garcia Ortiz

专题命中 Agent评测 :agent(abstract);autonomous agent(abstract);分类 cs.AI、cs.LG

Comments 27 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.11788 2019-07-30 cs.LG cs.AI stat.ML 73%

On Hard Exploration for Reinforcement Learning: a Case Study in Pommerman

Chao Gao, Bilal Kartal, Pablo Hernandez-Leal, Matthew E. Taylor

专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG

Comments AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment (AIIDE) 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.06508 2019-07-16 cs.AI cs.LG stat.ML 73%

General Board Game Playing for Education and Research in Generic AI Game Learning

Wolfgang Konen

专题命中 Agent评测 :agent(abstract);AI agent(abstract);分类 cs.AI、cs.LG

Comments 8 pages, for: Conference on Games (CoG), London, 2019. Index Terms: game learning, general game playing, AI, temporal difference learning, board games, n-tuple systems

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.08809 2019-06-24 cs.LG cs.AI stat.ML 73%

A Deep Reinforcement Learning Approach for Global Routing

Haiguang Liao, Wentai Zhang, Xuliang Dong, Barnabas Poczos, Kenji Shimada, Levent Burak Kara

专题命中 Agent评测 :agent(abstract);planning(abstract);分类 cs.AI、cs.LG

Comments Preprint submitted to ASME JMD

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.03967 2019-06-11 cs.LG cs.AI cs.NE cs.RO stat.ML 73%

Autonomous Goal Exploration using Learned Goal Spaces for Visuomotor Skill Acquisition in Robots

Adrien Laversanne-Finot, Alexandre Péré, Pierre-Yves Oudeyer

专题命中 Agent评测 :agent(abstract);autonomous agent(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1902.03765 2019-02-12 cs.LG cs.AI cs.CV cs.RO stat.ML 73%

Latent Space Reinforcement Learning for Steering Angle Prediction

Qadeer Khan, Torsten Schön, Patrick Wenzel

专题命中 Agent评测 :agent(abstract);autonomous agent(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.05098 2018-11-07 cs.LG cs.AI cs.NE 73%

DiCE: The Infinitely Differentiable Monte-Carlo Estimator

Jakob Foerster, Gregory Farquhar, Maruan Al-Shedivat, Tim Rocktäschel, Eric P. Xing, Shimon Whiteson

专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1705.05172 2017-11-30 cs.LG cs.AI cs.HC cs.RO stat.ML 73%

Emotion in Reinforcement Learning Agents and Robots: A Survey

Thomas M. Moerland, Joost Broekens, Catholijn M. Jonker

专题命中 Agent评测 :agent(abstract);AI agent(abstract);分类 cs.AI、cs.LG

Comments To be published in Machine Learning Journal

Journal ref Machine Learning 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1304.2024 2014-03-18 cs.LG cs.AI cs.MA stat.ML 73%

A General Framework for Interacting Bayes-Optimally with Self-Interested Agents using Arbitrary Parametric Model and Model Prior

Trong Nghia Hoang, Kian Hsiang Low

专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG

Comments 23rd International Joint Conference on Artificial Intelligence (IJCAI 2013), Extended version with proofs, 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24010 2026-07-28 cs.LG 新提交 72%

When Should Active RAG Retrieve? A Budget-Aware Evaluation of Utility, Calibration, and Cost

主动检索生成式人工智能何时应进行检索?效用、校准和成本的预算感知评估

Pin Qian, Su Wang, Chong Peng, Junxian You, Lifei Liu, Haoran Yu, Yihang Chen, Xiaochong Jiang

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Glasgow(格拉斯哥大学) Georgia Institute of Technology(佐治亚理工学院)

专题命中 Agent评测 :agentic(abstract,abstract_cn);分类 cs.LG

AI总结 研究主动检索生成式人工智能何时检索,通过将其重述为效用估计进行预算感知评估,并分离出相关三个问题,利用多种方法实现,在多数据集和模型中验证,强调评估应报告多方面指标。

Comments Accepted at the ACM SIGKDD KDD 2026 Workshop on Evaluation and Trustworthiness of Agentic AI; 7 pages, 1 figure, and 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23975 2026-07-28 cs.AI q-bio.QM 新提交 72%

Plato-Bio: verification-first biological novelty screening with temporal rediscovery and structural benchmarks

柏拉图生物:通过时间再发现和结构基准进行验证优先的生物新奇性筛选

Stefan G. Creadore

机构 * Praxa Labs(普拉克斯实验室)

专题命中 Agent评测 :agent(abstract,comments);workflow(abstract);分类 cs.AI

AI总结 研究开发柏拉图生物,扩展开放架构并结合多种功能,修复评估缺陷。通过Python套件等验证其有效性,评估两个用例,如历史再发现任务及蛋白质结构比较,提供可重复软件契约和筛选基准,为生物研究提供支持。

Comments 16 pages, 6 figures, 3 tables. Companion code and data: https://github.com/Eldergenix/Plato-Scientific-Research-Autonomous-Agent. This fork-specific validation study cites, but does not duplicate, arXiv:2510.26887

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.08986 2026-07-21 cs.AI cs.LO math-ph math.AP math.MP 版本更新 72%

A Formalization of the Mean-Field Derivation of the Vlasov Equation

弗拉索夫方程平均场推导的形式化:作为策略游戏的人工智能辅助精益形式化

Joseph K. Miller

机构 * Stanford University(斯坦福大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 Agent评测 :agent(abstract,comments);AI agent(abstract);分类 cs.AI

AI总结 该研究以数学家指导AI在Lean 4中形式化研究成果为案例,将其构建为形式化游戏。通过此方式对非线性弗拉索夫方程适定性完整形式化,展示了开发过程及成果,还介绍了最优传输机制的分离情况及开发时间等,为形式化研究提供了新方法。

Comments 26 pages, 4 figures. Lean 4 development, blueprint site, and agent logs: https://github.com/Hydrodynamical/Vlasov_Meanfield_Formalization

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11619 2026-07-16 cs.AI 版本更新 72%

When Agents Disagree With Themselves: Behavioral Consistency as an Uncertainty Signal for LLM Agents

当智能体与自身意见相左:测量基于LLM的智能体的行为一致性

Aman Mehta

机构 * Aman Mehta

专题命中 Agent评测 :agentic(abstract,comments);agent(abstract);分类 cs.AI

AI总结 研究发现基于LLM的智能体在相同任务上运行结果不一致,且这种不一致与任务成功率密切相关,通过监控行为一致性可提升智能体可靠性。

Comments Accepted at the ICML 2026 Workshop on Statistical Frameworks for Uncertainty in Agentic Systems. 12 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11672 2026-06-11 cs.CR cs.AI 新提交 72%

Can Open-Source LLM Agents Replace Static Application Security Testing Tools? An Empirical Assessment

开源LLM代理能否取代静态应用安全测试工具?一项实证评估

Derek Yohn, Luke Flancher, Mirajul Islam, Khaled Slhoub

机构 * College of Engineering and Science, Florida Institute of Technology(工程学院与科学学院,佛罗里达理工学院)

专题命中 Agent评测 :agentic(abstract,comments);agent(abstract);分类 cs.AI

AI总结 评估基于开源LLM的代理在静态应用安全测试中的性能,与SAST工具Bandit对比,发现当前不适合实际应用。

Comments Keywords: Agentic AI, Cybersecurity, Large Language Models, Static Application Security Testing, Model performance evaluation

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04229 2025-10-20 cs.HC cs.AI 72%

When AI Gets Persuaded, Humans Follow: Inducing the Conformity Effect in Persuasive Dialogue

Rikuo Sasaki, Michimasa Inaba

机构 * The University of Electro-Communications(电子通信大学)

专题命中 Agent评测 :agent(abstract,comments);AI agent(abstract);分类 cs.AI

Comments 23 pages, 19 figures. International Conference on Human-Agent Interaction (HAI 2025), November 10-13, 2025, Yokohama, Japan

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09780 2025-10-03 cs.AI 72%

AgentDAM: Privacy Leakage Evaluation for Autonomous Web Agents

Arman Zharmagambetov, Chuan Guo, Ivan Evtimov, Maya Pavlova, Ruslan Salakhutdinov, Kamalika Chaudhuri

机构 * FAIR at Meta(Meta 的公平性研究部门)

专题命中 Agent评测 :agent(abstract,comments);AI agent(abstract);分类 cs.AI

Comments Accepted to NeurIPS 2025 (D&B track), project page: https://github.com/facebookresearch/ai-agent-privacy

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.17918 2025-02-17 cs.AI 72%

AgentStudio: A Toolkit for Building General Virtual Agents

Longtao Zheng, Zhiyuan Huang, Zhenghai Xue, Xinrun Wang, Bo An, Shuicheng Yan

专题命中 Agent评测 :agent(abstract,comments);function calling(abstract);分类 cs.AI

Comments ICLR 2025. Project page: https://ltzheng.github.io/agent-studio

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.06399 2023-12-20 cs.AI 72%

Designing Behavior Trees from Goal-Oriented LTLf Formulas

Aadesh Neupane, Eric G Mercer, Michael A. Goodrich

专题命中 Agent评测 :autonomous agent(abstract,comments);agent(abstract);分类 cs.AI

Comments Accepted as "Most Visionary Paper" in Autonomous Robots and Multirobot Systems (ARMS) 2023 workshop affiliated with the 22nd International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.03450 2023-05-26 cs.LG cs.MA 72%

DeepFreight: Integrating Deep Reinforcement Learning and Mixed Integer Programming for Multi-transfer Truck Freight Delivery

Jiayu Chen, Abhishek K. Umrawal, Tian Lan, Vaneet Aggarwal

专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.LG;planning(comments)

Comments Citing the ICAPS version is preferred: Chen, Jiayu, Abhishek K. Umrawal, Tian Lan, and Vaneet Aggarwal. "DeepFreight: A Model-free Deep-reinforcement-learning-based Algorithm for Multi-transfer Freight Delivery." In Proceedings of the International Conference on Automated Planning and Scheduling, vol. 31, pp. 510-518. 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.02918 2020-11-06 cs.AI 72%

Domain-independent generation and classification of behavior traces

Daniel Borrajo, Manuela Veloso

专题命中 Agent评测 :planning(abstract,comments);agent(abstract);分类 cs.AI

Comments A version of this paper appears in the Pre-prints of the Workshop in Planning for Financial Services (FinPlan) at ICAPS'20. arXiv admin note: text overlap with arXiv:2011.01826

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17142 2026-08-19 eess.SY cs.SY 新提交 71%

A Hybrid Discrete-Event and Agent-Based Simulation Approach to Model Circular Supply Chains in Healthcare: A Case Study of Laparoscopic Scissors

用于建模医疗领域循环供应链的混合离散事件与基于智能体的仿真方法:以腹腔镜剪刀为例

Mohd Shoaib, Antuela Tako, Shahin Rahimifard

专题命中 Agent评测 :agent(title)

AI总结 本文以腹腔镜剪刀供应链为案例,采用混合离散事件与基于智能体的仿真方法,评估医疗领域引入循环医疗器械设计的影响,为医院提供最优库存策略,指出采用循环产品可降低环境影响但需大量前期投资。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00867 2026-08-19 cs.HC cs.MA 版本更新 71%

Interactionalism: Re-Designing Higher Learning for the Large Language Agent Era

互动主义:为大语言智能体时代重新设计高等学习

Mihnea C. Moldoveanu, George Siemens

专题命中 Agent评测 :agent(title)

AI总结 该研究提出互动主义作为与生成式AI协同的高等学习实践蓝图,定义互动智能为核心认知任务可由大语言模型智能体自动化时的关键技能,拆解为元认知与元情感组件,用于培养学习者。

Comments 37 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08277 2026-08-11 cs.NI 新提交 71%

WirelessOpsAgent: A Benchmark and Agent Design for Action Assurance in Wireless Networks

WirelessOpsAgent:无线网络行动保障的基准测试与智能体设计

Zijian Lu, Yiping Zuo, Hao Xu, Weicong Chen, Xin He, Jiajia Guo, Shi Jin

专题命中 Agent评测 :agent(title)

AI总结 针对现有无线网络运维LLM智能体基准未测试执行阶段支持检查的问题,本文提出WirelessOptBench基准与WirelessOpsAgent智能体,使后者在600片段评估中精确行动准确率达0.983,不安全执行率大幅下降。

Comments 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05485 2026-08-07 cs.CV 新提交 71%

VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing

VideoArgus:基于智能体评分规则的视频生成与编辑统一评估框架

Ziyun Zeng, Zixuan Wang, Yongsheng Yu, Hang Hua, Jiebo Luo

机构 * University of Rochester(罗切斯特大学)

专题命中 Agent评测 :agentic(title)

AI总结 针对现有视频评估基准的局限,本文提出VideoArgus框架,构建含1026个实例的VideoArgus-Bench,其在与人类判断的相关性上优于基准特定评估器,且模型排名稳定。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15696 2026-07-20 cs.IR 新提交 71%

PCTD: Preference-Guided Counterfactual Task Decomposition for Agent Tool Retrieval

PCTD:用于智能体工具检索的偏好引导反事实任务分解

Chu Zhao, Lei Tang, Minghang Li, Jianzhe Zhao, Guibing Guo, Zhengzong Chen, Yuanyuan Zhao, Fei Huang

专题命中 Agent评测 :agent(title)

AI总结 研究针对任务分解中因直接用检索指标作奖励易致奖励作弊及削弱域外泛化能力的问题,提出偏好引导反事实任务分解框架PCTD,通过反事实和偏好奖励改进,构建基准MTDTool,实验证明其在多方面超越现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.00889 2026-07-07 econ.TH 版本更新 71%

Mr.Keynes and the... Complexity! A suggested model for the General Theory

凯恩斯先生与……复杂性!《通论》的一个建议模型

Alessio Emanuele Biondo

专题命中 Agent评测 :agent(title)

AI总结 提出凯恩斯《就业、利息和货币通论》的数学模型,核心方法是依据原文,主要贡献是表明流动性偏好、资本边际效率和边际消费倾向决定既定利率下的总收益与就业水平。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18038 2026-06-19 cs.CY 版本更新 71%

Acceleration AI Ethics and the Telus GenAI Conversational Agent

加速AI伦理与Telus生成式AI对话代理

James Brusseau

专题命中 Agent评测 :agent(title)

AI总结 本文阐述加速伦理学的理论框架,并通过Telus公司的生成式AI语言工具案例,展示加速AI伦理如何在创新与安全之间平衡,以最大化社会责任。

Journal ref Law Ethics Technol. 2026(2):0006

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08636 2026-06-09 eess.SY cs.SY 新提交 71%

Cooperative Guidance and Control for Active Asset Protection with Time-Varying Agent Speeds

时变速度下主动资产保护的协同制导与控制

Ram Milan Kumar Verma, Shashi Ranjan Kumar, Hemendra Arya

专题命中 Agent评测 :agent(title)

AI总结 提出一种资产与防御者协同的制导控制策略,通过时变速度与航向协调,结合视线率归零、防御者保持在视线线上以及基于剩余时间引导的三种几何与时间目标,实现对机动威胁的拦截。

详情

展开后加载摘要…

URL PDF HTML 收藏