arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

共收录 15801 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. Agent评测 15801 篇

2508.03728 2025-08-07 cs.CL 77%

WINELL: Wikipedia Never-Ending Updating with LLM Agents

Revanth Gangi Reddy, Tanay Dixit, Jiaxin Qin, Cheng Qian, Daniel Lee, Jiawei Han, Kevin Small, Xing Fan, Ruhi Sarikaya, Heng Ji

专题命中 Agent评测 :agent(abstract);agentic(abstract);multi-agent(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10911 2025-07-16 cs.AI 77%

Lessons Learned from Evaluation of LLM based Multi-agents in Safer Therapy Recommendation

Yicong Wu, Ting Chen, Irit Hochberg, Zhoujian Sun, Ruth Edry, Zhengxing Huang, Mor Peleg

专题命中 Agent评测 :agent(abstract);planning(abstract);multi-agent(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09389 2025-07-15 cs.AI cs.CY cs.IR 77%

Knowledge Conceptualization Impacts RAG Efficacy

Chris Davis Jaldi, Anmol Saini, Elham Ghiasi, O. Divine Eziolise, Cogan Shimizu

机构 * Wright State University(怀特州立大学)

专题命中 Agent评测 :agent(abstract);AI agent(abstract);agentic(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17560 2025-06-24 cs.MA cs.AI 77%

Towards Zero-Shot Coordination between Teams of Agents: The N-XPlay Framework

Ava Abderezaei, Chi-Hui Lin, Joseph Miceli, Naren Sivagnanadasan, Stéphane Aroca-Ouellette, Jake Brawer, Alessandro Roncone

专题命中 Agent评测 :agent(abstract);autonomous agent(abstract);multi-agent(abstract);分类 cs.AI

Comments Accepted to RSS Workshop on Scalable and Resilient Multi-Robot Systems: Decision-Making, Coordination, and Learning 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12110 2025-06-17 econ.GN cs.AI q-fin.EC 77%

EconGym: A Scalable AI Testbed with Diverse Economic Tasks

Qirui Mi, Qipeng Yang, Zijun Fan, Wentian Fan, Heyang Ma, Chengdong Ma, Siyu Xia, Bo An, Jun Wang, Haifeng Zhang

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, Chinese Academy of Sciences(中国科学院人工智能学院) Nanyang Technological University(南洋理工大学) Nanjing University of Posts and Telecommunications(南京邮电大学) University of Chinese Academy of Sciences, Nanjing(中国科学院大学(南京)) Nanjing Artificial Intelligence Research of IA(南京人工智能研究所) Peking University(北京大学) University of International Business and Economics(对外经济贸易大学) University College London(伦敦大学学院)

专题命中 Agent评测 :agent(abstract);AI agent(abstract);multi-agent(abstract);分类 cs.AI

Comments 28 pages, 7 figures, 17 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11019 2025-06-16 cs.SE 77%

Mind the Metrics: Patterns for Telemetry-Aware In-IDE AI Application Development using the Model Context Protocol (MCP)

Vincent Koc, Jacques Verre, Douglas Blank, Abigail Morgan

专题命中 Agent评测 :agent(abstract);autonomous agent(abstract);workflow(abstract);分类 cs.SE

Comments 16 pages, 5 figures, conference preprint submission. Conceptual systems architecture paper on telemetry-driven prompt optimization and IDE design patterns for AI development. Builds on Opik MCP open-source architecture and Comet trace infrastructure

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07320 2025-05-29 cs.HC cs.AI 77%

When Trust Collides: Decoding Human-LLM Cooperation Dynamics through the Prisoner's Dilemma

Guanxuan Jiang, Shirao Yang, Yuyang Wang, Pan Hui

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 Agent评测 :agent(abstract);AI agent(abstract);autonomous agent(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02391 2025-05-27 cs.CL 77%

Attacking Vision-Language Computer Agents via Pop-ups

Yanzhe Zhang, Tao Yu, Diyi Yang

机构 * Georgia Tech(佐治亚理工学院) The University of Hong Kong(香港大学) Stanford University(斯坦福大学)

专题命中 Agent评测 :agent(abstract);autonomous agent(abstract);agentic(abstract);分类 cs.CL

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01039 2025-04-03 cs.CY cs.AI 77%

One Person, One Bot

Liat Lavi

专题命中 Agent评测 :agent(abstract);AI agent(abstract);agentic(abstract);分类 cs.AI

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10050 2025-03-05 cs.IR cs.AI 77%

A Survey on LLM-powered Agents for Recommender Systems

Qiyao Peng, Hongtao Liu, Hua Huang, Qing Yang, Minglai Shao

专题命中 Agent评测 :agent(abstract);planning(abstract);multi-agent(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11645 2025-02-18 cs.GT cs.CL cs.MA stat.OT 77%

Deviation Ratings: A General, Clone-Invariant Rating Method

Luke Marris, Siqi Liu, Ian Gemp, Georgios Piliouras, Marc Lanctot

专题命中 Agent评测 :agent(abstract);agentic(abstract);multi-agent(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23242 2025-01-06 cs.AI 77%

A little less conversation, a little more action, please: Investigating the physical common-sense of LLMs in a 3D embodied environment

Matteo G. Mecattaf, Ben Slater, Marko Tešić, Jonathan Prunty, Konstantinos Voudouris, Lucy G. Cheke

专题命中 Agent评测 :agent(abstract);tool use(abstract);agentic(abstract);分类 cs.AI

Comments 25 pages, 4 figures; v2: Added AFMR Acknowledgment

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22553 2024-10-31 cs.AI 77%

ML Research Benchmark

Matthew Kenney

专题命中 Agent评测 :agent(abstract);AI agent(abstract);planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03225 2024-10-29 cs.CR cs.AI 77%

AutoPenBench: Benchmarking Generative Agents for Penetration Testing

Luca Gioacchini, Marco Mellia, Idilio Drago, Alexander Delsanto, Giuseppe Siracusano, Roberto Bifulco

专题命中 Agent评测 :agent(abstract);AI agent(abstract);autonomous agent(abstract);分类 cs.AI

Comments Codes for the benchmark: https://github.com/lucagioacchini/auto-pen-bench Codes for the paper experiments: https://github.com/lucagioacchini/genai-pentest-paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08940 2024-07-16 cs.CL 77%

Large Language Models as Biomedical Hypothesis Generators: A Comprehensive Evaluation

Biqing Qi, Kaiyan Zhang, Kai Tian, Haoxiang Li, Zhang-Ren Chen, Sihang Zeng, Ermo Hua, Hu Jinfang, Bowen Zhou

专题命中 Agent评测 :agent(abstract);tool use(abstract);multi-agent(abstract);分类 cs.CL

Comments Accepted to COLM 2024. This is an extended version of the paper at arXiv:2311.05965

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03575 2024-06-21 cs.AI cs.HC 77%

Toward Human-AI Alignment in Large-Scale Multi-Player Games

Sugandha Sharma, Guy Davidson, Khimya Khetarpal, Anssi Kanervisto, Udit Arora, Katja Hofmann, Ida Momennejad

专题命中 Agent评测 :agent(abstract);AI agent(abstract);multi-agent(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.11865 2024-06-19 cs.AI 77%

Large Language Models Play StarCraft II: Benchmarks and A Chain of Summarization Approach

Weiyu Ma, Qirui Mi, Yongcheng Zeng, Xue Yan, Yuqiao Wu, Runji Lin, Haifeng Zhang, Jun Wang

专题命中 Agent评测 :agent(abstract);AI agent(abstract);planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.16100 2024-03-26 cs.AI 77%

Specifying Agent Ethics (Blue Sky Ideas)

Louise A. Dennis, Michael Fisher

专题命中 Agent评测 :agent(title,comments);分类 cs.AI;multi-agent(comments)

Comments To appear in Coordination, Organizations, Institutions, Norms and Ethics for Governance of Multi-Agent Systems 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06853 2023-12-14 cs.AI 77%

LLF-Bench: Benchmark for Interactive Learning from Language Feedback

Ching-An Cheng, Andrey Kolobov, Dipendra Misra, Allen Nie, Adith Swaminathan

专题命中 Agent评测 :agent(abstract);AI agent(abstract);planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08710 2023-10-16 cs.RO cs.LG 77%

Waymax: An Accelerated, Data-Driven Simulator for Large-Scale Autonomous Driving Research

Cole Gulino, Justin Fu, Wenjie Luo, George Tucker, Eli Bronstein, Yiren Lu, Jean Harb, Xinlei Pan, Yan Wang, Xiangyu Chen, John D. Co-Reyes, Rishabh Agarwal, Rebecca Roelofs, Yao Lu, Nico Montali, Paul Mougin, Zoey Yang, Brandyn White, Aleksandra Faust, Rowan McAllister, Dragomir Anguelov, Benjamin Sapp

专题命中 Agent评测 :agent(abstract);planning(abstract);multi-agent(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.01605 2023-02-06 cs.AI 77%

Learning Zero-Shot Cooperation with Humans, Assuming Humans Are Biased

Chao Yu, Jiaxuan Gao, Weilin Liu, Botian Xu, Hao Tang, Jiaqi Yang, Yu Wang, Yi Wu

专题命中 Agent评测 :agent(abstract);workflow(abstract);multi-agent(abstract);分类 cs.AI

Comments The first two authors share equal contributions. This paper is accepted by ICLR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.11703 2019-07-30 cs.LG cs.MA stat.ML 77%

Action Guidance with MCTS for Deep Reinforcement Learning

Bilal Kartal, Pablo Hernandez-Leal, Matthew E. Taylor

专题命中 Agent评测 :agent(abstract);planning(abstract);multi-agent(abstract);分类 cs.LG

Comments AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment (AIIDE'19). arXiv admin note: substantial text overlap with arXiv:1904.05759, arXiv:1812.00045

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19576 2026-07-30 cs.AI cs.CL cs.SE 版本更新 76%

Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries

库漂移:在自我演化的LLM技能库中诊断和修复一种无声的失败模式

Xing Zhang, Yanwei Cui, Guanghui Wang, Ziyuan Li, Wei Qiu, Bing Zhu, Peiyang He

机构 * AWS Generative AI Innovation Center(AWS生成式AI创新中心) HSBC Holdings Plc., HSBC Technology Center, China(汇丰控股有限公司,汇丰技术中心,中国)

专题命中 Agent评测 :agent(abstract,abstract_cn);分类 cs.AI、cs.CL、cs.SE;agentic(comments)

AI总结 本文研究了自我演化的LLM技能库中的一种无声失败模式——库漂移,通过可重复触发实验、细粒度诊断和验证修复方法,揭示了技能积累无序导致检索退化、假阳性注入和性能停滞的问题,并提出了一种经过验证的修复方案,显著提升了技能库的性能。

Comments Accepted to the ICML 2026 Workshop on Failure Modes in Agentic AI (FAGEN@ICML 2026), Seoul, South Korea. https://github.com/amazon-science/Self-Evolving-Agents-Ratchet

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20493 2026-02-25 cs.NI cs.MA 76%

AWCP: A Workspace Delegation Protocol for Deep-Engagement Collaboration across Remote Agents

AWCP:用于远程代理深度协作的工件委托协议

Xiaohang Nie, Zihan Guo, Youliang Chen, Yuanjian Zhou, Weinan Zhang

专题命中 Agent评测 :agent(abstract,comments);autonomous agent(abstract);agentic(abstract)

AI总结 AWCP通过工件委托协议实现远程代理的深度协作,提供开源实现以提升代理间协作的互操作性。

Comments 16 pages, 7 figure, tech report of Agent Workspace Collaboration Protocol

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07744 2024-02-15 cs.AI cs.CL cs.LG 76%

Towards Unified Alignment Between Agents, Humans, and Environment

Zonghan Yang, An Liu, Zijun Liu, Kaiming Liu, Fangzhou Xiong, Yile Wang, Zeyuan Yang, Qingyuan Hu, Xinrui Chen, Zhenhe Zhang, Fuwen Luo, Zhicheng Guo, Peng Li, Yang Liu

专题命中 Agent评测 :agent(abstract,comments);autonomous agent(abstract);分类 cs.AI、cs.CL、cs.LG

Comments Project webpage: https://agent-force.github.io/unified-alignment-for-agents.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20376 2026-08-24 cs.CL cs.LG 新提交 76%

TH-GNN: Heterogeneous Temporal Graph Neural Networks for LLM-Agent Shilling Attack Detection

TH-GNN:用于检测大语言模型智能体刷单攻击的异质性时序图神经网络

Shivam Swarup, Divya Prakash Shrivastava, Rakesh Thakur

机构 * JAIN (Deemed to be University)(JAIN(被认定大学)) Zayed University(扎耶德大学)

专题命中 Agent评测 :agent(title);分类 cs.CL、cs.LG

AI总结 本文提出TH-GNN,一种异质性时序图神经网络,联合建模时序、结构与语义信号,在5类攻击族和4个基准数据集上检测LLM智能体刷单攻击,总体平均F1值达0.870,性能优于纯文本基线模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04772 2026-08-19 cs.CL cs.AI 版本更新 76%

Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent

以指南为神谕:眼科电话分诊智能体的零标注训练

Chenyu Wang, Yi Liu, Baoqing Li, Min Tu, Diping Song

专题命中 Agent评测 :agent(title);分类 cs.AI、cs.CL

AI总结 本研究提出以指南为神谕(GAO)方法,将美国眼科学会指南转化为训练监督信号,零标注训练出GAO-Triage智能体,大幅提升眼科电话分诊的一致性与紧急案例召回率,且性能优于7个通用系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15654 2026-08-18 cs.CL cs.AI 新提交 76%

When Stories Evolve: Benchmarking LLM Storytelling Across Agent Architectures in Open-Ended World Simulations

当故事演变时:在开放世界模拟中针对智能体架构对大语言模型故事讲述能力进行基准测试

Yuqi Chen, Sixuan Li, Yunfeng Cai, Xueai Li, Ka Man Yan, Ying Li

机构 * The University of Hong Kong(香港大学) Peking University(北京大学) Tsinghua University(清华大学) Beijing Institute of Mathematical Sciences and Applications (BIMSA)(北京数学科学与应用研究院)

专题命中 Agent评测 :agent(title);分类 cs.AI、cs.CL

AI总结 该研究推出WSE-bench基准,评估开放世界模拟中不同智能体架构的大语言模型故事讲述的持续生成、规范一致性和有意义发展,发现三者存在竞争关系,模型规模仅提升持续生成能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15286 2026-08-18 cs.LG cs.AI 新提交 76%

No Task Fails Every Time: Why One-Shot Audits Are Structurally Blind to Agent Damage

没有任务每次都失败:为什么单次审计在智能体损伤问题上存在结构性盲区

Shiven Khurdi

机构 * Northeastern University(东北大学)

专题命中 Agent评测 :agent(title);分类 cs.AI、cs.LG

AI总结 该研究提出AgentRelBench工具,发现单次智能体审计易遗漏损伤对,模型能力提升会减少损伤任务,部分模型存在宣称弃权却执行不可逆操作的情况,且所有发现均按预注册标准执行。

Comments 25 pages, 4 figures, 16 tables, 6 appendices. Code, task suite, released per-run verdicts, and a one-command reproduction of every reported number: https://github.com/shivenkk/agentrelbench

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02636 2026-08-05 cs.SE cs.AI 新提交 76%

Rethinking Self-Evolving Agent Skills: Feedback Dynamics over Multiple Rounds

反思自进化智能体技能:多轮次的反馈动态

Yuxuan Liu, Zhaochen Su, Yuhao Zhang, Jiahe Guo, Zhongwei Xie, Huihao Jing, Lingyun Xie, Qing Zong, Yauwai Yim, Zhixiong Zhang, Haoran Li, Yangqiu Song

专题命中 Agent评测 :agent(title);分类 cs.AI、cs.SE

AI总结 本研究提出受控评估框架,发现自进化智能体技能是稀疏的验证过滤搜索,收益依赖模型与基准,失败轨迹反馈对技能选择关键,测试时计算难以完全恢复其收益。

详情

展开后加载摘要…

URL PDF HTML 收藏