arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

共收录 15729 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. Agent评测 15729 篇

1403.7221 2015-06-19 astro-ph.IM 67%

DIRAC framework evaluation for the $\boldsymbol{Fermi}$-LAT and CTA experiments

Luisa Arrabito, Johann Cohen-Tanugi, Ricardo Graciani Diaz, Francesco Longo, Michael Kuss, Frédéric Piron, Matthieu Renaud, Vincent Rolland, Matvey Sapunov, Andreï Tsaregorodtsev, Stephan Zimmer

专题命中 Agent评测 :agent(abstract);workflow(abstract)

Comments proceedings to CHEP 2013 conference : http://www.chep2013.org/

详情

展开后加载摘要…

URL PDF HTML 收藏
1411.4687 2014-11-19 math.OC 67%

Control to flocking of the kinetic Cucker-Smale model

Benedetto Piccoli, Francesco Rossi, Emmanuel Trélat

专题命中 Agent评测 :agent(abstract);autonomous agent(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1406.6404 2014-10-28 math.OC 67%

A Class of Randomized Primal-Dual Algorithms for Distributed Optimization

Jean-Christophe Pesquet, Audrey Repetti

专题命中 Agent评测 :agent(abstract);multi-agent(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1409.8035 2014-10-13 cs.CR 67%

Detecting Behavioral and Structural Anomalies in MediaCloud Applications

Guido Schwenk, Sebastian Bach

专题命中 Agent评测 :agent(abstract);planning(abstract)

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1407.8454 2014-08-01 physics.ao-ph 67%

Conceptual design of a tropical cyclone UAV based on the AR-6 Endeavor aircraft

Chung-Kiak Poh, Chung-How Poh, Mei-Ling Yeh, Tien-Yin Chou

专题命中 Agent评测 :agent(abstract);multi-agent(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1404.7717 2014-07-21 cs.MA 67%

A Checklist for the Evaluation of Pedestrian Simulation Software Functionalities

Mizar Luca Federici, Lorenza Manenti, Sara Manzoni

专题命中 Agent评测 :agent(abstract);planning(abstract)

Comments 20 pages, 3 figures, 1 Table

详情

展开后加载摘要…

URL PDF HTML 收藏
1401.8132 2014-02-03 cs.MA 67%

Heterogeneous Speed Profiles in Discrete Models for Pedestrian Simulation

Stefania Bandini, Luca Crociani, Giuseppe Vizzari

专题命中 Agent评测 :agent(abstract);multi-agent(abstract)

Comments Poster at the 93rd Transportation Research Board annual meeting, Washington, January 2014 - Committee number AHB45 - TRB Committee on Traffic Flow Theory and Characteristics

详情

展开后加载摘要…

URL PDF HTML 收藏
1302.5952 2013-02-26 cond-mat.soft cond-mat.stat-mech nlin.AO 67%

Swarming, swirling and stasis in sequestered bristle-bots

L. Giomi, N. Hawley-Weld, L. Mahadevan

专题命中 Agent评测 :agent(abstract);autonomous agent(abstract)

Comments 21 pages, 15 figures, supplementary videos at https://www.youtube.com/user/BristleBotChannel

Journal ref Proc. R. Soc. A 469, 20120637 (2013)

详情

展开后加载摘要…

URL PDF HTML 收藏
0810.4000 2009-12-01 q-fin.TR cs.GL 67%

Le trading algorithmique

Victor Lebreton

专题命中 Agent评测 :agent(abstract);autonomous agent(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
0710.2659 2009-12-01 cs.MA cs.DM 67%

Rigidity and persistence for ensuring shape maintenance of multiagent meta formations (ext'd version)

Julien M. Hendrickx, Changbin Yu, Baris Fidan, Brian D. O. Anderson

专题命中 Agent评测 :agent(abstract);autonomous agent(abstract)

Comments 1 zip file containing 1 .tex files, and 39 .eps files. The paper (including the appendix) contains 13 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
cs/0702091 2009-12-01 cs.MA 67%

Observable Graphs

Raphael M. Jungers, Vincent D. Blondel

专题命中 Agent评测 :agent(abstract);autonomous agent(abstract)

Comments 15 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
cs/0410058 2009-12-01 cs.CL cs.AI cs.HC cs.MA cs.SE 67%

Robust Dialogue Understanding in HERALD

Vincenzo Pallotta, Afzal Ballim

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL、cs.SE

Comments 6 pages

Journal ref Proceedings of RANLP 2001 - EuroConference on Recent Advances in Natural Language Processing, September 5-7, 2001, Tzigov-Chark, Bulgaria

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07436 2026-08-19 cs.AI cs.CL cs.CR 版本更新 66%

The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents

盲选策展人:有偏见的评判者如何在自我进化的智能体中悄然阻碍技能淘汰

Xing Zhang, Yanwei Cui, Guanghui Wang, Ziyuan Li, Wei Qiu, Bing Zhu, Peiyang He

机构 * AWS Generative AI Innovation Center(亚马逊云科技生成式人工智能创新中心) HSBC Holdings Plc., HSBC Technology Center, China(汇丰控股有限公司,汇丰科技中心,中国)

专题命中 Agent评测 :agent(abstract,comments);分类 cs.AI、cs.CL

AI总结 研究自我进化智能体中,有偏见的评判者对技能淘汰的影响。通过损坏奖励分析等方法发现,“误判通过”偏差会导致技能淘汰机制失效,此为行为安全结果。还提出廉价审计可判断评判者是否越过阈值影响智能体。

Comments Published at COLM 2026 Workshop on Agent Behavior

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.19449 2026-07-23 cs.LG cs.AI cs.CR 新提交 66%

Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents

作为替罪羊的护栏:审计工具增强型大语言模型智能体中不忠实的安全拒绝行为

Aarushi Singh

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG;agentic(comments)

AI总结 研究工具增强型大语言模型智能体安全拒绝行为审计问题,引入轻量级黑盒审计框架,将智能体响应分类。实验发现伪造行为占主导,不忠实安全拒绝行为在基线时少,增强安全语言会显著增加该行为,还提出检测方法及治理影响。

Comments 10 pages, 3 figures. Accepted at the ACM KDD 2026 Workshop on Evaluation and Trustworthiness of Agentic AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00392 2026-07-07 cs.SE cs.AI 版本更新 66%

Beyond Task Completion: A Verification-vs.-Conformance Gap in Tool-Evolving Agents

EvolveTool-Bench:评估LLM生成的工具库作为软件 artifact 的质量

Alibek Kaliyev, Artem Maryanskyy

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校) Uber Technologies(优步科技公司)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.SE;agentic(comments)

AI总结 本文提出EvolveTool-Bench,通过评估LLM生成的工具库在软件工程流程中的质量,揭示任务完成率之外的软件质量风险,强调需将工具库视为首要软件 artifact。

Comments 11 pages, 4 figures; accepted at KDD 2026 Workshop on Agentic AI Evaluation and Trustworthiness

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15322 2026-03-10 cs.AI cs.CL 66%

Replayable Financial Agents: A Determinism-Faithfulness Assurance Harness for Tool-Using LLM Agents

可重播的金融代理:用于工具使用LLM代理的确定性-忠实性保证框架

Raffi Khatchadourian

专题命中 Agent评测 :agentic(abstract,comments);分类 cs.AI、cs.CL

AI总结 本文提出DFAH框架,用于评估金融代理的确定性和忠实性,发现决策确定性与准确性无显著相关性,强调需独立测量两者的多维评估方法。

Comments 27 pages, 5 figures, 9 tables | Code and data: https://github.com/ibm-client-engineering/output-drift-financial-llms | To appear in the 2nd ICLR Workshop on Advances in Financial AI: Towards Agentic and Responsible Systems (ICLR 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14321 2026-02-17 cs.GT cs.AI cs.LG cs.MA 66%

Offline Learning of Nash Stable Coalition Structures with Possibly Overlapping Coalitions

离线学习纳什稳定的联盟结构(可能有重叠的联盟)

Saar Cohen

机构 * Bar Ilan University(巴伊兰大学) University of Oxford(牛津大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG;autonomous agent(comments)

AI总结 本文提出了一种允许重叠联盟的离线学习模型,通过分析代理级和联盟级效用反馈,设计样本高效的算法以推断纳什稳定的联盟结构。

Comments To Appear in the 25th International Conference on Autonomous Agents and Multiagent Systems (AAMAS), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22809 2025-09-08 cs.CL cs.AI cs.HC 66%

First Steps Towards Overhearing LLM Agents: A Case Study With Dungeons & Dragons Gameplay

Andrew Zhu, Evan Osgood, Chris Callison-Burch

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL;AI agent(comments)

Comments 9 pages, 5 figures. COLM 2025 Workshop on AI Agents

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20368 2025-04-30 cs.MA cs.AI cs.LG 66%

AKIBoards: A Structure-Following Multiagent System for Predicting Acute Kidney Injury

David Gordon, Panayiotis Petousis, Susanne B. Nicholas, Alex A. T. Bui

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG;autonomous agent(comments)

Comments Accepted at International Conference on Autonomous Agents and Multiagent Systems (AAMAS) Workshop, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14565 2024-09-24 cs.HC cs.AI cs.LG cs.MA cs.RO 66%

Combating Spatial Disorientation in a Dynamic Self-Stabilization Task Using AI Assistants

Sheikh Mannan, Paige Hansen, Vivekanand Pandey Vimal, Hannah N. Davies, Paul DiZio, Nikhil Krishnaswamy

专题命中 Agent评测 :AI agent(abstract);分类 cs.AI、cs.LG;agent(comments)

Comments 10 pages, To be published in the International Conference on Human-Agent Interaction (HAI '24) proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.07099 2024-04-11 cs.LG cs.AI 66%

Rethinking Out-of-Distribution Detection for Reinforcement Learning: Advancing Methods for Evaluation and Detection

Linas Nasvytis, Kai Sandbrink, Jakob Foerster, Tim Franzmeyer, Christian Schroeder de Witt

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG;autonomous agent(comments)

Comments Accepted as a full paper to the 23rd International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.07193 2019-05-20 cs.LG cs.AI cs.RO stat.ML 66%

MaMiC: Macro and Micro Curriculum for Robotic Reinforcement Learning

Manan Tomar, Akhil Sathuluri, Balaraman Ravindran

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG;autonomous agent(comments)

Comments To appear in the Proceedings of the 18th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2019). (Extended Abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08220 2026-08-11 cs.AI 新提交 65%

Metanormative Theory for RL-Based Moral Agents

基于强化学习的道德智能体的元规范理论

Aleks Knoks, Marija Slavkovik

机构 * University of Luxembourg(卢森堡大学) University of Bergen(卑尔根大学)

专题命中 Agent评测 :agent(abstract,comments);分类 cs.AI;multi-agent(comments)

AI总结 本文针对强化学习(RL)道德智能体设计中哲学文献被边缘化的问题,提炼元规范理论的相关理念审视RL架构,以明确RL智能体道德行为的判定标准,为评估相关RL方法奠定基础。

Comments The 14th International Workshop on Engineering Multi-Agent Systems (EMAS 2026) held May 25-26, 2026 Co-located with AAMAS 2026 Paphos, Cyprus

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23420 2026-06-04 cs.AI 65%

Bilevel Autoresearch: Meta-Autoresearching Itself

双层自动研究:元自动研究自身

Yaonan Qu, Meng Lu

机构 * Independent Researcher(独立研究者)

专题命中 Agent评测 :agent(abstract);分类 cs.AI;AI agent(comments);agentic(comments)

AI总结 提出双层自动研究框架,外层循环通过读取内层循环代码和轨迹、识别瓶颈并注入可执行Python搜索机制来改进内层循环,在GPT预训练基准上实现5倍改进。

Comments 16 pages, 5 figures, 3 tables. v2 expands the framing as mechanism-level agentic self-improvement and updates related work and limitations; core method and experiments unchanged. This paper was primarily drafted by AI agents with human oversight and direction

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03472 2026-03-02 cs.RO cs.AI cs.MA 65%

Destination-to-Chutes Task Mapping Optimization for Multi-Robot Coordination in Robotic Sorting Systems

多机器人协同中的目的地到传送口任务映射优化

Yulun Zhang, Alexandre O. G. Barbosa, Federico Pecora, Jiaoyang Li

机构 * Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所)

专题命中 Agent评测 :planning(abstract);分类 cs.AI;agent(comments);multi-agent(comments)

AI总结 本文提出基于进化算法和混合整数线性规划的任务映射优化方法,用于提升多机器人分拣系统的吞吐量。

Comments Accepted to IEEE International Symposium on Multi-Robot and Multi-Agent Systems (MRS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13783 2025-06-24 physics.soc-ph cs.LG 65%

Infected Smallville: How Disease Threat Shapes Sociality in LLM Agents

Soyeon Choi, Kangwook Lee, Oliver Sng, Joshua M. Ackerman

机构 * Department of Psychology, University of Wisconsin-Madison(威斯康星大学麦迪逊分校心理学系) Department of Electrical and Computer Engineering, University of Wisconsin-Madison(威斯康星大学麦迪逊分校电气与计算机工程系) Department of Psychological Science, University of California, Irvine(加州大学伊文斯分校心理学科学系) Department of Psychology, University of Michigan, Ann Arbor(密歇根大学安娜堡分校心理学系)

专题命中 Agent评测 :agent(abstract,comments);分类 cs.LG;multi-agent(comments)

Comments ICML 2025 Workshop on Multi-Agent Systems in the Era of Foundation Models: Opportunities, Challenges and Futures

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.04250 2023-05-05 cs.LG stat.ML 65%

Learning How to Infer Partial MDPs for In-Context Adaptation and Exploration

Chentian Jiang, Nan Rosemary Ke, Hado van Hasselt

专题命中 Agent评测 :agent(abstract,comments);分类 cs.LG;multi-agent(comments)

Comments In proceedings of the Reincarnating Reinforcement Learning (RRL) Workshop at ICLR 2023 and the Neuro-Symbolic AI for Agent and Multi-Agent Systems (NeSyMAS) Workshop at AAMAS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
1507.02563 2017-11-23 cs.AI cs.MA 65%

Managing Autonomous Mobility on Demand Systems for Better Passenger Experience

Wen Shen, Cristina Lopes

专题命中 Agent评测 :agent(abstract,journal_ref);分类 cs.AI;multi-agent(journal_ref)

Journal ref Proceedings of the 18th International Conference on Principles and Practice of Multi-Agent Systems (PRIMA 2015). pp 20-35. Lecture Notes in Computer Science, vol 9387. Springer

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23955 2026-08-21 cs.AI cs.DC cs.LG cs.SI q-fin.CP 版本更新 62%

From Accuracy to Auditability: A Survey of Determinism in Financial AI Systems

从准确性到可审计性:金融AI系统中的确定性综述

Ruizhe Zhou, Xiaoyang Liu, Gaoyuan Du, Yi Zheng, Shouxi Ren, Deepayan Chakrabarti, Dengdu Jiang

专题命中 Agent评测 :agentic(abstract);分类 cs.AI、cs.LG

AI总结 本文从系统视角综述了金融AI中表格模型、图网络和基于LLM的智能体工作流三种模态的不可重现性问题,通过实验量化了确定性指标并提出了分层评估框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19073 2026-08-20 cs.AI cs.LG stat.ML 新提交 62%

Robust Risk Under Evolving Uncertainty: A Wasserstein Counterpart of the Entropic Value-at-Risk

演化不确定性下的鲁棒风险:熵值风险(Entropic Value-at-Risk)的Wasserstein对应

Deep Kumar Ganguly, Jan Křetínský

机构 * Technical University of Munich(慕尼黑工业大学) Masaryk University(马萨里克大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

AI总结 该研究针对演化不确定性下的鲁棒风险问题,提出Wasserstein熵值风险,弥补熵值风险无法对冲名义模型认定不可能的灾难的缺陷,经数值验证其变分对偶性,并构造出随信念变化的闭式鲁棒动态规划算子。

Comments Best Paper Award at the 2nd Workshop on Safe AI at UAI (SafeAI@UAI 2026, non-archival), Amsterdam

详情

展开后加载摘要…

URL PDF HTML 收藏