arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

2026-03-05 至 2026-03-05 共收录 114 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. 记忆与上下文管理 11 篇

2603.03290 2026-03-05 cs.CL cs.AI cs.IR cs.LG 67%

AriadneMem: Threading the Maze of Lifelong Memory for LLM Agents

AriadneMem: 为LLM代理梳理终身记忆的迷宫

Wenhui Zhu, Xiwen Chen, Zhipeng Wang, Jingjing Wang, Xuanzhao Dong, Minzhou Huang, Rui Cai, Hejian Sang, Hao Wang, Peijie Qiu, Yueyue Deng, Prayag Tiwari, Brendan Hogan Rappazzo, Yalin Wang

机构 * Arizona State University(亚利桑那州立大学) Morgan Stanley(摩根大通) Rice University(里奇大学) Clemson University(克莱姆森大学) Northwestern University(西北大学) UC Davis(加州大学戴维斯分校) Iowa State University(爱荷华州立大学) Washington University in St. Louis(圣路易斯华盛顿大学) Columbia University(哥伦比亚大学) Halmstad University(哈马格大学)

专题命中 记忆与上下文管理 :planning(abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 AriadneMem通过解耦的两阶段流程提升LLM代理在长期对话中的多跳和平均F1表现,同时减少运行时间。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06531 2026-03-05 cs.LG cs.AI 62%

Unraveling the Complexity of Memory in RL Agents: an Approach for Classification and Evaluation

揭示强化学习智能体记忆复杂性:一种分类与评估的方法

Egor Cherepanov, Nikita Kachaev, Artem Zholus, Alexey K. Kovalev, Aleksandr I. Panov

机构 * AXXX MIRIAI ITMO University AI Talent Hub(ITMO大学AI人才中心) Mila – Quebec AI Institute(魁北克AI研究所) Polytechnique Montréal(蒙特利尔理工学院)

专题命中 记忆与上下文管理 :agent(abstract);分类 cs.AI、cs.LG

AI总结 本文提出通过定义智能体记忆类型(如长期vs短期、陈述性vs程序性记忆)来简化强化学习中的记忆概念,并通过实验方法评估不同智能体的记忆能力。

Comments 20 pages, 6 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.09459 2026-03-05 cs.LG cs.AI 62%

Recurrent Action Transformer with Memory

具有记忆的递归动作变换器

Egor Cherepanov, Alexey Staroverov, Alexey K. Kovalev, Aleksandr I. Panov

专题命中 记忆与上下文管理 :agent(abstract);分类 cs.AI、cs.LG

AI总结 本文提出具有记忆的递归动作变换器(RATE),通过引入递归记忆机制提升离线强化学习中对长序列和内存密集型任务的处理能力。

Comments 29 pages, 22 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03296 2026-03-05 cs.CL cs.AI cs.IR 62%

PlugMem: A Task-Agnostic Plugin Memory Module for LLM Agents

PlugMem: 一种通用插件记忆模块用于LLM代理

Ke Yang, Zixi Chen, Xuan He, Jize Jiang, Michel Galley, Chenglong Wang, Jianfeng Gao, Jiawei Han, ChengXiang Zhai

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Tsinghua University(清华大学) Microsoft Research(微软研究院)

专题命中 记忆与上下文管理 :agent(abstract);分类 cs.AI、cs.CL

AI总结 PlugMem是一种通用记忆模块,通过知识中心的记忆图结构提升LLM代理在复杂任务中的记忆检索与推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03781 2026-03-05 cs.AI 57%

LifeBench: A Benchmark for Long-Horizon Multi-Source Memory

LifeBench: 一个用于长周期多源记忆的基准

Zihao Cheng, Weixin Wang, Yu Zhao, Ziyang Ren, Jiaxuan Chen, Ruiyang Xu, Shuai Huang, Yang Chen, Guowei Li, Mengshi Wang, Yi Xie, Ren Zhu, Zeren Jiang, Keda Lu, Yihong Li, Xiaoliang Wang, Liwei Liu, Cam-Tu Nguyen

机构 * Institute for Clarity in Documentation(清晰文档研究所) Inria Paris-Rocquencourt(巴黎-罗克琴库特研究所) Rajiv Gandhi University(拉贾夫·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕勒尔研究实验室) State Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学) School of Artificial Intelligence, Nanjing University(人工智能学院,南京大学) Huawei Technologies Co., Ltd.(华为技术有限公司)

专题命中 记忆与上下文管理 :AI agent(abstract);分类 cs.AI

AI总结 Lifebench是一个用于长周期多源记忆的基准,通过密集连接的长周期事件模拟,推动AI代理在多样且时间延长的上下文中整合声明性和非声明性记忆推理。

Comments A total of 28 pages, 8 pages of main text, and 15 figures and tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03633 2026-03-05 cs.CR cs.AI 57%

Goal-Driven Risk Assessment for LLM-Powered Systems: A Healthcare Case Study

面向目标的风险评估用于LLM驱动系统:一个医疗案例研究

Neha Nagaraja, Hayretdin Bahsi

机构 * School of Informatics, Computing, and Cyber Systems(信息学、计算与网络系统学院) Northern Arizona University(北亚利桑那大学) Department of Software Science(软件科学系) Tallinn University of Technology(塔林技术大学)

专题命中 记忆与上下文管理 :agent(abstract);分类 cs.AI

AI总结 本文提出了一种面向目标的风险评估方法,用于评估LLM驱动系统的安全风险,通过攻击树详细分析威胁向量和攻击路径,以提升医疗系统等复杂系统的安全性。

Comments To appear in the HealthSec Workshop at the 2025 IEEE Annual Computer Security Applications Conference (ACSAC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02862 2026-03-05 cs.LG 57%

Learning in Markov Decision Processes with Exogenous Dynamics

在外生动态中学习马尔可夫决策过程

Davide Maran, Davide Salaorni, Marcello Restelli

机构 * Department of DEIB, Politecnico di Milano, Milan, Italy(米兰理工大学DEIB部门)

专题命中 记忆与上下文管理 :agent(abstract);分类 cs.LG

AI总结 本文提出了一种在具有外生动态的马尔可夫决策过程中学习的方法,通过利用外生状态组件的独立性,提高了学习保证和样本效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04254 2026-03-05 cs.CV 50%

EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding

EmbodiedSplat: 在线前馈语义3DGS用于开放词汇3D场景理解

Seungjun Lee, Zihan Wang, Yunsong Wang, Gim Hee Lee

机构 * National University of Singapore(新加坡国立大学)

专题命中 记忆与上下文管理 :agent(abstract)

AI总结 EmbodiedSplat通过在线前馈3DGS实现开放词汇3D场景理解,结合CLIP编码和3D U-Net提升语义重建效率与泛化能力。

Comments CVPR 2026, Project Page: https://0nandon.github.io/EmbodiedSplat/

详情

展开后加载摘要…

URL PDF HTML 收藏

2. Agent评测 15 篇

2602.09937 2026-03-05 cs.AI cs.DC 83%

Why Do AI Agents Systematically Fail at Cloud Root Cause Analysis?

为何AI代理在云根因分析中系统性失败?

Taeyoon Kim, Woohyeok Park, Hoyeong Yun, Kyungyong Lee

机构 * Hanyang University(汉阳大学) OKESTRO Co., Ltd.(OKESTRO公司)

专题命中 Agent评测 :AI agent(title);agent(abstract);autonomous agent(abstract);分类 cs.AI

AI总结 本文分析了基于LLM的云根因分析代理在推理、通信和环境交互中的系统性故障,揭示了幻觉数据和不完全探索等普遍问题,并通过改进通信协议减少故障率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03640 2026-03-05 cs.RO 82%

MistyPilot: An Agentic Fast-Slow Thinking LLM Framework for Misty Social Robots

MistyPilot: 一种面向Misty社交机器人的代理快速-缓慢思考LLM框架

Xiao Wang, Lu Dong, Jingchen Sun, Ifeoma Nwogu, Srirangaraj Setlur, Venu Govindaraju

机构 * State University of New York at Buffalo(纽约州立大学布法罗分校)

专题命中 Agent评测 :agentic(title,abstract);agent(abstract)

AI总结 MistyPilot通过代理驱动的LLM框架,实现社交机器人中工具选择、协调及参数配置的高效执行,提升任务效率与情感一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03630 2026-03-05 cs.IR cs.MA 82%

Behind the Prompt: The Agent-User Problem in Information Retrieval

提示背后:信息检索中的代理-用户问题

Saber Zerhoudi, Michael Granitzer, Dang Hai Dang, Jelena Mitrovic, Florian Lemmerich, Annette Hautli-Janisz, Stefan Katzenbeisser, Kanishka Ghosh Dastidar

专题命中 Agent评测 :agent(title,abstract);AI agent(abstract)

AI总结 研究揭示了信息检索中代理-用户问题的结构性挑战,指出基于人类意图假设的模型在代理存在时可能失效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07757 2026-03-05 cs.AI cs.LG 81%

Emotion-Gradient Metacognitive RSI (Part I): Theoretical Foundations and Single-Agent Architecture

情绪梯度元认知递归自我改进(第一部分):理论基础与单体代理架构

Rintaro Ando

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI、cs.LG

AI总结 EG-MRSI框架整合反思元认知、情绪内在动机和递归自我修改,通过可微奖励函数和安全机制实现单体代理的理论基础,为安全AGI提供严谨基础。

Comments Withdrawn due to a critical error discovered in the stability and convergence proofs (specifically Lemma 2, Theorem 12, and Proposition 10) in Section 3. The identified flaw invalidates the core theoretical guarantees regarding capability growth and system stability

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03761 2026-03-05 cs.AI cs.IR 79%

AgentSelect: Benchmark for Narrative Query-to-Agent Recommendation

AgentSelect: 用于叙述查询到代理推荐的基准测试

Yunxiao Shi, Wujiang Xu, Tingwei Chen, Haoning Shang, Ling Yang, Yunfeng Wan, Zhuo Cao, Xing Zi, Dimitris N. Metaxas, Min Xu

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI

AI总结 AgentSelect是一个用于叙述查询到代理推荐的基准测试,通过统一的数据和评估基础设施,填补了代理推荐领域的研究空白,展示了基于能力匹配的代理推荐方法。

Comments under review by conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03762 2026-03-05 cs.CV 78%

Seeing as Experts Do: A Knowledge-Augmented Agent for Open-Set Fine-Grained Visual Understanding

如同专家所见:一个知识增强的代理用于开放集细粒度视觉理解

Junhan Chen, Zilu Zhou, Yujun Tong, Dongliang Chang, Yitao Luo, Zhanyu Ma

专题命中 Agent评测 :agent(title,abstract)

AI总结 KFRA通过知识增强推理代理实现开放集细粒度视觉理解,提升推理准确率和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03671 2026-03-05 q-fin.CP q-fin.TR 78%

Is an investor stolen their profits by mimic investors? Investigated by an agent-based model

投资者是否因模仿者而损失了利润?通过基于代理的模型进行调查

Takanobu Mizuta, Isao Yagi

专题命中 Agent评测 :agent(title,abstract)

AI总结 本文通过基于代理的模型研究了投资者因模仿者而损失利润的问题,发现增加基本策略代理使市场稳定从而减少利润,而增加技术策略代理使市场波动从而增加利润。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03881 2026-03-05 cs.CR cs.AI cs.CL cs.CY cs.HC 62%

On the Suitability of LLM-Driven Agents for Dark Pattern Audits

关于LLM驱动代理在暗模式审计中的适用性

Chen Sun, Yash Vekaria, Rishab Nithyanand

机构 * University of Iowa(爱荷华大学) University of California, Davis(加州大学戴维斯分校)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL

AI总结 本文研究了LLM驱动代理在暗模式审计中的适用性,通过分析456个数据经纪人网站,评估了代理在识别和分类暗模式方面的性能及局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10550 2026-03-05 cs.LG cs.AI cs.RO 62%

Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement Learning

记忆、基准与机器人:一种用于通过强化学习解决复杂任务的基准

Egor Cherepanov, Nikita Kachaev, Alexey K. Kovalev, Aleksandr I. Panov

机构 * AXXX MIRIAI ITMO University(ITMO大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

AI总结 本文提出MIKASA基准,旨在通过强化学习解决复杂任务,包含32个记忆密集型任务,用于评估智能体在桌面机器人操作中的记忆能力。

Comments 57 pages, 29 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03484 2026-03-05 cs.LG cs.AI 62%

Optimal trajectory-guided stochastic co-optimization for e-fuel system design and real-time operation

最优轨迹引导的随机共优化用于e-燃料系统设计和实时运行

Jeongdong Kim, Minsu Kim, Jonggeol Na, Junghwan Kim

机构 * Department of Chemical Engineering(化学工程系) Massachusetts Institute of Technology(麻省理工学院) Department of Chemical and Biomolecular Engineering(化学与生物分子工程系) Yonsei University(延世大学) Department of Chemical Engineering and Materials Science(化学工程与材料科学系) Ewha Womans University(成均馆大学) Institute for Multiscale Matter and Systems (IMMS)(多尺度物质与系统研究所)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

AI总结 MasCOR框架通过轨迹引导的随机共优化,实现e-燃料系统设计与实时运行的高效优化,适用于不同场景下的碳中性生产与成本控制。

Comments 29 pages, 6 figures. Supplementary Information included

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03045 2026-03-05 quant-ph cs.AI 57%

QFlowNet: Fast, Diverse, and Efficient Unitary Synthesis with Generative Flow Networks

QFlowNet: 快速、多样且高效的单位元合成与生成流网络

Inhoe Koo, Hyunho Cha, Jungwoo Lee

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) Seoul National University(首尔国立大学) NextQuantum

专题命中 Agent评测 :agent(abstract);分类 cs.AI

AI总结 QFlowNet通过结合生成流网络和Transformer,实现高效、多样化的单位元合成,解决了传统强化学习在稀疏奖励下的局限性。

Comments 7 pages, 6 figures, IEEE International Conference on Quantum Communications, Networking, and Computing (QCNC 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23124 2026-03-05 cs.CL 57%

Non-Collaborative User Simulators for Tool Agents

非协作用户模拟器用于工具代理

Jeonghoon Shim, Woojung Song, Cheyon Jin, Seungwon KooK, Yohan Jo

机构 * Graduate School of Data Science, Seoul National University(数据科学研究生院,首尔国立大学)

专题命中 Agent评测 :agent(abstract);分类 cs.CL

AI总结 本文提出一种非协作用户模拟器,用于测试工具代理在面对非协作用户时的鲁棒性,揭示了代理在复杂对话中的弱点。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06743 2026-03-05 cs.RO cs.AI 57%

TPK: Trustworthy Trajectory Prediction Integrating Prior Knowledge For Interpretability and Kinematic Feasibility

TPK: 通过整合先验知识实现可解释性和运动学可行性轨迹预测

Marius Baden, Ahmed Abouelazm, Christian Hubschneider, Yin Wu, Daniel Slieter, J. Marius Zöllner

机构 * FZI Research Center for Information Technology(弗赖堡信息科技研究中心) Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) CARIAD SE

专题命中 Agent评测 :agent(abstract);分类 cs.AI

AI总结 TPK通过整合车辆、行人和自行车的交互和运动学先验知识,提高轨迹预测的可解释性和物理可行性。

Comments First and Second authors contributed equally; Accepted in the 36th IEEE Intelligent Vehicles Symposium (IV 2025) for oral presentation; Winner of the best paper award

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10131 2026-03-05 cs.LO 50%

Proof Strategy Extraction from LLMs for Enhancing Symbolic Provers

从LLMs中提取证明策略以增强符号证明器

Jian Fang, Yican Sun, Yingfei Xiong

专题命中 Agent评测 :agent(abstract)

AI总结 本文提出Strat2Rocq方法,通过提取LLM的证明策略来增强符号证明器的能力,提升自动化证明效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14516 2026-03-05 cs.RO 50%

Event-LAB: Towards Standardized Evaluation of Neuromorphic Localization Methods

Event-LAB: 向神经形态定位方法的标准化评估迈进

Adam D. Hines, Alejandro Fontan, Michael Milford, Tobias Fischer

机构 * QUT Centre for Robotics, School of Electrical Engineering and Robotics, Queensland University of Technology(昆士兰理工大学机器人中心,电气与机器人工程学院,昆士兰理工大学) School of Natural Sciences, Macquarie University(自然科学院,麦觉理大学)

专题命中 Agent评测 :workflow(abstract)

AI总结 Event-LAB提供了一个统一框架,用于在多个数据集上运行多种基于事件的定位方法,通过一致的参数设置实现公平比较。

Comments 8 pages, 6 figures, accepted to the IEEE International Conference on Robotics and Automation (ICRA) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他Agent 1 篇

2507.15796 2026-03-05 cs.AI 80%

From Privacy to Trust in the Agentic Era: A Taxonomy of Challenges in Trustworthy Federated Learning Through the Lens of Trust Report 2.0

从隐私到信任在代理时代:通过信任报告2.0的视角,对可信联邦学习中挑战的分类

Nuria Rodríguez-Barroso, Mario García-Márquez, M. Victoria Luzón, Francisco Herrera

机构 * Department of Computer Science and Artificial Intelligence, Andalusian Research Institute in Data Science and Computational Intelligence (DaSCI) University of Granada(计算机科学与人工智能系,数据科学与计算智能安达卢西亚研究 institute,格拉纳达大学) Department of Software Engineering, Andalusian Research Institute in Data Science and Computational Intelligence (DaSCI) University of Granada(软件工程系,数据科学与计算智能安达卢西亚研究 institute,格拉纳达大学)

专题命中 其他Agent :agentic(title,abstract);分类 cs.AI

AI总结 本文通过信任报告2.0提出可信联邦学习的挑战分类,强调信任作为持续维持的操作条件,并引入协调蓝图以处理跨要求的权衡和治理对齐。

Comments Already published in Information Fusion

Journal ref Rodríguez-Barroso, et. al. (2026). From Privacy to Trust in the Agentic Era: A Taxonomy of Challenges in Trustworthy Federated Learning Through the Lens of Trust Report 2.0. Information Fusion, 104236

详情

展开后加载摘要…

URL PDF HTML 收藏