arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

2026-03-05 至 2026-03-05 共收录 15 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. Agent评测 15 篇

2602.09937 2026-03-05 cs.AI cs.DC 83%

Why Do AI Agents Systematically Fail at Cloud Root Cause Analysis?

为何AI代理在云根因分析中系统性失败?

Taeyoon Kim, Woohyeok Park, Hoyeong Yun, Kyungyong Lee

机构 * Hanyang University(汉阳大学) OKESTRO Co., Ltd.(OKESTRO公司)

专题命中 Agent评测 :AI agent(title);agent(abstract);autonomous agent(abstract);分类 cs.AI

AI总结 本文分析了基于LLM的云根因分析代理在推理、通信和环境交互中的系统性故障,揭示了幻觉数据和不完全探索等普遍问题,并通过改进通信协议减少故障率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03640 2026-03-05 cs.RO 82%

MistyPilot: An Agentic Fast-Slow Thinking LLM Framework for Misty Social Robots

MistyPilot: 一种面向Misty社交机器人的代理快速-缓慢思考LLM框架

Xiao Wang, Lu Dong, Jingchen Sun, Ifeoma Nwogu, Srirangaraj Setlur, Venu Govindaraju

机构 * State University of New York at Buffalo(纽约州立大学布法罗分校)

专题命中 Agent评测 :agentic(title,abstract);agent(abstract)

AI总结 MistyPilot通过代理驱动的LLM框架,实现社交机器人中工具选择、协调及参数配置的高效执行,提升任务效率与情感一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03630 2026-03-05 cs.IR cs.MA 82%

Behind the Prompt: The Agent-User Problem in Information Retrieval

提示背后:信息检索中的代理-用户问题

Saber Zerhoudi, Michael Granitzer, Dang Hai Dang, Jelena Mitrovic, Florian Lemmerich, Annette Hautli-Janisz, Stefan Katzenbeisser, Kanishka Ghosh Dastidar

专题命中 Agent评测 :agent(title,abstract);AI agent(abstract)

AI总结 研究揭示了信息检索中代理-用户问题的结构性挑战,指出基于人类意图假设的模型在代理存在时可能失效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07757 2026-03-05 cs.AI cs.LG 81%

Emotion-Gradient Metacognitive RSI (Part I): Theoretical Foundations and Single-Agent Architecture

情绪梯度元认知递归自我改进(第一部分):理论基础与单体代理架构

Rintaro Ando

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI、cs.LG

AI总结 EG-MRSI框架整合反思元认知、情绪内在动机和递归自我修改,通过可微奖励函数和安全机制实现单体代理的理论基础,为安全AGI提供严谨基础。

Comments Withdrawn due to a critical error discovered in the stability and convergence proofs (specifically Lemma 2, Theorem 12, and Proposition 10) in Section 3. The identified flaw invalidates the core theoretical guarantees regarding capability growth and system stability

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03761 2026-03-05 cs.AI cs.IR 79%

AgentSelect: Benchmark for Narrative Query-to-Agent Recommendation

AgentSelect: 用于叙述查询到代理推荐的基准测试

Yunxiao Shi, Wujiang Xu, Tingwei Chen, Haoning Shang, Ling Yang, Yunfeng Wan, Zhuo Cao, Xing Zi, Dimitris N. Metaxas, Min Xu

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI

AI总结 AgentSelect是一个用于叙述查询到代理推荐的基准测试,通过统一的数据和评估基础设施,填补了代理推荐领域的研究空白,展示了基于能力匹配的代理推荐方法。

Comments under review by conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03762 2026-03-05 cs.CV 78%

Seeing as Experts Do: A Knowledge-Augmented Agent for Open-Set Fine-Grained Visual Understanding

如同专家所见:一个知识增强的代理用于开放集细粒度视觉理解

Junhan Chen, Zilu Zhou, Yujun Tong, Dongliang Chang, Yitao Luo, Zhanyu Ma

专题命中 Agent评测 :agent(title,abstract)

AI总结 KFRA通过知识增强推理代理实现开放集细粒度视觉理解,提升推理准确率和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03671 2026-03-05 q-fin.CP q-fin.TR 78%

Is an investor stolen their profits by mimic investors? Investigated by an agent-based model

投资者是否因模仿者而损失了利润?通过基于代理的模型进行调查

Takanobu Mizuta, Isao Yagi

专题命中 Agent评测 :agent(title,abstract)

AI总结 本文通过基于代理的模型研究了投资者因模仿者而损失利润的问题,发现增加基本策略代理使市场稳定从而减少利润,而增加技术策略代理使市场波动从而增加利润。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03881 2026-03-05 cs.CR cs.AI cs.CL cs.CY cs.HC 62%

On the Suitability of LLM-Driven Agents for Dark Pattern Audits

关于LLM驱动代理在暗模式审计中的适用性

Chen Sun, Yash Vekaria, Rishab Nithyanand

机构 * University of Iowa(爱荷华大学) University of California, Davis(加州大学戴维斯分校)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL

AI总结 本文研究了LLM驱动代理在暗模式审计中的适用性,通过分析456个数据经纪人网站,评估了代理在识别和分类暗模式方面的性能及局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10550 2026-03-05 cs.LG cs.AI cs.RO 62%

Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement Learning

记忆、基准与机器人:一种用于通过强化学习解决复杂任务的基准

Egor Cherepanov, Nikita Kachaev, Alexey K. Kovalev, Aleksandr I. Panov

机构 * AXXX MIRIAI ITMO University(ITMO大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

AI总结 本文提出MIKASA基准,旨在通过强化学习解决复杂任务,包含32个记忆密集型任务,用于评估智能体在桌面机器人操作中的记忆能力。

Comments 57 pages, 29 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03484 2026-03-05 cs.LG cs.AI 62%

Optimal trajectory-guided stochastic co-optimization for e-fuel system design and real-time operation

最优轨迹引导的随机共优化用于e-燃料系统设计和实时运行

Jeongdong Kim, Minsu Kim, Jonggeol Na, Junghwan Kim

机构 * Department of Chemical Engineering(化学工程系) Massachusetts Institute of Technology(麻省理工学院) Department of Chemical and Biomolecular Engineering(化学与生物分子工程系) Yonsei University(延世大学) Department of Chemical Engineering and Materials Science(化学工程与材料科学系) Ewha Womans University(成均馆大学) Institute for Multiscale Matter and Systems (IMMS)(多尺度物质与系统研究所)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

AI总结 MasCOR框架通过轨迹引导的随机共优化,实现e-燃料系统设计与实时运行的高效优化,适用于不同场景下的碳中性生产与成本控制。

Comments 29 pages, 6 figures. Supplementary Information included

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03045 2026-03-05 quant-ph cs.AI 57%

QFlowNet: Fast, Diverse, and Efficient Unitary Synthesis with Generative Flow Networks

QFlowNet: 快速、多样且高效的单位元合成与生成流网络

Inhoe Koo, Hyunho Cha, Jungwoo Lee

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) Seoul National University(首尔国立大学) NextQuantum

专题命中 Agent评测 :agent(abstract);分类 cs.AI

AI总结 QFlowNet通过结合生成流网络和Transformer,实现高效、多样化的单位元合成,解决了传统强化学习在稀疏奖励下的局限性。

Comments 7 pages, 6 figures, IEEE International Conference on Quantum Communications, Networking, and Computing (QCNC 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23124 2026-03-05 cs.CL 57%

Non-Collaborative User Simulators for Tool Agents

非协作用户模拟器用于工具代理

Jeonghoon Shim, Woojung Song, Cheyon Jin, Seungwon KooK, Yohan Jo

机构 * Graduate School of Data Science, Seoul National University(数据科学研究生院,首尔国立大学)

专题命中 Agent评测 :agent(abstract);分类 cs.CL

AI总结 本文提出一种非协作用户模拟器,用于测试工具代理在面对非协作用户时的鲁棒性,揭示了代理在复杂对话中的弱点。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06743 2026-03-05 cs.RO cs.AI 57%

TPK: Trustworthy Trajectory Prediction Integrating Prior Knowledge For Interpretability and Kinematic Feasibility

TPK: 通过整合先验知识实现可解释性和运动学可行性轨迹预测

Marius Baden, Ahmed Abouelazm, Christian Hubschneider, Yin Wu, Daniel Slieter, J. Marius Zöllner

机构 * FZI Research Center for Information Technology(弗赖堡信息科技研究中心) Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) CARIAD SE

专题命中 Agent评测 :agent(abstract);分类 cs.AI

AI总结 TPK通过整合车辆、行人和自行车的交互和运动学先验知识,提高轨迹预测的可解释性和物理可行性。

Comments First and Second authors contributed equally; Accepted in the 36th IEEE Intelligent Vehicles Symposium (IV 2025) for oral presentation; Winner of the best paper award

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10131 2026-03-05 cs.LO 50%

Proof Strategy Extraction from LLMs for Enhancing Symbolic Provers

从LLMs中提取证明策略以增强符号证明器

Jian Fang, Yican Sun, Yingfei Xiong

专题命中 Agent评测 :agent(abstract)

AI总结 本文提出Strat2Rocq方法,通过提取LLM的证明策略来增强符号证明器的能力,提升自动化证明效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14516 2026-03-05 cs.RO 50%

Event-LAB: Towards Standardized Evaluation of Neuromorphic Localization Methods

Event-LAB: 向神经形态定位方法的标准化评估迈进

Adam D. Hines, Alejandro Fontan, Michael Milford, Tobias Fischer

机构 * QUT Centre for Robotics, School of Electrical Engineering and Robotics, Queensland University of Technology(昆士兰理工大学机器人中心,电气与机器人工程学院,昆士兰理工大学) School of Natural Sciences, Macquarie University(自然科学院,麦觉理大学)

专题命中 Agent评测 :workflow(abstract)

AI总结 Event-LAB提供了一个统一框架,用于在多个数据集上运行多种基于事件的定位方法,通过一致的参数设置实现公平比较。

Comments 8 pages, 6 figures, accepted to the IEEE International Conference on Robotics and Automation (ICRA) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏