arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

共收录 15819 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. Agent评测 15819 篇

2607.21635 2026-07-27 cs.LG 新提交 61%

Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions

面向时间干预下个人语言模型代理的用户条件评估

Pin Qian, Su Wang, Yihang Chen, Qiaolin Yu, Xiaoyuan Wang, Zhitong Guo, Zhicheng Wang, Junxian You

机构 * Carnegie Mellon University(卡内基梅隆大学) Georgia Institute of Technology(佐治亚理工学院) Cornell University(康奈尔大学) University of Glasgow(格拉斯哥大学)

专题命中 Agent评测 :agent(abstract);分类 cs.LG;agentic(comments)

AI总结 研究个人语言模型代理在时间干预下的用户条件评估,提出需在不同用户条件状态下重放时间干预并测量故障传播的协议,形式化四个条件,审查发现无满足条件的公开协议,进而提出最小基准设计和候选报告指标。

Comments 9 pages, 2 figures, and 8 tables. Accepted for oral presentation at the ACM SIGKDD KDD 2026 Workshop on Personal Intelligence in the Agentic AI Era (PILA 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18242 2026-07-22 cs.AI cs.MA cs.NI 新提交 61%

AI Tool Discovery at Scale: All You Need is DNS

大规模人工智能工具发现:你所需要的只是DNS

Enhao Chen, Yulin Shao

机构 * The University of Hong Kong(香港大学)

专题命中 Agent评测 :AI agent(abstract);分类 cs.AI;agent(comments)

AI总结 针对自主人工智能代理时代工具发现难题,提出ToolDNS框架,通过嵌入意图和信任将语义搜索转为轻量级名称解析,引入三项增强。经大规模基准测试验证,其在减少搜索空间和降低延迟方面效果显著,证明可通过利用现有基础设施实现可扩展的AI互操作性。

Comments keywords: AI tool discovery, ToolDNS, Agent, DNS

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26076 2026-07-10 cs.CL 版本更新 61%

IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation

IMProofBench:在研究级数学证明生成上对人工智能进行基准测试

Johannes Schmitt, Gergely Bérczi, Jasper Dekoninck, Jeremy Feusi, Tim Gehrunger, Raphael Appenzeller, Pieter Belmans, Alessio Bottini, Jim Bryan, João Camarneiro, Ana Cannas da Silva, Niklas Canova, Ana-Maria Castravet, Timo de Wolff, Claudio Fontanari, Filippo Gaia, Baran Hashemi, Daniel Holmes, David Holmes, Aitor Iribar Lopez, Victor Jaeck, Martina Jørgensen, Steven Kelk, Martijn Kool, Stefan Kuhlmann, Adam Kurpisz, Johannes Lengler, Chiara Meroni, Ingmar Metzler, Martin Möller, Samuel Muñoz-Echániz, David Muñoz-Lahoz, Robert Nowak, Georg Oberdieck, Daniel Platt, Dylan Possamaï, Gabriel Ribeiro, Aluna Rizzoli, Daria Sakhanda, Raúl Sánchez Galán, Zheming Sun, Diaaeldin Taha, Josef Teichmann, Richard P. Thomas, Henk van der Pol, Michel van Garrel, Charles Vial, Ignacio Barros, Benjamin Doerr, Peter Grünwald, Henry Liu, David Martins, Aleksandar Mijatović, Sergej Monavari, Marc Roth, Patrick Schnider, Yannik Schuler, Pim Spelier, Yuuji Tanaka, Ronald van Luijk

机构 * ETH Zurich(苏黎世联邦理工学院) Aarhus University(奥胡斯大学)

专题命中 Agent评测 :agentic(abstract,comments);分类 cs.CL

AI总结 针对大语言模型数学能力评估,引入IMProofBench基准测试,由专家开发77个经同行评审问题,涵盖详细证明及子问题,模拟现实研究环境,结果显示当前LLMs能解决不少研究级问题,该基准测试将持续发展。

Comments v2: benchmark expanded from 39 to 77 problems; evaluation extended to 14 models including GPT-5.4, Gemini 3.1 Pro, and Claude Opus 4.6; new analyses (IRT-based score aggregation, inter-rater reliability, tool/token usage, non-agentic ablation); contributor author list updated

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26539 2026-05-27 cs.SE cs.CR 61%

FuzzPilot: Plateau-Triggered Recipe Validation for Structured Text Fuzzing

FuzzPilot: 用于结构化文本模糊测试的高原触发式配方验证

Zhiyi Yao

专题命中 Agent评测 :agent(abstract,comments);分类 cs.SE

AI总结 FuzzPilot 通过高原触发式配方验证,在覆盖率停滞时利用快照和微测试评估变异配方,以提升 AFL++ 的模糊测试效率。

Comments 41 pages, 7 figures, 7 tables. Preliminary cJSON-only evaluation (N=5 main, N=3 ablation; descriptive statistics, no significance claims). Code and 25-run artifacts at https://github.com/Qiao-Zhiyi/fuzz_agent (tag paper01-arxiv-v1). Venue-version Stage-1 pilot on libxml2, sqlite3, openssl_x509 currently in flight; v2 will report those results

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27819 2026-05-01 cs.AI 61%

MCPHunt: An Evaluation Framework for Cross-Boundary Data Propagation in Multi-Server MCP Agents

MCPHunt:多服务器MCP代理跨边界数据传播的评估框架

Haonan Li, Tianjun Sun, Yongqing Wang, Qisheng Zhang

机构 * Key Laboratory of Intraplate Volcanoes and Earthquakes (China University of Geosciences, Beijing), Ministry of Education, Beijing 100083, China(中国大陆地质大学(北京)构造运动与地震重点实验室,教育部,北京100083,中国) School of Geophysics and Information Technology, China University of Geosciences, Beijing 100083, China(中国大陆地质大学(北京)地球物理与信息技术学院,北京100083,中国)

专题命中 Agent评测 :workflow(abstract);分类 cs.AI;agent(comments)

AI总结 本文提出MCPHunt框架,通过三种方法论贡献评估多服务器MCP代理中非对抗性跨边界凭证传播问题,发现政策违规传播率高达11.5%-41.3%,并验证了提示引导的数据流足以避免凭证传播。

Comments 21 pages, 1 figure, 16 tables. Code: https://github.com/lihaonan0716/MCPHunt Data: https://huggingface.co/datasets/lihaonan0716/mcphunt-agent-traces

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24191 2026-03-02 cs.GT cs.AI cs.LO 61%

Resilient Strategies for Stochastic Systems: How Much Does It Take to Break a Winning Strategy?

随机系统中的稳健策略:打破获胜策略需要多大的代价?

Kush Grover, Markel Zubia, Debraj Chakraborty, Muqsit Azeem, Nils Jansen, Jan Kretinsky

机构 * Ruhr University Bochum Germany Nanyang Technological University, Singapore Technical University of Munich \& University of Konstanz Germany Ruhr University Bochum \& Radboud University Nijmegen Germany Masaryk University Czech Republic Ruhr University Bochum Technical University of Munich \& University of Konstanz Ruhr University Bochum \& Radboud University Nijmegen Masaryk University

专题命中 Agent评测 :agent(abstract);分类 cs.AI;autonomous agent(comments)

AI总结 本文研究了随机系统中稳健策略的定义与问题,探讨了马尔可夫决策过程和随机游戏中的可达性和安全目标,提出了一种基于干扰频率的稳健性度量方法。

Comments To appear in Proc. of the 25th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2026), Paphos, Cyprus, May 25-29, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07243 2026-02-10 cs.RO cs.AI cs.GR 61%

Realistic Synthetic Household Data Generation at Scale

大规模生成逼真合成家庭数据

Siddharth Singh, Ifrah Idrees, Abraham Dauhajre

机构 * Siddharth Singh(1 西雅图)

专题命中 Agent评测 :agent(abstract);分类 cs.AI;agentic(comments)

AI总结 本文提出了一种大规模生成逼真合成家庭数据的方法,通过松耦合生成人机交互与环境数据,实现双向影响建模,提升家庭智能设备的开发与测试能力。

Comments Accepted at Agentic AI Benchmarks and Applications for Enterprise Tasks workshop at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01047 2026-01-21 cs.NI cs.AI 61%

A Deep Reinforcement Learning-Based TCP Congestion Control Algorithm: Design, Simulation, and Evaluation

基于深度强化学习的TCP拥塞控制算法:设计、仿真与评估

Efe Ağlamazlar, Emirhan Eken, Harun Batur Geçici

机构 * Aydın Science High School(艾丁科学高中) ASELSAN Technical Anatolian High School(ASELSAN技术安纳托利亚高中) Erman Ilıcak Science High School(埃尔曼·伊利卡克科学高中)

专题命中 Agent评测 :agent(abstract,comments);分类 cs.AI

AI总结 本文提出一种基于深度强化学习的TCP拥塞控制算法,通过动态调整拥塞窗口提升吞吐量与降低延迟,优于传统TCP算法。

Comments 11 pages, 4 figures. Presents a DRL agent that mitigates bufferbloat and achieves near-zero packet loss. Validated via NS-3 simulations under a strict training-testing protocol. Code: https://github.com/aglamazlarefe/DRL-TCP

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24695 2025-10-29 cs.CL 61%

AgentFrontier: Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesis

Xuanzhong Chen, Zile Qiao, Guoxin Chen, Liangcai Su, Zhen Zhang, Xinyu Wang, Pengjun Xie, Fei Huang, Jingren Zhou, Yong Jiang

机构 * Tongyi Lab(通义实验室) Alibaba Group(阿里巴巴集团)

专题命中 Agent评测 :agent(abstract,comments);分类 cs.CL

Comments https://tongyi-agent.github.io/blog/introducing-tongyi-deep-research/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13557 2025-10-16 cs.CV cs.AI 61%

Modeling Cultural Bias in Facial Expression Recognition with Adaptive Agents

David Freire-Obregón, José Salas-Cáceres, Javier Lorenzo-Navarro, Oliverio J. Santana, Daniel Hernández-Sosa, Modesto Castrillón-Santana

专题命中 Agent评测 :agent(abstract);分类 cs.AI;agentic(comments)

Comments Accepted for presentation at the International Symposium on Agentic Artificial Intelligence Systems (AAIS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05329 2025-07-04 cs.LG cs.RO 61%

Knowledge Transfer in Model-Based Reinforcement Learning Agents for Efficient Multi-Task Learning

Dmytro Kuzmenko, Nadiya Shvai

机构 * National University of Kyiv-Mohyla Academy(基辅莫希拉大学)

专题命中 Agent评测 :agent(abstract);分类 cs.LG;autonomous agent(journal_ref)

Comments Preprint of an extended abstract accepted to AAMAS 2025

Journal ref Proceedings of the 24th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2025), pp. 2597-2599, ACM, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05891 2025-03-14 cs.CL 61%

MastermindEval: A Simple But Scalable Reasoning Benchmark

Jonas Golde, Patrick Haller, Fabio Barth, Alan Akbik

专题命中 Agent评测 :agentic(abstract);分类 cs.CL;planning(comments)

Comments 9 pages, 2 figures, 4 tables. In: ICLR 2025 Workshop on Reasoning and Planning for Large Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13930 2024-07-16 cs.LG cs.SY eess.SY 61%

Enhancing Reinforcement Learning Agents with Local Guides

Paul Daoudi, Bogdan Robu, Christophe Prieur, Ludovic Dos Santos, Merwan Barlier

专题命中 Agent评测 :agent(abstract);分类 cs.LG;autonomous agent(journal_ref)

Journal ref AAMAS '23: Proceedings of the 2023 International Conference on Autonomous Agents and Multiagent Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.02637 2024-05-28 cs.AI cs.MA cs.RO 61%

Surge Routing: Event-informed Multiagent Reinforcement Learning for Autonomous Rideshare

Daniel Garces, Stephanie Gil

专题命中 Agent评测 :agent(abstract);分类 cs.AI;autonomous agent(comments)

Comments 10 pages, 7 figures, 4 tables, 23rd International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.00003 2023-01-03 cs.HC cs.AI 61%

Emotion in Cognitive Architecture: Emergent Properties from Interactions with Human Emotion

Junya Morita

专题命中 Agent评测 :agent(abstract,comments);分类 cs.AI

Comments Presented at HAI'22 workshop on Cognitive Human-agent Interaction (https://sites.google.com/view/chai-workshop22)

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.05762 2022-01-17 cs.HC cs.LG cs.MM 61%

Multimodal analysis of the predictability of hand-gesture properties

Taras Kucherenko, Rajmund Nagy, Michael Neff, Hedvig Kjellström, Gustav Eje Henter

专题命中 Agent评测 :agent(abstract);分类 cs.LG;autonomous agent(comments)

Comments Accepted at the International Conference on Autonomous Agents and Multiagent Systems (AAMAS) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.10502 2021-08-18 cs.AI cs.GT econ.TH 61%

Weighted Envy-Freeness in Indivisible Item Allocation

Mithun Chakraborty, Ayumi Igarashi, Warut Suksompong, Yair Zick

专题命中 Agent评测 :agent(abstract);分类 cs.AI;autonomous agent(comments)

Comments A preliminary version appears in Proceedings of the 19th International Conference on Autonomous Agents and Multiagent Systems (AAMAS), 2020

Journal ref ACM Transactions on Economics and Computation, 9(3):18 (2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.12302 2021-02-25 cs.HC cs.GR cs.LG 61%

A Framework for Integrating Gesture Generation Models into Interactive Conversational Agents

Rajmund Nagy, Taras Kucherenko, Birger Moell, André Pereira, Hedvig Kjellström, Ulysses Bernardet

专题命中 Agent评测 :agent(abstract);分类 cs.LG;autonomous agent(comments)

Comments Rajmund Nagy and Taras Kucherenko contributed equally to this work. To be published in the Proceedings of the 20th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2021), Online, May 3-7, 2021, IFAA-MAS, 3 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.09888 2020-10-26 cs.CV cs.HC cs.LG cs.SD eess.AS eess.IV stat.ML 61%

Let's Face It: Probabilistic Multi-modal Interlocutor-aware Generation of Facial Gestures in Dyadic Settings

Patrik Jonell, Taras Kucherenko, Gustav Eje Henter, Jonas Beskow

专题命中 Agent评测 :agent(abstract,comments);分类 cs.LG

Comments Best Paper Award. 8 pages, 4 figures, IVA '20: Proceedings of the 20th ACM International Conference on Intelligent Virtual Agent

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.07745 2020-03-18 cs.AI cs.RO 61%

Learning to Optimize Autonomy in Competence-Aware Systems

Connor Basich, Justin Svegliato, Kyle Hollins Wray, Stefan Witwicki, Joydeep Biswas, Shlomo Zilberstein

专题命中 Agent评测 :agent(abstract);分类 cs.AI;autonomous agent(comments)

Comments To be published in Proceedings of the 19th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2020). 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1606.01467 2017-07-14 cs.LG cs.NE 61%

Deep Q-Networks for Accelerating the Training of Deep Neural Networks

Jie Fu

专题命中 Agent评测 :agent(abstract,comments);分类 cs.LG

Comments We choose to withdraw this paper. The DQN itself has too many hyperparameters, which makes it almost impossible to be applied to reasonably large datasets. In the later versions (from v4) with SGDR experiments, it seems that the agent only performs random actions

详情

展开后加载摘要…

URL PDF HTML 收藏
0903.1137 2012-04-18 cs.AI cs.CC cs.MA 61%

Complexity of Terminating Preference Elicitation

Toby Walsh

专题命中 Agent评测 :agent(abstract);分类 cs.AI;autonomous agent(comments)

Comments 7th International Joint Conference on Autonomous Agents and Multiagent Systems (AAMAS 2008)

Journal ref AAMAS 2008: 967-974

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11444 2026-03-13 cs.CY cs.HC cs.MA 60%

EducaSim: Interactive Simulacra for CS1 Instructional Practice

EducaSim:用于CS1教学实践的交互仿像

Cameron Mohne, Nicholas Vo, Dora Demszky, Chris Piech

专题命中 Agent评测 :agent(abstract,comments);multi-agent(comments)

AI总结 EducaSim通过生成代理为教师培训提供交互仿像,实现低成本大规模教学实践。

Comments 7 pages, 3 figures, 2 tables. Presents a multi-agent generative architecture for educational simulations intended for instructor training

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22814 2025-09-30 cs.CR 60%

Model Context Protocol for Vision Systems: Audit, Security, and Protocol Extensions

Aditi Tiwari, Akshit Bhalla, Darshan Prasad

专题命中 Agent评测 :agent(abstract,comments);planning(comments)

Comments Accepted to NeurIPS 2025 Workshop on Bridging Language, Agent, and World Models for Reasoning and Planning (LAW 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00072 2025-06-03 cs.CY cs.AI cs.CL cs.LG 60%

Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs

Nariman Naderi, Zahra Atf, Peter R Lewis, Aref Mahjoub far, Seyed Amir Ahmad Safavi-Naini, Ali Soroush

专题命中 Agent评测 :分类 cs.AI、cs.CL、cs.LG;agent(comments);multi-agent(comments)

Comments This paper was accepted for presentation at the 7th International Workshop on EXplainable, Trustworthy, and Responsible AI and Multi-Agent Systems (EXTRAAMAS 2025). Workshop website: https://extraamas.ehealth.hevs.ch/index.html

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.09880 2024-02-20 cs.GT 60%

Online Revenue Maximization for Server Pricing

Shant Boodaghians, Federico Fusco, Stefano Leonardi, Yishay Mansour, Ruta Mehta

专题命中 Agent评测 :agent(abstract,journal_ref);multi-agent(journal_ref)

Journal ref Auton Agent Multi-Agent Syst 36, 11 (2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.03253 2022-11-15 cs.MA 60%

About Digital Twins, agents, and multiagent systems: a cross-fertilisation journey

Stefano Mariani, Marco Picone, Alessandro Ricci

专题命中 Agent评测 :agent(abstract,comments);autonomous agent(journal_ref)

Comments Digital Twin, Agent, Multiagent system, Cyber-physical system;16 pages; accepted and presented at EMAS2022 (workshop of AAMAS: accepted" target="_blank" rel="noopener">https://emas.in.tu-clausthal.de/2022/#accepted)

Journal ref Autonomous Agents and Multiagent Systems. Best and Visionary Papers. AAMAS 2022. Lecture Notes in Computer Science(), vol 13441, pp. 114-129

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.03166 2022-11-02 cs.MA 60%

A Bayesian model of information cascades

Sriashalya Srivathsan, Stephen Cranefield, Jeremy Pitt

专题命中 Agent评测 :agent(abstract,comments);multi-agent(comments)

Comments 13 pages, 37 figures, Paper accepted for presentation A Bayesian model of information cascade in International Workshop on Coordination, Organizations, Institutions, Norms and Ethics for Governance of Multi-Agent Systems (COINE), co-located with AAMAS 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.14281 2021-11-10 cs.MA 60%

A Q-values Sharing Framework for Multiagent Reinforcement Learning under Budget Constraint

Changxi Zhu, Ho-fung Leung, Shuyue Hu, Yi Cai

专题命中 Agent评测 :agent(abstract,journal_ref);multi-agent(journal_ref)

Comments 31 pages, 16 figures, submitted to ACM Transactions on Autonomous and Adaptive Systems (TAAS)

Journal ref Changxi Zhu, Ho-fung Leung, Shuyue Hu, Yi Cai. A Q-values Sharing Framework for Multi-agent Reinforcement Learning under Budget Constraint. ACM Trans. Auton. Adapt. Syst. 15(2): 4:1-4:28 (2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.14177 2020-06-26 cs.GT 60%

Cost Sharing Security Information with Minimal Release Delay

Mingyu Guo, Yong Yang, Muhammad Ali Babar

专题命中 Agent评测 :agent(abstract,journal_ref);multi-agent(journal_ref)

Journal ref PRIMA 2018: Principles and Practice of Multi-Agent Systems

详情

展开后加载摘要…

URL PDF HTML 收藏