arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

共收录 5090 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. 工具调用 5090 篇

1511.00787 2015-11-04 cs.AI 83%

A Pareto Optimal D* Search Algorithm for Multiobjective Path Planning

Alexander Lavin

专题命中 工具调用 :planning(title,abstract);agent(abstract);分类 cs.AI

Comments arXiv admin note: substantial text overlap with arXiv:1505.05947

详情

展开后加载摘要…

URL PDF HTML 收藏
1410.6519 2014-10-27 cs.AI 83%

Justifying and Improving Meta-Agent Conflict-Based Search

David Tolpin

专题命中 工具调用 :agent(title,abstract);multi-agent(abstract);分类 cs.AI

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02464 2026-08-04 cs.AI cs.LG cs.SE 新提交 83%

Real-Time Detection and Repair of LLM Agent Failures

LLM智能体故障的实时检测与修复

Sunny Dubey

专题命中 工具调用 :agent(title,abstract);分类 cs.AI、cs.LG、cs.SE

AI总结 该研究提出基于LLM智能体步骤遥测的实时故障检测与修复系统,结合单类回声状态网络集成与确定性验证层,可高效检测并修复故障,提升任务成功率且成本极低。

Comments 16 pages, 5 figures. Code, data and demo: github.com/sunnydubey1111/agent-trajectory-sentinel Walkthrough: youtu.be/a05n_000klE

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05915 2023-10-10 cs.CL cs.AI cs.LG 83%

FireAct: Toward Language Agent Fine-tuning

Baian Chen, Chang Shu, Ehsan Shareghi, Nigel Collier, Karthik Narasimhan, Shunyu Yao

专题命中 工具调用 :agent(title,abstract);分类 cs.AI、cs.CL、cs.LG

Comments Code, data, and models are available at https://fireact-agent.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22793 2026-08-25 cs.CL cs.AI 新提交 82%

TRACE: A Self-Evolving Skill Bank for Consistent, Limit-Aware LLM Agents

TRACE:面向一致性、极限感知的自进化技能库

Wenhao Wu, Menghao Zhang, Xin Wang, Zhi Wang, Kun Shao, Jian Luan

机构 * Xiaomi Inc.(小米公司) Nanjing University(南京大学) Beijing University of Posts and Telecommunications(北京邮电大学) Tsinghua University(清华大学)

专题命中 工具调用 :agent(abstract);tool use(abstract);tool-use(abstract);agentic(abstract)

AI总结 TRACE是一种无需修改模型权重的自进化技能库方法,通过轨迹对比进化优化技能,在CAR-bench任务上显著提升LLM智能体的一致性性能,缩小潜在与可靠性能的差距。

Comments 9 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01236 2026-08-11 cs.CL cs.AI 版本更新 82%

Safeguarding LLM Agents from Misalignment through Provenance Analysis

通过溯源分析保护LLM智能体免受失调影响

Yining She, Yiliang Liang, Eunsuk Kang

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 工具调用 :agent(summary_cn,abstract);分类 cs.AI、cs.CL

AI总结 提出基于溯源分析的ProvenanceGuard框架,通过多阶段流水线检测工具调用前的三种失调类型,在Agent-SafetyBench和WorkBench上显著降低失调错误率并减少不必要的干预。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05159 2026-08-05 cs.CR cs.AI cs.LG 版本更新 82%

Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain

Agentland中的恶意行为:深入AI供应链后门问题

Léo Boisvert, Abhay Puri, Chandra Kiran Reddy Evuru, Nazanin Sepahvand, Nicolas Chapados, Quentin Cappart, Jason Stanley, Alexandre Lacoste, Krishnamurthy Dj Dvijotham, Alexandre Drouin

机构 * ServiceNow Research Mila - Qu\'ebec AI Institute

专题命中 工具调用 :agent(abstract);AI agent(abstract);tool use(abstract);agentic(abstract)

AI总结 研究探讨了在交互数据上微调AI代理时引入的安全漏洞,提出三种供应链威胁模型,证明通过污染少量演示即可使代理泄露用户信息。

Comments 27 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06716 2026-08-04 cs.AI cs.CL cs.CR 版本更新 82%

SIEVE: Selective Integrity Verification and Escalation for Defending LLM Agents against Indirect Prompt Injection

认知控制架构(CCA):一种用于鲁棒对齐AI代理的生命周期监督框架

Zhibo Liang, Tianze Hu, Zaiye Chen, Mingjie Tang

机构 * Sichuan University(四川大学)

专题命中 工具调用 :agent(abstract);tool-use(abstract);planning(abstract);agentic(abstract)

AI总结 本文提出认知控制架构(CCA),通过全生命周期认知监督框架,有效应对复杂IPI攻击,实现安全、功能与效率的平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22431 2026-07-03 cs.AI cs.CL cs.FL 82%

Monadic Context Engineering

单子上下文工程

Yifan Zhang, Yang Yuan, Mengdi Wang, Andrew Chi-Chih Yao

机构 * IIIS, Tsinghua University(清华大学人工智能研究院)

专题命中 工具调用 :agent(abstract);AI agent(abstract);autonomous agent(abstract);tool use(abstract)

AI总结 本文提出Monadic Context Engineering,利用函子、应用函子和单子的代数结构,为智能体设计提供形式基础,通过分层方法构建高效可靠的AI智能体。

Comments We found some issues in the categorical foundations of this work, so we respectfully withdraw it

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17946 2026-05-21 cs.AI cs.CV cs.LG 82%

SVFSearch: A Multimodal Knowledge-Intensive Benchmark for Short-Video Frame Search in the Gaming Vertical Domain

SVFSearch: 一种面向游戏垂直领域的多模态知识密集型短视频帧搜索基准

Lingtao Mao, Huangyu Dai, Xinyu Sun, Zihan Liang, Ben Chen, Chenyi Lei, Wenwu Ou

机构 * Kuaishou Technology(快手科技)

专题命中 工具调用 :agent(abstract);tool-use(abstract);workflow(abstract);agentic(abstract)

AI总结 本文提出SVFSearch,首个针对中文游戏领域短视频帧搜索的多模态知识密集型基准,通过5000个四选一测试示例和4198个辅助训练示例,评估了从直接问答到计划-行动-重新计划代理等多种方法在短视频帧搜索中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24459 2025-10-29 cs.AI cs.MA cs.SE 82%

Affordance Representation and Recognition for Autonomous Agents

Habtom Kahsay Gidey, Niklas Huber, Alexander Lenz, Alois Knoll

机构 * Technische Universität München(慕尼黑技术大学) Jessy Works(杰西工作)

专题命中 工具调用 :autonomous agent(title);agent(abstract,journal_ref);分类 cs.AI、cs.SE;multi-agent(journal_ref)

Journal ref The Second International Workshop on Hypermedia Multi-Agent Systems (HyperAgents 2025), in conjunction with the 28th European Conference on Artificial Intelligence (ECAI 2025); October 26, 2025, Bologna, Italy

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13310 2025-09-17 cs.CL 82%

Scaling Agents via Continual Pre-training

Liangcai Su, Zhen Zhang, Guangyu Li, Zhuo Chen, Chenxi Wang, Maojia Song, Xinyu Wang, Kuan Li, Jialong Wu, Xuanzhong Chen, Zile Qiao, Zhongwang Zhang, Huifeng Yin, Shihao Cai, Runnan Fang, Zhengwei Tao, Wenbiao Yin, Chenxiong Qian, Yong Jiang, Pengjun Xie, Fei Huang, Jingren Zhou

专题命中 工具调用 :agent(abstract,comments);tool use(abstract);tool-use(abstract);agentic(abstract)

Comments https://tongyi-agent.github.io/blog/introducing-tongyi-deep-research/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23283 2026-08-25 cs.AI cs.CL cs.LG 新提交 82%

Apodex 1.1: Scaling Agentic Intelligence for Complex Work

Apodex 1.1:为复杂工作扩展智能体智能

Apodex Team: B. An, B. Li, B. Wang, B. Zhang, B.L. Wang, C. Feng, C. Wei, C. Xue, C. Zhang, D. Ng, D. Ye, E. Min, F. Chen, F. Liu, F. Yang, F. Ye, H. Xu, H. Yang, H. Ye, H. Zhang, H. Zhao, J. Li, J. Lin, J. Xia, K. Jin, K. Wang, K. Yang, L. Bing, L. Lei, L. Su, Le. Wang, Lu. Wang, N. Wang, Q. Ren, Q. Yang, R. Li, S. Bai, S. Du, S. Li, S. Lin, S. Nie, S. Wang, S. Zhang, S.Z. Wang, Ta.Q. Fang, Ti.Q. Fang, W. Fang, W. Li, W. Zhang, X. Chen, X. Li, X. Tang, X. Wang, X. Xu, X. Zhang, X.Q. Wang, X.Y. Wang, Y. Deng, Y. Gao, Y. Hu, Y. Li, Y. Sui, Y. Wang, Y. Xiao, Y. Zhang, Z. Chen, Z. Cheng, Z. Feng, Z. Liang, Z. Zhang

专题命中 工具调用 :agentic(title,abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 Apodex 1.1 从环境扩展与智能体协调扩展两维度开发工作能力,在多领域复杂任务中性能领先,其 350亿参数的 Mini 版本可本地部署,助力构建长周期任务的重型求解器。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22751 2026-08-25 cs.IR 新提交 82%

Risk-Aware Reranking for Agentic Tool Retrieval

面向智能体工具检索的风险感知重排序

Qinfei Li, Xiaoxuan Dong, Jin Zhang, Dexu Yu, Wenhao Deng, Junchen Fu, Youhua Li, Hanwen Du, Chunxiao Li

专题命中 工具调用 :agentic(title);agent(abstract);tool-use(abstract)

AI总结 本文针对智能体工具检索的风险问题,提出轻量级风险感知重排序框架,通过权衡安全与效用改善相关性-安全 tradeoff,为安全关键部署提供保守操作点。

Comments Accepted by CIKM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20317 2026-08-21 cs.IR 新提交 82%

Projecting BrowseComp-Plus onto ClimbMix: Toward More Realistic Corpora for Agentic Search

将BrowseComp-Plus映射到ClimbMix:构建更符合智能体搜索需求的真实语料库

Sahel Sharifymoghaddam, Lingwei Gu, Yijun Ge, Jimmy Lin

专题命中 工具调用 :agentic(title,abstract);agent(abstract)

AI总结 该研究将BrowseComp-Plus的查询映射到NVIDIA的ClimbMix语料库,构建了与基准无关的映射流水线,生成57个有效查询,使智能体搜索难度向检索环节转移,发布了相关资源。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05493 2026-08-07 cs.PL cs.AI cs.CL cs.SE 新提交 82%

Learning Context-Free Grammars for Grammar-Constrained Decoding via Declarative Agentic Programming with Guarantees

通过带保障的声明式智能体编程学习用于语法约束解码的上下文无关文法

Kevin Cheang, Geoff Hulette, Rahul Kumar, Felipe R. Monteiro, Federico Mora, Robin Salkeld, Lin Tan, Serdar Tasiran

专题命中 工具调用 :agentic(title);agent(abstract);分类 cs.AI、cs.CL、cs.SE

AI总结 本研究提出名为Autogrammar的声明式智能体,可从文档和执行数据自动学习上下文无关文法,用于语法约束解码,在三种DSLs上的实验显示其生成的文法性能优于现有基线,能显著提升端到端LM的真实任务表现。

Comments 9 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02611 2026-08-05 cs.DC cs.AI cs.LG cs.SE 新提交 82%

KernelBrain: Coarse-to-Fine, Budget-Aware Search for Agentic GPU Kernel Optimization

KernelBrain:面向智能体的GPU内核优化的粗到细、预算感知搜索

Shuai Che, Gang Peng

专题命中 工具调用 :agentic(title);agent(abstract);分类 cs.AI、cs.LG、cs.SE

AI总结 KernelBrain是结合LLM引导变异等技术的GPU内核优化智能体,可提升内核质量与搜索效率,在Triton内核任务中获多倍加速且缩短优化时间。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02217 2026-08-04 cs.CV 新提交 82%

VC-Tooler: Learning Compositional and Adaptive Visual Tool Use

VC-Tooler:学习组合式与自适应视觉工具使用

Yizheng Wu, Jiashen Hua, Bing Deng, Jieping Ye

专题命中 工具调用 :tool use(title,abstract);agentic(abstract)

AI总结 VC-Tooler是一种学习组合式与自适应视觉工具使用的模型,通过分层合成轨迹库与两阶段训练,在通用和具身基准上达到开源模型最优性能,且具良好迁移能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23624 2026-07-30 cs.SE cs.AI cs.CL 版本更新 82%

Where Is the Cost of Third-Party API Routers in Agentic Software Development?

代理软件开发中第三方 API 路由器的成本在哪里?

Donghao Fu, Jingxin Li, Xue Jiang, Yihong Dong

专题命中 工具调用 :agentic(title);agent(abstract);分类 cs.AI、cs.CL、cs.SE

AI总结 研究第三方 API 路由器在代理软件开发中的影响,通过实证研究编码代理中路由器端注入的四个干预级别,开发 SIDEL 框架评估代理,发现路由器端干预难测,客户端缓解措施未完全恢复控制,强调需提供商端输出完整性保证。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25883 2026-07-29 cs.MA 新提交 82%

Towards a Systems Foundation for Agentic Cloud Management

迈向智能云管理的系统基础

Minghao Li, Ziqian Liu, Ziyu Mao, Daqian Ding, Yu Kang, Qingwei Lin, Tianyin Xu, Yiming Qiu

专题命中 工具调用 :agentic(title,abstract);agent(abstract)

AI总结 研究智能云管理中缺少系统基础的问题,核心方法是开发CloudWeaver智能管理基础架构,贡献在于能跨界面管理,确定会话上下文、协调并发操作,提供安全保证和可追溯反馈,经Azure API工作负载验证。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25565 2026-07-29 cs.CV 新提交 82%

ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition

ReDesign:通过智能分解从图像中恢复可编辑设计结构

Jooyeol Yun, Jintae Park, Hyesu Lim, Junha Hyung, Hyungjin Chung, Jaegul Choo

机构 * KAIST AI(韩国科学技术院人工智能研究所) Helmholtz Munich(慕尼黑亥姆霍兹中心) Korea University(韩国大学)

专题命中 工具调用 :agentic(title,abstract);tool use(abstract)

AI总结 研究旨在从图像恢复可编辑设计文件,提出ReDesign智能框架,通过跨模态选工具生成层层次结构,引入优雅验证与评估基准,在视觉保真度和编辑性上表现出色,超越基线和管道。

Comments Accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04569 2026-07-29 eess.SY cs.SY 版本更新 82%

LLMs for Agentic Home Energy Management

用于智能家庭能源管理的大语言模型

Sokipriala Jonah, Queen Moses, Abiola Babatunde, Michael Ajao-Olarinoye, Daniel Bammeke

专题命中 工具调用 :agentic(title);agent(abstract);function calling(abstract)

AI总结 研究家庭能源管理系统中,大语言模型智能体能否为多设备家庭能源调度提供自然语言接口。通过工具调用ReAct智能体及相关数据,对三个商业模型进行基准测试,结果显示模型表现各异,能实现高调度成功率和接近最优性。

Comments 15 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22975 2026-07-28 cond-mat.mtrl-sci 新提交 82%

AEcroscopyWave: Towards Self-Driving Characterization Platforms for Agentic AI

AEcroscopyWave:迈向用于智能AI的自动驾驶表征平台

Yongtao Liu, Jawad Chowdhury, Ganesh Narasimha, Ralph Bulanadi, Liam Collins, Ruben Millan Solsona, Marti Checa, Asraful Haque, Sumner B. Harris, Stephen Jesse, Rama Vasudevan

专题命中 工具调用 :agentic(title,abstract);AI agent(abstract)

AI总结 探讨电子材料表征的两种传统方式,介绍“自动驾驶”表征工具进展。基于AEcroscopyWave平台,通过创建控制硬件的API并整合AI方法,弥合行业规模自动化与高度定制系统差距,展示异构仪器供智能体使用的好处。

Comments 17 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.16372 2026-07-21 cs.SE cs.AI cs.LG cs.PL 新提交 82%

AoA: Theorem Proving Agent over Abstract Syntax Tree of Redesigned Language

AoA:基于重新设计语言抽象语法树的定理证明智能体

Qiyuan Xu, Joshua Ong Jun Leang, Renxi Wang, Wenda Li, Haonan Li, Luke Ong, Conrad Watt

机构 * Nanyang Technological University Singapore(南洋理工大学新加坡分校) Imperial College London London, UK(伦敦帝国理工学院伦敦分校) University of Edinburgh Edinburgh, UK(爱丁堡大学爱丁堡分校)

专题命中 工具调用 :agent(title,abstract);分类 cs.AI、cs.LG、cs.SE

AI总结 研究针对交互式定理证明中人工操作限制可扩展性及基于LLM的证明智能体成本高的问题,提出将智能体从源文本提升到抽象语法树的方法,实现了AoA,在多个方面有显著提升且解决更多难题。

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15518 2026-07-20 physics.ed-ph physics.comp-ph 新提交 82%

A Tool-Invariant Framework for Teaching and Assessing Computational Methods in the Age of Agentic AI

在智能人工智能时代用于教授和评估计算方法的工具不变框架

Larry Engelhardt

专题命中 工具调用 :agentic(title,abstract);agent(abstract)

AI总结 探讨在智能人工智能时代计算方法教学与评估,提出工具不变框架,指出验证是关键技能,阐述评估变化,介绍针对小班课程的实际应对,强调计算物理课程中学生解释和捍卫计算工件能力的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12650 2026-07-15 cs.LG cs.AI cs.CY cs.SE 新提交 82%

Evidence-Grounded Verified Agentic Reasoning: A Path Toward Eliminating LLM Hallucination in Empirical Inference via Tool-Attested Kernel Proofs

基于证据的可验证智能推理:通过工具验证内核证明消除经验推理中语言模型幻觉的途径

Junyu Ren

专题命中 工具调用 :agentic(title,abstract);分类 cs.AI、cs.LG、cs.SE

AI总结 研究旨在消除语言模型经验推理中的幻觉,提出基于Lean 4的EG-VAR工具调用架构,通过工具验证公理等生成可验证声明,经实验在数值推理等测试中表现良好,定位为高风险经验声明的技术治理接口,可审计相关条件并转化错误为审计目标。

Comments Accepted at the ICML 2026 TAIGR workshop. System name: EG-VAR

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02716 2026-07-07 cs.MA 新提交 82%

Evaluating Large Language Models for Decision-Making in Agent-Based Urban Mobility Simulations

在基于代理的城市交通模拟中评估用于决策的大语言模型

Bruno Cascaes Alves, Míriam Blank Born, Ulisses Gilioli Francescatto Júnior, Felipe Moura Goulart, Letícia Brandão Caldas, Marilton Sanchotene de Aguiar

专题命中 工具调用 :agent(title,abstract);multi-agent(abstract)

AI总结 研究在多智能体模拟中集成大语言模型作为决策组件,提出混合架构,通过API连接GAMA平台与外部基于大语言模型的模块,能指导智能体重规划行为,比较不同场景下效果,显示其可丰富行为表示。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31935 2026-07-01 econ.EM 新提交 82%

Delegation Rights: Property, Agency, and Investment Incentives in the Age of AI Agents

委托权:AI代理时代中的财产、代理与投资激励

Yukun Zhang, Kemu Xu

专题命中 工具调用 :AI agent(title,abstract);agent(abstract)

AI总结 本文定义委托权为账户持有人授权代理执行的可撤销、身份保留、范围受限且模式特定的权限,通过三方不完全契约模型分析平台控制、用户控制与认证委托三种机制对投资激励和风险分配的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31665 2026-07-01 cs.MA 新提交 82%

ForecastAgentSearch: Towards a Multi-Expert Agent Search System for Geopolitical Event Forecasting

ForecastAgentSearch:面向地缘政治事件预测的多专家智能体搜索系统

Miaomiao Cai, He Chang, Yunshan Ma, See-kiong Ng

专题命中 工具调用 :agent(title,abstract);multi-agent(abstract)

AI总结 提出ForecastAgentSearch框架,将地缘政治事件预测建模为多专家智能体搜索问题,通过检索、排序和协调专业智能体生成预测,并讨论关键设计挑战与评估方案。

Journal ref SIGIR 2026 AgentSearch Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31504 2026-07-01 cs.CV 新提交 82%

SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search

SimpleSearch-VL:多模态智能深度搜索的简单配方

Ming Dai, Zhihong Lu, Jinjie Gu, Jiedong Zhuang, Yefeng Liu, Wankou Yang, Jian Wang, Chunhua Shen

机构 * Southeast University(东南大学) Ant Group(蚂蚁集团)

专题命中 工具调用 :agentic(title,abstract);agent(abstract)

AI总结 提出SimpleSearch-VL框架,通过因子化自适应展开和证据验证推理提升多模态智能搜索的效率与可靠性,仅用少量数据即显著超越基线。

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏