arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

野外的智能体:研究与部署的交汇点

Agents in the Wild: Where Research Meets Deployment

Grace Hui Yang, Pranav N. Venkit, Hooman Sedghamiz, Enrico Santus, Victor Dibia, Ioana Baldini

arXiv 2607.19336首次发表:更新:

发表机构

Georgetown University; Salesforce; Bayer; Bloomberg; Microsoft Research(乔治敦大学; Salesforce公司; 拜耳公司; 彭博公司; 微软研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

探讨基于大语言模型的智能体系统从研究到部署的转变,通过案例研究分析成功设计模式及失败缓解策略,为与会者提供跨行业安全可靠部署的全面视角、设计模式、评估清单和模板。

AI 中文摘要

基于大语言模型(LLM)的智能体系统,能够进行推理、规划、行动,并与工具及其他智能体协调,正迅速从研究原型向软件工程、科学发现和金融等领域的生产规模部署转变。学术工作侧重于基准测试和算法创新,而部署带来了关于鲁棒性、安全性和可靠性的新挑战。本教程汇聚研究人员和从业者,探讨推理与规划、多智能体协调及评估方面的进展,强调部署经验带来的开放挑战。通过药物发现和金融系统的应用案例研究,分析使智能体系统成功的常见设计模式,讨论针对失败模式的实际缓解策略,如验证管道、回退机制和人工监督。与会者将全面了解该领域,以及具体的设计模式、评估清单和跨行业安全可靠部署的模板。

英文摘要

Agentic systems large language model (LLM) based architectures capable of reasoning, planning, acting, and coordinating with tools and other agents are rapidly transitioning from research prototypes to production scale deployments across domains such as software engineering, scientific discovery, and finance. While academic work has emphasized benchmarks and algorithmic innovation, deployment raises new challenges around robustness, safety, and reliability. This tutorial brings together researchers and practitioners to explore advances in reasoning and planning, multi agent coordination, and evaluation, highlighting open challenges arising from deployment experience. Through applied case studies in pharmaceutical discovery and financial systems, we analyze common design patterns that make agentic systems successful, and discuss practical mitigation strategies for failure modes, such as verification pipelines, fallback mechanisms, and human in the loop supervision. Attendees will gain a comprehensive view of the field along with concrete design patterns, evaluation checklists, and templates for safe and reliable deployment across industries.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑