arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01189cs.SE

MADE:用于自主模型部署的信念驱动双智能体协调机制

MADE: Belief-Driven Dual-Agent Coordination for Autonomous Model Deployment

Yicheng Liu, Bolin Zhang, Weiran Liu, Yakun Zhang, Yangqin Jiang, Zhiying Tu, Dianhui Chu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对自动模型部署任务提出MADE双智能体协调系统,构建M2ABench基准测试集,实验显示MADE部署成功率达68.85%,优于SWE-agent与OpenHands。

中文摘要 AI 辅助

基于大语言模型(LLM)的智能体现已具备强大的通用能力,但在特定领域任务中仍存在不足,这促使人们通过集成外部工具来拓展其能力。开源社区提供了大量AI模型,通常以异构研究产物的形式发布,而将这些模型转换为可直接调用的应用程序编程接口(API)成本高昂且耗时费力。因此,自动模型部署对于弥合模型资源与工具可用性之间的差距至关重要,但该任务属于长周期多阶段任务,尚未得到充分探索。为应对这一挑战,我们提出了模型自动部署引擎(MADE),这是一个双智能体协调系统。具体而言,给定一个模型资源,MADE会迭代构建并验证部署产物,根据执行反馈更新其部署信念,并重新审视无效的上游产物,直至模型成功部署为可被其他智能体调用的API。我们还提出了M2ABench,这是一个用于将模型转换为可直接调用API任务的基准测试集,包含122个真实世界模型及用于评估的标准化测试用例。实验结果表明,MADE的部署成功率达68.85%,分别以13.93和44.26个百分点的优势优于SWE-agent和OpenHands。我们的代码和数据集可在该https链接获取。

英文摘要

LLM-based agents now have strong general capabilities. However, they still struggle with domain-specific tasks, motivating the integration of external tools to broaden their capabilities. The open-source community offers a vast array of AI models typically released as heterogeneous research artifacts, whereas transforming them into ready-to-call APIs is costly and labor-intensive. Automated model deployment is therefore essential for bridging the gap between model resources and tool usability, yet it remains a long-horizon, multi-stage task that has not been sufficiently explored. To tackle this challenge, we introduce Model Automated Deployment Engine (MADE), a dual-agent coordination system. Specifically, given a model resource, MADE iteratively constructs and validates the deployment artifacts, updates its deployment belief based on execution feedback, and revisits invalid upstream artifacts until the model is successfully served as a ready-to-call API that can then be used by other agents. We further introduce M2ABench, a benchmark for the task of transforming Models to ready-to-call APIs. M2ABench comprises 122 real-world models with standardized test cases for evaluation. Experimental results demonstrate that MADE achieves a deployment success rate of 68.85%, outperforming SWE-agent and OpenHands by 13.93 and 44.26 percentage points, respectively. Our code and dataset are publicly available at https://github.com/HITDiSC/MADE.

↑