MEGA:基于智慧图的自进化智能体优化基础设施
MEGA: Self-Evolving Agent Optimization Infrastructure via Wisdom Graph
浏览论文内容
中文总结 AI 辅助
MEGA是一种自进化智能体优化基础设施,通过三层架构实现智慧积累、组合推理与自进化,可系统性优化异构智能体系统,将优化与知识进化融为一体。
中文摘要 AI 辅助
随着编码智能体越来越多地处理实现工作,核心挑战已从构建单个智能体转向构建能系统性改进智能体的基础设施。现有方法存在三大不足:优化智能体系统时无法积累可迁移知识、积累知识后无法对其进行组合推理、缺乏通过操作证据实现知识自进化的机制。MEGA(元评估驱动的自适应)作为一种自进化基础设施,解决了上述缺口:每个优化周期都会产生可持久化的资产,对这些资产的组合推理会指导后续优化,操作证据则会同时优化积累的智慧与支配推理的规则。第一层通过行为模式聚类和实证A/B验证,从智能体会话中提炼可复用智慧,将每个过程转化为可持久化资产;第二层将这些资产分解为类型化智慧图中的原子PCR(主-上下文-结果)单元,执行演绎、归纳、溯因推理以拓展隐式关系,再通过组合检索组装特定上下文的执行计划,该计划能呈现仅靠嵌入相似度无法获取的桥接知识;第三层针对异构智能体工作流(代码节点、大语言模型调用、工具使用智能体)执行多智能体协同优化,通过消除数据方差的受控评估,将改进效果归因于特定策略变更。第三层反馈的证据会驱动支配智慧组合的筛选策略与跨运行积累的优化轨迹的自进化,最终形成优化智能体系统与进化指导优化的知识为同一过程的基础设施。
英文摘要
As coding agents increasingly handle implementation, the central challenge shifts from building individual agents to building an infrastructure that systematically improves them. Current approaches optimize agent systems without accumulating transferable knowledge, accumulate knowledge without compositional reasoning over it, and lack a mechanism for that knowledge to self-evolve through operational evidence. MEGA (Meta Evaluation-Grounded Adaptation) addresses these gaps as a self-evolving infrastructure: each optimization cycle produces durable assets, compositional reasoning over those assets guides subsequent optimization, and operational evidence refines both the accumulated wisdom and the reasoning that governs it. Layer 1 distills reusable wisdom from agent sessions through behavioral-pattern clustering and empirical A/B validation, transforming each process into a durable asset. Layer 2 decomposes these assets into atomic PCR (Primary-Context-Resultant) units within a typed Wisdom Graph and performs deductive, abductive, and inductive reasoning to expand implicit relations; it then assembles context-specific execution plans through compositional retrieval that surfaces bridging knowledge unreachable by embedding similarity alone. Layer 3 performs multi-agent collaborative optimization over heterogeneous agent workflows (code nodes, LLM calls, and tool-using agents), attributing improvement effects to specific strategy changes through controlled evaluation that eliminates data variance. Evidence fed back from Layer 3 drives the self-evolution of both the curation strategies that govern wisdom composition and the optimization trajectories accumulated across runs. The result is an infrastructure in which optimizing an agent system and evolving the knowledge that guides optimization are one and the same process.