arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MOAE:基于帕累托保持搜索的多目标智能体进化

MOAE: Multi-Objective Agent Evolution with Pareto-Preserving Search

Hengle Jiang, Qijun Cai, Ziying Luo, Ke Tang

arXiv 2609.05992首次发表:更新:

发表机构

Southern University of Science and Technology; Guangdong Provincial Key Laboratory of Brain-inspired Intelligent Computation(南方科技大学; 广东省类脑智能计算重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多目标智能体优化中固定标量化导致候选丢失的问题,提出MOAE方法,通过帕累托保持的进化搜索维护非支配候选,在实验中提升任务性能与轨迹质量并保持安全性。

AI 中文摘要

随着基于大语言模型的智能体不断进步,对其评估也变得越来越多元化:一个能力强的智能体不仅要实现高任务完成准确率,还需在交互质量、安全性和效率方面表现出色,这引出了一个核心问题:这些目标能否同时优化?现有方法已考虑多目标,但许多方法将异质度量压缩为固定标量分数。这种标量化依赖于指标归一化和偏好权重,并可能丢弃代表有用部署权衡的候选方案。我们提出多目标智能体进化(MOAE),它将迭代式上下文内精炼组织为对完整智能体轨迹的帕累托保持进化搜索。在有限的轨迹预算下,MOAE维护一个非支配候选的经验档案,使用目标特定诊断来指导后代生成,并仅在部署时应用约束感知选择。这将搜索过程中的候选保留与用于返回最终解决方案的偏好分离开来。该方法无需参数更新,并允许每个目标替换为任何可测量属性,我们将其实例化为任务性能、轨迹质量和安全性。在TravelPlanner和AgentDojo上的实验表明,在匹配的轨迹预算下,MOAE持续提高任务性能和轨迹质量,同时保持强大的安全性。搜索行为分析进一步表明,帕累托保持扩展了可达目标区域,并增加了联合改进的频率。这些结果证明了帕累托保持上下文内进化在优化多个智能体属性方面的潜力,而无需在搜索过程中承诺固定标量化。

英文摘要

As LLM-based agents continue to advance, their evaluation has become increasingly multifaceted: a capable agent must not only achieve high task completion accuracy but also perform well in interaction quality, safety, and efficiency, raising a central question: can these objectives be optimized simultaneously? Existing methods have considered multiple objectives, but many collapse heterogeneous measurements into a fixed scalar score. Such scalarization depends on metric normalization and preference weights and may discard candidates that represent useful deployment trade-offs. We introduce Multi-Objective Agent Evolution (MOAE), which organizes iterative in-context refinement as a Pareto-preserving evolutionary search over complete agent rollouts. Given a limited rollout budget, MOAE maintains an empirical archive of non-dominated candidates, uses objective-specific diagnostics to guide offspring generation, and applies constraint-aware selection only at deployment. This separates candidate preservation during search from the preference used to return a final solution. The procedure requires no parameter updates and allows each objective to be replaced by any measurable property, which we instantiate as task performance, trajectory quality, and safety. Experiments on TravelPlanner and AgentDojo show that MOAE consistently improves task performance and trajectory quality while maintaining strong safety under matched rollout budgets. Search-behavior analysis further shows that Pareto preservation expands the attainable objective region and increases the frequency of joint improvement. These results demonstrate the potential of Pareto-preserving in-context evolution for optimizing multiple agent properties without committing to a fixed scalarization during search.

Comments18 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑