arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02996cs.NE

面向可信人工智能的进化计算:从攻击与防御到自进化时代

Evolutionary Computation for Trustworthy AI: From Attacks and Defenses to Self-Evolving Era

Junhao Dong, Chenkai Wang, Xuanhui Lin, Mingrong Gong, Siyu Wang, Yuqing Wen, Jiao Liu, Catherine Huang, Gary G. Yen, Xin Yao, Yew-Soon Ong

首次发表
浏览论文内容

中文总结 AI 辅助

本文综述进化计算在可信AI中的应用,涵盖进化攻击、进化防御和可信自进化系统,通过共同进化视角连接这些方向,并讨论评估方法、基准及未来挑战。

中文摘要 AI 辅助

随着人工智能(AI)从特定任务模型演变为基础模型和智能体,可信AI的范围已从模型级鲁棒性扩展到更广泛AI系统的可靠性和安全性。这一演变也将攻击面从单个模型扩展到更广泛的系统级交互,包括工具使用、上下文以及与动态环境的交互轨迹。因此,在不断变化或故意操纵的条件下维持可靠和安全的行为变得越来越具有挑战性。寻找有效的攻击和防御通常依赖于黑盒反馈,以在单词、动作、系统组件或其组合之间进行离散选择。多目标和昂贵的候选评估进一步限制了可探索的范围。进化计算(EC)凭借其基于种群、无梯度的搜索以及灵活的变异和选择机制,非常适合这些场景。本综述回顾了EC如何应用于可信AI的三个方向:进化攻击、进化防御和可信的自进化AI系统。与先前将可信AI、EC和自进化系统大致分开处理的综述不同,我们通过共同的进化视角将这些线索联系起来。对于自进化AI,我们考察可信性如何支配塑造后续适应的更新的生成和保留。我们进一步从可信性和进化搜索两个角度综合了评估方法和基准资源。最后,我们讨论了在可信AI中更有效和可靠地使用EC的关键挑战和未来研究方向。

英文摘要

As Artificial Intelligence (AI) has evolved from task-specific models to foundation models and agents, the scope of trustworthy AI has expanded from model-level robustness to the reliability and safety of broader AI systems. This evolution has also expanded the attack surface from individual models to broader system-level interactions, including tool use, context, and interaction trajectories with dynamic environments. As a result, maintaining reliable and safe behavior under changing or deliberately manipulated conditions has become increasingly challenging. The search for effective attacks and defenses often relies on black-box feedback to navigate discrete choices among words, actions, system components, or their combinations. Multiple objectives and expensive candidate evaluations further limit what can be explored. Evolutionary Computation (EC), with its population-based, gradient-free search and flexible variation and selection mechanisms, is well suited to these settings. This survey reviews how EC has been applied to trustworthy AI across three directions: evolutionary attacks, evolutionary defenses, and trustworthy self-evolving AI systems. Unlike prior reviews that treat trustworthy AI, EC, and self-evolving systems largely separately, we connect these lines through a common evolutionary perspective. For self-evolving AI, we examine how trustworthiness governs the generation and retention of updates that shape subsequent adaptation. We further synthesize evaluation methods and benchmark resources from both trustworthiness and evolutionary-search perspectives. Finally, we discuss key challenges and future research directions toward more effective and reliable use of EC in trustworthy AI.

发表机构

  • Nanyang Technological University, Singapore(南洋理工大学)
  • Center for Frontier AI Research, Institute of High Performance Computing, A*STAR, Singapore(新加坡国立研究科学院前沿人工智能研究中心高性能计算研究所)
  • Google(谷歌)
  • Sichuan University, China(四川大学)
  • Lingnan University, Hong Kong, China(岭南大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑