arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34553cs.LGcs.SE

用强化学习验证神经网络

Verifying Neural Networks with Reinforcement Learning

Hai Duong, Thanh Le, ThanhVu Nguyen

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出RSB,一种基于强化学习的框架,通过训练演员-评论家模型优化分支启发式,在600个实例上多解决11%并减少50%分支探索,提升DNN形式化验证效率。

中文摘要 AI 辅助

形式化验证在确保部署于安全关键系统中的深度神经网络(DNNs)的可靠性方面可发挥关键作用。现代DNN验证器采用分支定界框架,该框架在分支(将问题拆分为更小的子问题)与定界(剪枝子问题)之间交替进行,以高效探索验证空间。然而,现有的分支启发式方法基于静态评分函数做出贪婪决策,它们不预见长期效率,也未利用日益增长的验证数据来提升性能。本工作引入RSB,一个强化学习框架,学习优化基线分支启发式方法。它训练一个演员-评论家架构,以最大化累积未来奖励而非即时分数。演员从原始神经元特征和学习到的图嵌入的观察中生成注意力权重,这些权重重新缩放基线启发式分数以指导神经元分支。在600个具有挑战性的实例上的评估表明,RSB始终优于最先进的分支启发式方法,多解决11%的实例,同时将分支探索减少50%。

英文摘要

Formal verification can play a key role in ensuring the reliability of Deep Neural Networks (DNNs) deployed in safety-critical systems. Modern DNN verifiers employ a branch-and-bound framework, which alternates between branching (splitting into smaller subproblems) and bounding (pruning subproblems) to efficiently explore the verification space. However, existing branching heuristics make greedy decisions based on static scoring functions. They do not anticipate long-term efficiency or leverage the growing availability of verification data to improve performance. This work introduces RSB, a reinforcement learning framework that learns to refine baseline branching heuristics. It trains an actor-critic architecture to maximize cumulative future rewards rather than immediate scores. The actor generates attention weights from observations of raw neuron features and learned graph embeddings, which rescale baseline heuristic scores to guide neuron branching. Evaluation on 600 challenging instances demonstrates that RSB consistently outperforms state-of-the-art branching heuristics, solving 11% more instances while reducing branch exploration by 50%.

发表机构

  • George Mason University(乔治梅森大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑