arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14579cs.AI

SKILL:用于逻辑优化的自校正知识引导迭代大语言模型智能体

SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization

  • University of California, Riverside(加州大学河滨分校)

机构由 AI 辅助整理,请以论文原文为准。

Rui Yang

AI总结:

针对逻辑综合优化的挑战,提出SKILL智能体,结合多LLM推理与RL交互,经基准测试实现12.4%的PDA提升及86.3%的50万门级逻辑系统成功率。

AI中文摘要:

逻辑综合优化面临重大挑战,原因在于搜索空间呈指数级增长、奖励信号稀疏以及逻辑结构多样。传统专家设计的流程缺乏适应性,而强化学习(RL)方法通常存在样本效率低、可解释性有限的问题。我们提出SKILL,即自校正知识引导迭代大语言模型智能体,它将多智能体LLM推理与基于RL的环境交互相结合,用于自动化综合优化。SKILL协调三个专用LLM:用于战略规划的GPT-4o、用于详细推理的Claude Sonnet 4、用于高效分析的Gemini 2.5 Pro,以及一个基于PPO的RL智能体,该智能体通过与综合工具直接交互学习可执行策略。一个新颖的自校正模块监测环境反馈(PDA指标),检测次优行为并调用LLM引导的恢复策略。在IWLS、OpenCores和EPFL基准测试上的评估显示,SKILL相较于专家流程实现了12.4%的PDA提升,在规模达50万门的逻辑系统上达到86.3%的成功率。

英文摘要:

Logic synthesis optimization poses significant challenges due to exponentially growing search spaces, sparse reward signals, and diverse logic structures. Traditional expert-designed flows lack adaptability, while reinforcement learning (RL) methods often suffer from low sample efficiency and limited interpretability. We introduce SKILL, a Self-correcting Knowledge-guided Iterative Large Language Model Agent that unifies multi-agent LLM reasoning and RL-based environment interaction for automated synthesis optimization. SKILL coordinates three specialized LLMs: GPT-4o for strategic planning, Claude Sonnet 4 for detailed reasoning, and Gemini 2.5 Pro for efficient analysis with a PPO-based RL agent that learns actionable policies through direct interaction with synthesis tools. A novel self-correcting module monitors environment feedback (PDA metrics), detects suboptimal behaviors, and invokes LLM-guided recovery strategies. Evaluations on IWLS, OpenCores, and EPFL benchmarks show SKILL achieves a 12.4 % PDA improvement over expert flows and 86.3% success rate on logic systems up to 500K gates.

↑