arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23363cs.AI

TicTacBench:基准测试编码智能体的时序收敛能力

TicTacBench: Benchmarking Timing Closure Capabilities of Coding Agents

Bowei Wang, Zhigang Fang, Zhijie Yang, Renzhi Chen, Shanshan Li, Lei Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出TicTacBench基准,评估编码智能体在RTL时序收敛上的能力,发现最佳智能体仅收敛53.3%任务,并提出TicTacSkill方法将收敛率提升9%。

中文摘要 AI 辅助

近年来,大型语言模型(LLM)的进展催生了能够执行复杂工程任务的编码智能体,包括寄存器传输级(RTL)设计与优化。现有的RTL基准主要评估生成RTL设计的功能正确性以及性能、功耗和面积(PPA),而智能体在时序收敛方面的能力评估不足。我们提出了TicTacBench,一个专门设计用于在布局布线后(post-PnR)评估下评估编码智能体RTL级时序收敛能力的基准。TicTacBench包含30个多样化的任务,每个任务提供一个次优的RTL设计、现实的时序约束、功能等价性验证和时序报告。通过由8个前沿LLM驱动的超过300次编码智能体运行,我们发现即使最好的智能体也只能收敛53.3%的任务,平均面积延迟积(ADP)退化7.18%,能量延迟平方积(EDDP)改善8.83%。我们识别了常见的失败类别,解释了智能体为何无法收敛时序。然后我们提出了TicTacSkill,一种引导智能体遵循标准时序收敛流程的新方法,将时序收敛率提高了9%。这些结果表明,虽然编码智能体在RTL设计方面取得了显著进展,但其时序收敛能力仍有很大的提升空间。

英文摘要

Recent advances in large language models (LLMs) have led to the emergence of coding agents capable of performing complex engineering tasks, including register-transfer level (RTL) design and optimization. Existing RTL benchmarks mainly evaluate functional correctness and performance, power, and area (PPA) of the generated RTL designs, leaving agents' ability for \emph{timing closure} under-evaluated. We propose TicTacBench, a benchmark specifically designed to evaluate coding agents' capabilities for RTL-level timing closure under post-place-and-route (post-PnR) evaluation. TicTacBench contains 30 diverse tasks, each provided with a suboptimal RTL design, realistic timing constraints, functional equivalence verification, and timing reports. With over 300 runs of coding agents driven by 8 frontier LLMs, we find that even the best agent can only close 53.3\% of tasks with 7.18\% area-delay product (ADP) degradation and 8.83\% energy-delay-squared product (EDDP) improvement on average. We identify common failure categories that explain why agents fail to close timing. Then we propose TicTacSkill, a new method that guides agents to follow standard timing-closure procedures and improves the Timing Closure Rate by 9\%. These results suggest that while coding agents have made significant progress in RTL design, their timing-closure capability still has substantial room for improvement.

发表机构

  • National University of Defense Technology(国防科技大学)
  • Academy of Military Science(军事科学院)
  • Qiyuan Lab(启元实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑