AI 中文总结
VeriTrace是一种多智能体系统,通过赋予Inspector智能体完整调试动作空间实现类人时间探索,在VerilogEval-V2基准测试上首次达到100% Pass@1,准确率较最强复现基线提升5.1%。
AI 中文摘要
大型语言模型在自动化Verilog RTL生成方面展现出潜力,但最先进的多智能体系统在标准基准测试上的准确率停滞在约95%。我们将这一上限归因于不完整的调试动作空间:现有系统限制了智能体可检查的信号、可查询的时间窗口,或同时限制两者,将调试简化为对电路行为的狭窄预设视图的模式匹配,而非基于假设的根本原因分析。我们提出VeriTrace,一种多智能体系统,其Inspector智能体在完整的调试动作空间上运行,对信号选择、时间窗口边界和迭代深度具有独立控制权。这种我们称为智能体时间探索的能力,使智能体能够形成关于故障原因的假设、查询波形以获取证据并迭代完善其理解,模拟人类验证工程师的探索过程。VeriTrace在VerilogEval-V2上实现了100%的Pass@1,是首个在该基准测试上达到完美功能正确性的系统。在共享Claude Sonnet 4.0主干上,VeriTrace比最强的复现基线高出5.1%,表明调试智能性缩小了最终的准确率差距。
英文摘要
Large language models have shown promise for automated Verilog RTL generation, yet state-of-the-art multi-agent systems plateau at ~95% accuracy on standard benchmarks. We trace this ceiling to an incomplete debugging action space: existing systems restrict which signals the agent can inspect, which time windows it can query, or both, reducing debugging to pattern matching on a narrow, predetermined view of circuit behavior rather than hypothesis-driven root-cause analysis. We present VeriTrace, a multi-agent system whose Inspector agent operates over a complete debugging action space, with independent control over signal selection, time-window bounds, and iteration depth. This capability, which we term Agentic Temporal Exploration, enables the agent to form hypotheses about failure causes, query the waveform for evidence, and refine its understanding iteratively, mirroring the exploratory process of human verification engineers. VeriTrace achieves 100\% Pass@1 on VerilogEval-V2, the first system to attain perfect functional correctness on this benchmark. On a shared Claude Sonnet 4.0 backbone, VeriTrace outperforms the strongest reproduced baseline by +5.1%, demonstrating that debugging agency closes the final accuracy gap.
CommentsICLAD 2026, Long Oral