WaveformQA:对数字波形上的大语言模型时间推理进行基准测试
WaveformQA: Benchmarking LLM Temporal Reasoning on Digital Waveforms
浏览论文内容
中文总结 AI 辅助
研究大语言模型对数字波形的时间推理能力,提出开源问答基准WaveformQA,含360个不同难度问题,波形由开源设计生成。评估发现模型在简单查询有准确率,但复杂问题上因窗口限制等性能下降,JSON表示可提高准确率,框架支持扩展实验。
中文摘要 AI 辅助
大语言模型在代码生成和推理方面展现出强大能力,但其对数字波形数据进行时间推理的能力很大程度上未被探索。尽管数字波形推理是设计验证的关键瓶颈,现有基准主要评估硬件描述语言代码生成,仅将波形用作补充上下文。本文提出WaveformQA,一个用于评估大语言模型对数字波形时间推理的开源问答基准。该基准包含360个问题,涵盖八个不同难度类别,有针对多信号相关性和事件排序的问题。波形由开源设计实现生成,确保可重复性并基于实际硬件行为。对前沿大语言模型的评估表明,模型在简单查询上有合理准确率,但在复杂时间和多步问题上因上下文窗口限制和推理困难而性能下降。此外,波形的事件时间JSON表示比标准化值变化转储格式提高了大语言模型推理准确率。开源框架支持扩展到新问题类别和导入新波形源,使研究人员能快速进行时间推理实验原型。
英文摘要
Large Language Models (LLMs) have demonstrated strong capabilities in code generation and reasoning, yet their ability to perform temporal reasoning over digital waveform data remains largely unexplored. Although reasoning over digital waveforms is a critical bottleneck in design verification, existing benchmarks primarily evaluate hardware description language (HDL) code generation and use waveforms only as supplementary context. This paper presents WaveformQA, an open-source question-answering benchmark for evaluating LLM temporal reasoning over digital waveforms. The benchmark comprises 360 questions with programmatically generated ground truths across eight categories of varying difficulty, including questions targeting multi-signal correlation and event ordering. Waveforms are generated from open-source design implementations, ensuring reproducibility and grounding the benchmark in real hardware behavior. Evaluation of frontier LLMs reveals that while models achieve reasonable accuracy on simple queries, performance degrades due to context window limitations and reasoning difficulties on complex temporal and multi-step questions. In addition, we show that an event-time JSON representation of waveforms improves LLM reasoning accuracy versus the standardized value change dump (VCD) format. The open-source framework supports extending to new question categories and importing new waveform sources, enabling researchers to rapidly prototype temporal reasoning experiments.