arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当大语言模型过度回答时:测量和缓解基于大语言模型的硬件描述语言问答中的质量问题

When LLMs Over-Answer: Measuring and Mitigating Quality Issues in LLM-Based Hardware Description Language Question Answering

Ziteng Hu, Jiachi Chen, Wenhao Lv, Huan Zhang, Yingjie Xia

arXiv 2607.17063首次发表:更新:

AI 中文总结

研究基于大语言模型的硬件描述语言问答质量问题,收集相关问答帖子构建数据集并开展用户研究,发现LLMs存在过度回答倾向。据此提出多智能体框架,用特定指标评估答案质量,该框架有效提升了主流LLMs的回答质量。

AI 中文摘要

大语言模型(LLMs)的快速发展使从业者越来越依赖它们来回答有关硬件描述语言(HDLs)的问题。由于HDL最终会被合成到物理硬件中,不准确或冗余的答案可能会导致时序违规或不可合成的逻辑,这些问题只会在设计流程后期出现,因此HDL答案的质量尤为重要。然而,LLM生成的回答质量,特别是与人类专家提供的答案相比,仍不清楚。为了研究这个问题,我们从Stack Overflow收集了6246个有公认答案的HDL问答帖子,并将它们整理成一个数据集,分为四个主要类别(概念、调试、生成和优化)和十个子类别。使用这个数据集,我们对19名有一到三年经验的HDL工程师进行了一项用户研究。我们的发现揭示了一种普遍的过度回答倾向:LLMs提供了正确的内容,但却将其埋没在冗余的替代方案(65.7%)和冗长的填充内容(69.1%)之下,而近一半的答案(49.0%)未能与专家答案完全一致,但参与者仍然更喜欢LLM回答的可读性(58.3%)。基于这些发现,我们提出了一个多智能体框架来改进基于LLM的HDL问答。我们使用一个LLM作为评判器和两个结构指标来评估答案质量:核心答案的数量,它反映了冗余,因为LLMs经常提供多个替代解决方案,以及非核心内容的长度,它反映了冗长。在四个主流LLMs上进行评估,我们的框架将平均核心答案质量分数从3.71提高到4.67(+0.96),非核心内容质量从3.72提高到4.23(+0.51),满分为五分。

英文摘要

The rapid advancement of large language models (LLMs) has led practitioners to increasingly rely on them for answering questions about hardware description languages (HDLs). Because HDL is ultimately synthesized into physical hardware, an imprecise or redundant answer can propagate into timing violations or non-synthesizable logic that surface only late in the design flow, making the quality of HDL answers especially consequential. However, the quality of LLM-generated responses, particularly in comparison with answers provided by human experts, remains unclear. To investigate this question, we collect 6,246 HDL Q&A posts with accepted answers from Stack Overflow and curate them into a dataset, organized into a taxonomy of four main categories (Conceptual, Debugging, Generation, and Optimization) and ten subcategories. Using this dataset, we design a user study conducted with 19 HDL engineers with one to three years of experience. Our findings reveal a pervasive over answering tendency: LLMs supply correct content but bury it under redundant alternatives (65.7%) and verbose padding (69.1%), while nearly half of answers (49.0%) fail to fully align with expert answers yet participants still preferred LLM responses for readability (58.3%). Motivated by these findings, we propose a multi-agent framework for improving LLM-based HDL question answering. We evaluate answer quality using an LLM-as-Judge and two structural metrics: the number of core answers, which reflects redundancy since LLMs often provide multiple alternative solutions, and the length of non-core content, which reflects verbosity. Evaluated on the four mainstream LLMs, our framework increases the average core-answer quality score from 3.71 to 4.67 (+0.96) and the non-core content quality from 3.72 to 4.23 (+0.51), on a five-point scale.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑