从片段到功能证据:面向EDA文档问答的功能感知检索
From Chunks to Functional Evidence: Function-Aware Retrieval for EDA Documentation QA
- Zhejiang University(浙江大学)
- Shanghai Innovation Institute(上海创新研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对EDA文档问答中查询与知识组织不匹配的问题,提出将类型化工件组织为功能单元作为检索单元,结合单元与片段检索并统一重排序,显著提升ROUGE-L指标。
AI中文摘要:
检索增强生成(RAG)被广泛用于将答案基于文档。然而,对于复杂的技术文档,主要瓶颈往往不是模型推理,而是查询与知识组织方式之间的不匹配。这种不匹配在电子设计自动化(EDA)文档中尤为突出,因为回答所需的信息分散在异构但紧密耦合的工件中。因此,我们重新设计了RAG的基本检索单元。我们不再处理孤立的片段或二元关系,而是将类型化工件收集到EDA功能单元中。每个单元记录为一条超边,并链接到其源片段。然后,我们训练一个编码器,将查询与功能单元对齐,并将单元检索与直接片段检索相结合。在将选定的单元映射回其源后,一个统一的重排序器选择提供给生成器的证据。在新构建的EDADocEval-QA数据集上,我们的方法相比Chunk RAG将ROUGE-L提高了37.1%,相比最强的图基线提高了55.6%。在公开的ORD-MMBench基准上,它相比最强基线将ROUGE-L提高了30.0%。这些结果支持在所评估的EDA文档设置中采用功能感知的证据组织。
英文摘要:
Retrieval-Augmented Generation (RAG) is widely used to ground answers in documents. For complex technical documentation, however, the primary bottleneck is often not model reasoning but a mismatch between a query and the way knowledge is organized for retrieval. This mismatch is pronounced in Electronic Design Automation (EDA) documentation, where the information needed for an answer is scattered across heterogeneous yet tightly coupled artifacts. We therefore redesign the basic retrieval unit of RAG. Instead of operating on isolated chunks or binary relations, we collect typed artifacts into EDA functional units. Each unit is recorded as a hyperedge with links to its source chunks. We then train an encoder to align queries with functional units and combine unit retrieval with direct chunk retrieval. After mapping the selected units back to their sources, a unified reranker chooses the evidence given to the generator. On the newly constructed EDADocEval-QA dataset, our method improves ROUGE-L by 37.1% over Chunk RAG and 55.6% over the strongest graph baseline. On the public ORD-MMBench benchmark, it improves ROUGE-L by 30.0% over the strongest baseline. These results support function-aware evidence organization in the evaluated EDA documentation settings.