AI 中文总结
研究针对可分块FPGA架构的LUT映射问题,提出迭代双输出感知LUT映射框架,集成于ABC映射器优化过程。通过交替单输出割选和双输出匹配,利用稀疏支持索引生成候选对并验证选择,反馈信息调整成本,实验表明该框架能有效降低LUT面积、加速并减少深度和面积。
AI 中文摘要
现代现场可编程门阵列采用可分块查找表(LUT),每个LUT可在共享物理LUT站点内实现单个大输入函数或两个较小函数。传统技术映射流程通常在单输出LUT映射后才进行双输出打包,阻碍了潜在配对机会对割选的影响。本文提出一种迭代双输出感知LUT映射框架,集成到伯克利ABC映射器的延迟、面积流和精确面积优化过程中。该方法在单输出割选和有界双输出匹配之间交替。利用基于稀疏支持的索引高效生成候选对,根据参数化架构、依赖性和输出特定的时序约束进行验证,并通过启发式基于分数的匹配过程进行选择。然后通过兼容性感知割成本调整将结果伙伴信息反馈到后续映射轮次。为保持时序精度,输入的物理并集仅用于架构合法性检查,而每个输出保留其自己的逻辑时序支持。在两个代表性架构模型下对EPFL组合基准套件进行的实验表明,相对于普通ABC,报告的LUT面积指标平均降低了34.96%和23.39%。与之前报道的最佳方法相比,所提出的框架实现了15.8倍的加速,同时进一步将深度降低了约5%,面积降低了1%。
英文摘要
Modern field-programmable gate arrays employ fracturable lookup tables (LUTs), each of which can implement either a single large-input function or two smaller functions within a shared physical LUT site. Conventional technology-mapping flows typically perform dual-output packing only after single-output LUT mapping, thereby preventing potential pairing opportunities from influencing cut selection. This paper presents an iterative dual-output-aware LUT-mapping framework integrated into the delay, area-flow, and exact-area optimization passes of the Berkeley ABC mapper. The proposed method alternates between single-output cut selection and bounded dual-output matching. Candidate pairs are efficiently generated using a sparse support-based index, validated against parameterized architectural, dependency, and output-specific timing constraints, and selected through a heuristic score-based matching procedure. The resulting partner information is then fed back into subsequent mapping rounds through a compatibility-aware cut-cost adjustment. To preserve timing accuracy, the physical union of the inputs is used only for architectural legality checking, while each output retains its own logical timing support. Experiments on the EPFL combinational benchmark suite under two representative architecture models demonstrate average reductions of 34.96% and 23.39% in the reported LUT-area metric relative to vanilla ABC. Compared with the best previously reported method, the proposed framework achieves a 15.8 times speedup while further reducing depth by approximately 5% and area by 1%.