arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FAPO:基于LUT的FPGA的扇出感知映射后优化

FAPO: Fanout-Aware Post-Mapping Optimization for LUT-Based FPGAs

Xiaoyu Hao, Keren Zhu

arXiv 2610.06937首次发表:更新:

发表机构

State Key Lab of Integrated Chips and Systems, Fudan University(复旦大学集成芯片与系统国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FAPO是一种映射后优化器,通过联合优化汇点分配和LUT切割,在面积约束下减少扇出瓶颈,实现高达14.48%的ATP降低和7.04%的布线延迟降低。

AI 中文摘要

在FPGA技术映射过程中,降低逻辑深度并不一定能消除所生成LUT网络中的扇出瓶颈。重载LUT可以缓解这些瓶颈,但增加的副本会消耗面积并可能增加上游负载。我们提出FAPO,一种与映射器无关的映射后优化器,在最终面积约束下联合优化汇点分配和LUT切割。切割切换和恢复通过所得的活跃LUT网络进行评估。我们推导了延迟降低的必要全关键路径条件,并在受限的相同切割集合内通过最小割获得最小副本提案。这些提案补充了贪婪细化,通过全网络面积和延迟选择映射。在EPFL基准测试中,使用三种初始映射器,在5%最终面积约束下,扇出感知面积-延迟乘积(ATP)的几何平均降低高达14.48%。进一步评估显示,相对于ABC,布线关键路径延迟降低了7.04%。

英文摘要

Reducing logic depth during FPGA technology mapping does not necessarily eliminate fanout bottlenecks in the resulting LUT networks.Replicating heavily loaded LUTs can alleviate these bottlenecks, but the added copies consume area and may increase upstream loading. We propose FAPO, a mapper-independent post-mapping optimizer that jointly refines sink assignments and LUT cuts under a final-area constraint.Replication, cut switching, and recovery are evaluated by the resulting live LUT count.We derive a necessary all-critical-path condition for delay reduction and obtain minimum-copy proposals by minimum cut within a restricted same-cut family.These proposals complement greedy refinement, with mappings selected by full-network area and timing.Experiments on EPFL benchmarks with three initial mappers achieve geometric-mean reductions of up to 14.48% in fanout-aware area--timing product (ATP) with a 5% final-area allowance.VPR evaluation shows a 7.04% reduction in routed critical-path delay relative to ABC if.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑