Fovea:具备物理含义感知的晶圆级设计空间探索,结合决策域引导的跨保真度优化
Fovea: Physical-Implication-Aware Wafer-Scale DSE with Decision-Domain-Guided Cross-Fidelity Refinement
浏览论文内容
中文总结 AI 辅助
Fovea是一种晶圆级设计空间探索方法,通过物理含义感知的空间构建与决策域引导的跨保真度优化,在LLM训练工作负载的晶圆架构选择中实现了最优解恢复与显著加速。
中文摘要 AI 辅助
现代硅前设计空间探索(DSE)采用从粗到细的工作流:低成本评估器筛选候选空间,详细评估仅针对短名单。晶圆级系统对这两个阶段均造成压力。架构选择会产生与掩模版合规性、晶圆拼接、裸片面积、裸片间(D2D)能力、边界访问和布局相关的耦合物理含义,因此设计空间不能视为无约束的笛卡尔积。同时,详细评估成本过高,无法覆盖所得空间,而分析与参考排名的反转使得固定短名单不可靠。我们提出Fovea,这是一种针对特定工作负载的晶圆架构选择的可复用方法,而非固定晶圆模板。Fovea首先执行物理含义感知的设计空间构建,以构建独特的建模可行空间,同时保留跨维度权衡并仅应用保留评估器的局部缩减。随后执行决策域引导的跨保真度优化。配对域内校准估计特定工作负载和空间的分析与参考的不一致性,这会参数化与参考一致的性能区间,并诱导决策域以进行选择性指定参考评估。在有效的域范围不一致性边界下,该域包含指定参考的最优解;基于采样的实现通过穷举参考设计空间进行实证评估。在10个可参考验证的设计空间和7个大语言模型(LLM)训练工作负载上,采用10%配对校准的Fovea在所有70个评估对中均恢复了穷举指定参考的最优解,同时实现了平均4.13倍、最大7.80倍的端到端加速。
英文摘要
Modern pre-silicon design-space exploration (DSE) follows a coarse-to-fine workflow: low-cost evaluators screen candidate spaces, while detailed evaluation is reserved for a shortlist. Wafer-scale systems strain both stages. Architecture choices induce coupled physical implications for reticle compliance, wafer tiling, die area, D2D capability, boundary access, and placement, so the design space cannot be treated as an unconstrained Cartesian product. Meanwhile, detailed evaluation is too expensive to cover the resulting space, whereas analytical-to-reference ranking inversions make a fixed shortlist unreliable. We present Fovea, a reusable methodology for workload-specific wafer architecture selection rather than a fixed wafer template. Fovea first performs physical-implication-aware design-space formulation to construct a distinct modeled-feasible space while preserving cross-dimensional trade-offs and applying only evaluator-preserving local reductions. It then performs Decision-Domain-guided cross-fidelity refinement. Paired in-domain calibration estimates workload- and space-specific analytical-to-reference disagreement, which parameterizes reference-consistent performance intervals and induces a Decision Domain for selective designated-reference evaluation. Under a valid domain-wide disagreement bound, this domain contains the designated-reference optimum; the sampling-based implementation is evaluated empirically on exhaustive-reference design spaces. Across ten reference-verifiable design spaces and seven LLM-training workloads, Fovea with 10% paired calibration recovers the exhaustive designated-reference optimum in all 70 evaluated pairs while achieving 4.13x average and 7.80x maximum end-to-end speedup.