arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无每资源RHI跟踪的全局传递屏障:基于Blade的跨厂商研究

Global Pass Barriers Without Per-Resource RHI Tracking: A Cross-Vendor Study with Blade

Dzmitry Malyshau

arXiv 2607.26506首次发表:更新:

AI 中文总结

该研究提出无每资源RHI跟踪的Blade方案,通过全局传递屏障减少冗余,在多款GPU上降低了GPU跨度,为图形API优化提供了跨厂商参考。

AI 中文摘要

显式图形API会暴露内存依赖、每资源访问及图像布局。wgpu会重构并验证该状态;Blade则将Vulkan图像保持为GENERAL布局,不跟踪任何每资源状态,而是发布全局传递边界屏障。我们在Blade内部隔离屏障放置及阶段/访问范围,端到端对比匹配的wgpu程序,测量来自4家厂商的6款GPU,其中包含探索性的Apple/Metal结果。从16个独立计算传递中移除15个冗余屏障,可使RTX 5070的GPU跨度降低29.3%,RX 7900 XT的GPU跨度降低32.3%;在同一款独立GPU上,还可使独立渲染跨度分别降低32.4%和7.3%,但Radeon 780M在32个传递时跨度增加42.4%,超出12.3%的特定计数稳定性下限。无依赖链放置效应可满足本研究的特定单元稳定性标准。从每个边界周围的传递类型推导全局屏障范围,无需跟踪任何资源,可在NVIDIA图形链上节省5.0%,在AMD计算链上节省6.7%;未解决任何涉及AMD渲染的范围单元问题。在所有测量单元中,wgpu的记录与提交成本更高,但该端到端差异不能仅归因于跟踪。RADV源码解释了为何命令数无法预测这些成本:广泛的全局依赖会扩展为多个刷新和无效请求,驱动程序可能会部分省略这些请求。源码还显示,在测量的RDNA条件下,持久GENERAL布局会保留DCC,但会禁用FMASK。由此得到的Blade方向是:在无跟踪的RHI中采用轻量聚合传递类型状态:上游渲染图选择全局依赖切割,而别名、跨队列使用、特殊布局及任意每资源DAG边仍属于资源感知引擎的职责。

英文摘要

Explicit graphics APIs expose memory dependencies, per-resource accesses, and image layouts. wgpu reconstructs and validates this state; Blade keeps Vulkan images in GENERAL, tracks no per-resource state, and issues global pass-boundary barriers. We isolate barrier placement and stage/access scope within Blade, compare matched wgpu programs end-to-end, and measure six GPUs from four vendors, including exploratory Apple/Metal results. Removing fifteen redundant barriers from sixteen independent compute passes reduces GPU span by 29.3% on an RTX 5070 and 32.3% on an RX 7900 XT. It reduces independent-render span by 32.4% and 7.3% on the same discrete parts, but increases Radeon 780M span by 42.4% at 32 passes, beyond a 12.3% count-specific stability floor. No dependent-chain placement effect clears the study's cell-specific stability criterion. Deriving global barrier scope from the pass kinds around each boundary, without tracking any resource, saves 5.0% on an NVIDIA graphics chain and 6.7% on an AMD compute chain; no AMD render-involving scope cell resolves. wgpu's record-and-submit cost is higher in every measured cell, but this end-to-end difference cannot be attributed to tracking alone. RADV source shows why command counts do not predict these costs: broad global dependencies expand to several flush and invalidate requests that the driver may partly elide. It also shows that persistent GENERAL retains DCC under the measured RDNA conditions but disables FMASK. The resulting Blade direction is lightweight aggregate pass-kind state in a tracking-free RHI: an upstream render graph selects global dependency cuts, while aliasing, cross-queue use, exceptional layouts, and arbitrary per-resource DAG edges remain resource-aware engine responsibilities.

Comments26 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑