面向DOM源粒子特效的浏览器管线架构分解:工作线程卸载、WebGL与WebAssembly
Decomposing Browser Pipeline Architectures for DOM-Sourced Particle Effects: Worker Offload, WebGL, and WebAssembly
浏览论文内容
中文总结 AI 辅助
本文通过分解5种粒子管线及基准,测试了Web Workers、WebGL、WebAssembly在不同浏览器/GPU/主机配置下的性能,得出工作线程卸载等关键结果,提出瓶颈感知的架构测量方法。
中文摘要 AI 辅助
在分层式浏览器架构中,局部优化并不一定能实现端到端优化。团队常将Web Workers、WebGL和WebAssembly视为可互换的“提速”手段,但三者分别针对不同架构层。本文以DOM源粒子管线为具体 workload,提出一种受控架构分解方案,涵盖5种粒子管线(P1-P5)及额外的非粒子CSS层基准(P0),用于隔离线程放置、渲染器选择与模拟后端。借助可复现测试 harness,我们测量了交互 pacing、高负载端到端压力测试、纯模拟微基准,以及同主机跨浏览器/双GPU级别的复现(Firefox 153+Intel UHD;Chrome 138+NVIDIA NVK),还在Google Colab上进行了独立的第二主机切片测试(Chrome 150,Tesla T4,n=5)。四项关键结果:(1)工作线程卸载可提升Chrome(以主线程为主)的交互 pacing,FPS约为144,对比基准约52,Cliff's delta=1,n=10;(2)AssemblyScript可将Chrome更新内核的速度在主主机上提升约1.5-1.6倍,在Colab T4的纯模拟场景下提升约1.85倍;(3)在约25万WebGL粒子时,Chrome纯模拟场景下的WASM优势不会提升产品FPS(主主机/Firefox中P5≤P4;Colab T4中P5≈P4);工作线程内的CPU阶段计时器显示,模拟仍占主导预算,因此非翻译场景并非简单的“绘制主导CPU”;(4)渲染器排名与绝对差值随浏览器/GPU/主机配置变化。本文的贡献是瓶颈感知的架构测量:识别主导层并测试该层的优化是否会传递到用户可见的FPS。
英文摘要
Local optimization does not necessarily yield end-to-end optimization in layered browser architectures. Teams often treat Web Workers, WebGL, and WebAssembly as interchangeable ways to "make it faster," yet each lever targets a different layer. We present a controlled architectural decomposition--using DOM-sourced particle pipelines as a concrete workload--of five particle pipelines (P1-P5), with an additional non-particle CSS-layer baseline (P0), that isolates thread placement, renderer choice, and simulation backend. Using a reproducible harness we measure interactive pacing, high-load end-to-end stress, and a simulation-only microbenchmark, plus same-host cross-browser / dual-GPU-class replication (Firefox 153+Intel UHD; Chrome 138+NVIDIA NVK) and an independent second-host slice on Google Colab (Chrome 150, Tesla T4, n=5). Four results stand out. (1) Worker offload improves interactive pacing on paper-primary Chrome (approx. 144 vs approx. 52 FPS; Cliff's delta=1, n=10). (2) AssemblyScript speeds the Chrome update kernel by about 1.5-1.6x on the primary host and approx. 1.85x on Colab T4 (sim-only). (3) Under approx. 250k WebGL particles the Chrome sim-only WASM win need not raise product FPS (P5<=P4 on primary/Firefox; P5 approx. P4 on Colab T4); in-worker CPU phase timers show simulate still dominates the accounted budget, so the non-translation is not a simple "draw dominates CPU" story. (4) Renderer ranking and absolute margins vary across browser/GPU/host configurations. The contribution is bottleneck-aware architectural measurement: identify the dominant layer and test whether a layer win propagates to user-visible FPS.