arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ParaWeb:面向Web开发的并行编程模式

ParaWeb: Parallel Programming Patterns for Web Development

Suejb Memeti

arXiv 2608.19935首次发表:更新:

AI 中文总结

本文提出TypeScript库ParaWeb,为Web开发实现十种并行编程模式,含三种变体,实验显示其CPU变体最高加速11.6倍、GPU变体最高达414倍加速,性能优异。

AI 中文摘要

现代Web应用日益需要计算密集型处理,但作为Web主流语言的JavaScript传统上受限于单线程执行模型。HTML5 Worker Threads和浏览器Web Workers提供了并行执行的底层机制,但开发者缺乏能将重复并行结构封装为可复用模式的高层抽象。本文提出ParaWeb,这是一个TypeScript库,为服务器端Node.js、客户端浏览器环境及WebGPU计算着色器实现了十种并行编程模式。ParaWeb为每种模式提供三种实现变体:基于postMessage结构化克隆的消息传递(MP)变体、使用SharedArrayBuffer与类型化数组视图的共享缓冲区(Shared)变体,以及利用WebGPU计算着色器实现硬件加速执行的GPU变体。本文阐述了架构、设计决策及模式特定的实现策略,并在三种数据规模下评估了全部三十种实现的性能。实验评估结果显示,基于CPU的变体在计算密集型模式下使用16线程可实现最高11.6倍的加速,而GPU变体在算术强度较高的计算密集型模式(如Farm、Scatter、Reduce和Map)下可实现最高260倍的加速。针对五种图像卷积过滤器的案例研究进一步表明,GPU加速在非可分内核上相比单线程CPU可达到最高414倍的加速,且在1024×1024、2048×2048及4K图像上呈现出一致的扩展性。

英文摘要

Modern web applications increasingly require computationally intensive processing, yet JavaScript, the dominant language of the web, has traditionally been limited to a single-threaded execution model. Node.js Worker Threads and browser Web Workers provide low-level mechanisms for parallel execution, but developers lack high-level abstractions that capture recurring parallel structures as reusable patterns. In this paper, we present ParaWeb, a TypeScript library that implements ten parallel programming patterns for server-side Node.js, client-side browser environments, and WebGPU compute shaders. ParaWeb provides three implementation variants for each pattern: a message-passing (MP) variant based on structured cloning via postMessage, a shared-buffer (Shared) variant that uses SharedArrayBuffer with typed array views, and a GPU variant that uses WebGPU compute shaders for hardware-accelerated execution. We describe the architecture, design decisions, and pattern-specific implementation strategies, and we evaluate the performance of all thirty implementations across three data sizes. Experimental evaluation results show that the CPU-based variants achieve speedups of up to 11.6x with 16 threads for compute-bound patterns, while the GPU variants reach speedups of up to 260x for compute-bound patterns with high arithmetic intensity such as Farm, Scatter, Reduce, and Map. A case study on five image-convolution filters further shows that GPU acceleration reaches up to 414x speedup over single-threaded CPU on non-separable kernels, with consistent scaling across 1024x1024$, 2048x2048$, and 4K images.

CommentsPresented at HLPP 2026, 19th International Symposium on High-Level Parallel Programming and Applications, Paris, July 2026. Part of the HLPP 2026 proceedings (hal-05689350, arXiv:2607.12917)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑