arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SemBridge:将消费者观察编译为跨栈通信计划

SemBridge: Compiling Consumer Observations into Cross-Stack Communication Plans

Genlang Chen, Junyi Zhu, Yuanshan Lin

arXiv 2609.08231首次发表:更新:

发表机构

NingboTech University; Dalian Ocean University(宁波工程学院; 大连海洋大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SemBridge将图与运行时事实编译为类型化契约,生成并验证跨栈通信计划,显著减少流量并提升吞吐,确立消费者观察为新的语义层。

AI 中文摘要

分布式张量系统指定值所在的位置,而集合通信系统优化请求操作的执行方式。在无法共享原生通信器的供应商运行时之间的边界上,两种抽象均未说明远程消费者必须观察到什么。SemBridge通过将图和运行时事实编译为针对消费者可见结果及其交付义务的类型化契约来填补这一空白。该契约捕获来源、可替换性、完成性、权威性、需求性和原生域局部性。一个确定性的降级器构建后端无关的通信计划,一个符号检查器在执行前验证每个计划,涵盖CUDA/NCCL和CANN/HCCL。一个独立的仅布局规划器处理所有72种结构转换,但仅建立54种完整义务;一个仅字节的最小化器提出40个语义无效的候选方案,全部被SemBridge拒绝。在九个真实边上,SemBridge生成不同的观察感知计划,在所有九个边上减少启动次数,并在三个结果边上减少负载字节数。一次实时的CUDA/CANN运行从对数概率、仅令牌和所有者作用域请求中推导并执行全逻辑重建、源投影和所有者令牌交付。在实测的双主机1-GbE容量溢出部署中,源投影将结果流量减少超过99.97%,并在Dense、MoE和MiniMax工作负载上将吞吐量提高8.92%-80.20%。在并发度为1、8和16的所有18个MiniMax重启对中,源投影均占优。一个Qwen3-14B MLP切片还验证了按位激活分片交付和行并行部分和的HCCL完成。这些结果确立了消费者观察作为放置与集合执行之间的语义层。

英文摘要

Distributed-tensor systems specify where values reside, while collective systems optimize how requested operations execute. At a boundary between vendor runtimes that cannot share a native communicator, neither abstraction states what a remote consumer must observe. SemBridge fills this gap by compiling graph and runtime facts into a typed contract for the consumer-visible result and its delivery obligations. The contract captures provenance, substitutability, completion, authority, demand, and native-domain locality. A deterministic lowerer constructs backend-neutral communication plans, and a symbolic checker validates each plan before execution across CUDA/NCCL and CANN/HCCL. An independent layout-only planner handles all 72 structural transitions but establishes only 54 complete obligations; a byte-only minimizer proposes 40 semantically invalid candidates, all rejected by SemBridge. On nine real edges, SemBridge produces distinct observation-aware plans that reduce startups on all nine and payload bytes on the three result edges. A live CUDA/CANN run derives and executes full-logit reconstruction, source projection, and owner-token delivery from log-probability, token-only, and owner-scoped requests. On a measured two-host 1-GbE capacity-spillover deployment, source projection cuts result traffic by more than 99.97% and increases throughput by 8.92-80.20% across Dense, MoE, and MiniMax workloads. All 18 MiniMax restart pairs at concurrency 1, 8, and 16 favor source projection. A Qwen3-14B MLP slice additionally verifies bitwise activation-shard delivery and HCCL completion of row-parallel partials. These results establish consumer observation as a semantic layer between placement and collective execution.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑