Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
机构 * Department of XXX, University of YYY, Location, Country(XXX系,YYY大学,地点,国家) ; School of ZZZ, Institute of WWW, Location, Country(ZZZ学院,WWW研究所,地点,国家)
Comments 6 pages, 2 figures, 2 tables. Uses ICML 2025 style