AI 中文总结
本文提出将随机计算视为密集自适应量化器,通过比特流长度控制精度,实现按需精度的混合精度推理,并在视觉任务上验证其与固定格式量化相当的性能。
AI 中文摘要
矩阵乘法主导了现代基于Transformer的视觉模型的推理成本,然而现有的效率技术,如训练后量化和混合精度推理,在很大程度上局限于传统加速器所支持的少量固定宽度格式(INT4、INT8、BF16和FP16)。我们重新审视随机计算(SC),以此解除这一限制:将SC视为一种密集自适应量化器,它通过比特流长度L来控制精度,而非固定的数据路径,同时每次乘法仅简化为一个AND/XNOR门。我们构建了一个GPU库,大规模模拟SC矩阵乘法,将流长度作为一等内核参数暴露,并在图像分类、目标检测和实例分割、类条件图像生成以及视觉世界模型规划上端到端评估SC。在此基础之上,我们开发了一种动态的逐行混合精度策略,该策略在匹配的平均预算下为每个token或组分配流长度,无需重新训练,并在不同调度中使用相同的SC硬件。在各种任务中,在匹配的比特预算下,SC与固定格式INT量化相比保持竞争力,而逐行混合精度有助于在较低的平均流长度下维持准确性。这些结果提供了软件层面的可行性证据,表明SC可以作为现代视觉Transformer上细粒度混合精度推理的密集精度基础。
英文摘要
Matrix multiplications dominate the inference cost of modern transformer-based vision models, yet existing efficiency techniques such as post-training quantization and mixed-precision inference are largely limited to the small set of fixed-width formats (INT4, INT8, BF16, and FP16) supported by conventional accelerators. We revisit stochastic computing (SC) as a way to lift this constraint: viewed as a dense adaptive quantizer, SC controls precision by bit-stream length L rather than a fixed datapath, while each multiplication reduces to a single AND/XNOR gate. We build a GPU library that emulates SC matrix multiplication at scale, exposes stream lengths as first-class kernel arguments, and evaluates SC end-to-end on image classification, object detection and instance segmentation, class-conditional image generation, and visual world-model planning. On top of this substrate, we develop a dynamic per-row mixed-precision policy that assigns stream length per token or group at matched average budget, requires no retraining, and uses the same SC hardware across schedules. Across tasks, SC remains competitive with fixed-format INT quantization at matched bit budgets, while per-row mixed precision helps maintain accuracy at lower average stream lengths. These results provide software-level feasibility evidence that SC can serve as a dense-precision substrate for fine-grained mixed-precision inference on modern vision transformers.
CommentsNeurIPS 2026