路由-块成员选择打包AWQ算术:受控单一固定装置机制研究
Route-Block Membership Selects Packed-AWQ Arithmetic: A Controlled Single-Fixture Mechanism Study
浏览论文内容
中文总结 AI 辅助
该研究在Qwen3-Coder AWQ第6层固定装置上,证实路由-块成员可选择打包AWQ算术轨迹,揭示了MoE推理中路由与块对齐的因果机制。
中文摘要 AI 辅助
混合专家(MoE)推理首先将路由后的令牌对齐为填充的专家块,然后在这些块上执行打包量化矩阵乘法,此预处理常被视为簿记工作。在固定vLLM/Marlin构建的预指定Qwen3-Coder AWQ第6层固定装置、RTX 3090运行时环境中,我们展示了被测路由-块干预选择了精确的打包算术轨迹:两个固定预构建历史产生了不同的原生对齐和精确轨迹;注入相反对齐会转移W13、激活值、路由W2及最终输出;在一个块内置换两条路由保留各原生轨迹,而在专家106的40和41块边界交换两条先前数据选定的路由会转移完全相反的轨迹;源自源代码和二进制的调度几何将这些块映射为直接/全K和拆分/全局归约类;强制单切片200块网格使W13按位相等;稳定的规范构建使两个历史收敛到第三条精确轨迹;验证队列包含70个有效冷进程和7个必需的扰动拒绝。这是单一固定装置的因果机制结果,并非关于普遍性、分配器、可移植性或服务影响的主张。
英文摘要
Mixture-of-experts (MoE) inference first aligns routed tokens into padded expert blocks, then executes packed quantized matrix multiplication over those blocks. This preprocessing is often treated as bookkeeping. In one pre-specified Qwen3-Coder AWQ layer-6 fixture on a pinned vLLM/Marlin build and RTX 3090 runtime, we show that the tested route-block interventions select exact packed arithmetic trajectories. Two fixed preconstruction histories produced distinct native alignments and exact trajectories. Injecting the opposite alignment transferred W13, activation, routed-W2, and final outputs. Permuting two routes within one block preserved each native trajectory, while exchanging two prior-data-selected routes across the boundary between expert-106 blocks 40 and 41 transferred the complete opposite trajectory. Source- and binary-derived schedule geometry maps those blocks to direct/full-K and split/global-reduction classes. Forcing a single-slice 200-block grid made W13 bitwise equal. Stable canonical construction made both histories converge to a third exact trajectory. The confirmatory cohort contains 70 valid cold processes and seven required perturbation rejections. This is a causal mechanism result for one fixture, not a prevalence, allocator, portability, or serving-impact claim.