8 张卡、2 个节点。逐步看一层:QKV 切片 → AllReduce, 再 MoE Router + All-to-All dispatch / combine + Residual Add,最后 DP 再开一组。 Ring Attention 在 ring_attention.html。
8 张卡待命。点下一步,一条请求进 GPU0。
Ring AllReduce:T ≈ 2(N−1)α + 2(N−1)/N · n/β(Hockney;NCCL ring 把 AllReduce 拆成 ReduceScatter + AllGather)。
Decode 每层 AllReduce 的 payload n 很小,时间由延迟项 2(N−1)α 主导,增大 TP 即增大 N,单步线性变慢。
全链路见 journey.html。