Experiment 005 · DeepSeek-V4.1 Flash on SGLangSGLang 上的 DeepSeek-V4.1 Flash

GB300 vs MI355X, one token at a time

We reproduced SGLang's GB300 numbers and measured where MI355X's 2× per step comes from. Turning off GB300's side streams takes its 4.61 ms verify cycle to 7.34 ms, against MI355X's 9.15 ms: 60% of the gap is stream concurrency, which HIP graphs make expensive, and most of the rest is MoE and attention kernel time. A 30% swing on GB300 turned out to be CPU placement.

我们复现了 SGLang 的 GB300 数据, 并实测了 MI355X 每步慢 2 倍的原因。 关掉 GB300 的侧 stream 后, 它 4.61 ms 的 verify 周期变成 7.34 ms, MI355X 是 9.15 ms: 差距里 60% 是 stream 并发, 而 HIP graph 让并发变得很贵;其余大部分是 MoE 和 attention 的 kernel 耗时。 GB300 上 30% 的波动, 最后查明是 CPU 放置造成的。

012345678910ms per DSpark verify cycle, BS=1GB3004.61 msGB300, one stream7.34 msMI355X9.15 ms
Same checkpoint, prompt and client on both machines. GB300 is NUMA-bound, at its default stream layout and with every side stream off; MI355X runs at SGLang's default placement. Plain decode: 3.53, 5.65 and 6.65 ms per step.两台机器用同一个 checkpoint、同一个 prompt 和同一个客户端。 GB300 绑定 NUMA, 分别为默认 stream 布局和关闭所有侧 stream;MI355X 为 SGLang 默认放置。 plain decode 每步分别是 3.53、5.65 和 6.65 ms。
modelDeepSeek-V4.1 Flash @ dba1be0aDeepSeek-V4.1 Flash @ dba1be0a
requestBS=1 · 4,096 random ids · 1,024 outBS=1 · 4,096 个随机 id · 输出 1,024
parallelismTP4 / EP4 · cookbook cellsTP4 / EP4 · cookbook cell
GB3004× GB300 (1 NVL72 tray) · CUDA 13.2 · main ffac53d7794× GB300(1 个 NVL72 托盘)· CUDA 13.2 · main ffac53d779
MI355X4× MI355X · ROCm 7.2 · dsv41-amd-main e2e824dc584× MI355X · ROCm 7.2 · dsv41-amd-main e2e824dc58
runs70 timed launches · 2 real-text launches · 4 nsys profiles · 2026-09-24/2570 次计时启动 · 2 次真实文本启动 · 4 个 nsys profile · 2026-09-24/25

Prologue序

SGLang's DeepSeek-V4.1 Flash numbers were produced on 4×GB300. Before borrowing anything from that work we wanted three answers: how far MI355X is behind when the request, the client and the metric are identical; which part of the difference is the machine and which part is software we can port; and which parts of the published numbers are measurement artifacts. We reran the published GB300 protocol on a GB300 tray, profiled it with nsys, replayed the same prompt with the same client on four MI355X GPUs, and then went back to the GB300 to measure what its stream concurrency is worth.

SGLang 的 DeepSeek-V4.1 Flash 数据来自 4×GB300。 在借鉴这些工作之前, 我们想先回答三个问题: 请求、客户端和指标完全相同时, MI355X 落后多少;差距里哪部分来自机器, 哪部分是可以移植的软件;公开数字里哪些是测量方式造成的。 为此我们在一个 GB300 托盘上重跑了公开的测试流程并用 nsys 做了 profile, 再用同一个客户端、同一个 prompt 在四张 MI355X 上重放, 最后回到 GB300 上实测它的 stream 并发到底值多少。

Conclusions结论

The three questions above, answered by the measurements in sections 1 to 8.

上面三个问题的答案, 依据是第 1 到第 8 节的测量。

metric (BS=1, 4,096 in / 1,024 out)指标(BS=1, 输入 4,096 / 输出 1,024)GB300MI355Xfaster更快的一方
Plain decode stepplain decode 每步3.53 ms6.65 msGB300 1.88×
DSpark verify cycleDSpark verify 周期4.61 ms9.15 msGB300 1.98×
Output tokens/s, published protocol (sim 5.5)输出 tokens/s, 公开流程(sim 5.5)1,190.1603.1GB300 1.97×
Verify cycle, GB300 with every side stream offverify 周期, GB300 关闭所有侧 stream7.34 ms9.15 msGB300 1.25×
Plain decode step, GB300 with every side stream offplain decode 每步, GB300 关闭所有侧 stream5.65 ms6.65 msGB300 1.18×
Accepted tokens per verify, real text真实文本上每次 verify 接受的 token 数3.173.23–
4,096-token prefill (TTFT)4,096 token prefill(TTFT)208 ms143 msMI355X 1.45×
Graph cost of one dependent tiny kernelgraph 里一个相互依赖的微小 kernel 的开销0.64 µs1.63 µsGB300 2.54×
Graph cost per kernel with fork/join branches带 fork/join 分支时每个 kernel 的 graph 开销0.64 µs7.12 µsGB300 11.1×

How far behind is MI355X?MI355X 落后多少?

About 2× per step at BS=1, and ahead at prefill. With the same request, client and metric, MI355X needs 6.65 ms per plain decode step against 3.53 ms and 9.15 ms per DSpark verify cycle against 4.61 ms. The verify cost does not depend on the acceptance mode and moves by at most 0.5% with CPU placement on MI355X, so the factor belongs to the machines and their software, not to the protocol. On real text the two machines accept 3.17 and 3.23 tokens per verify, so acceptance does not change the picture. MI355X answers the 4,096-token prefill 1.45× faster because its cell captures breakable prefill graphs.

BS=1 下每步大约慢 2 倍, 但 prefill 更快。 请求、客户端和指标相同时, MI355X 的 plain decode 每步 6.65 ms, GB300 3.53 ms;DSpark 每个 verify 周期 9.15 ms 对 4.61 ms。 verify 的耗时与接受方式无关, MI355X 上 CPU 放置对它的影响最多 0.5%, 所以这个倍数属于机器和它的软件, 而不是测试流程。 真实文本上两台机器每次 verify 分别接受 3.17 和 3.23 个 token, 接受长度不改变结论。 MI355X 的 4,096 token prefill 快 1.45 倍, 因为它的 cell 捕获了 breakable prefill graph。

Where is the gap?差距在哪里?

The largest part is stream concurrency. With every side stream turned off, GB300's verify cycle rises from 4.61 to 7.34 ms, so 2.73 ms of the 4.53 ms gap (60%) is what GB300's overlapped layout buys (a little more than overlap alone, because without side streams SGLang also launches 196 more kernels per verify). 1.77 ms of it comes from the attention-preparation, mHC-statistics, routed-quantization and draft streams and 0.96 ms from running the shared experts beside the routed experts. Plain decode has the same shape: 2.11 of its 3.12 ms gap. MI355X cannot take this win today, because a HIP graph kernel behind a fork/join costs 7.1 µs against 0.64 µs in a CUDA graph (section 7).

最大的一块是 stream 并发。 关闭所有侧 stream 后, GB300 的 verify 周期从 4.61 ms 升到 7.34 ms, 所以 4.53 ms 差距里有 2.73 ms(60%)是 GB300 的重叠布局带来的(比纯粹的重叠收益略多, 因为关闭侧 stream 后 SGLang 每个 verify 还会多发射 196 个 kernel)。 其中 1.77 ms 来自 attention 准备、mHC 统计、routed 量化和 draft 这几条 stream, 0.96 ms 来自让 shared expert 与 routed expert 并行。 plain decode 的形状一样: 3.12 ms 差距里有 2.11 ms。 MI355X 今天拿不到这部分收益, 因为 HIP graph 里跟在 fork/join 后面的 kernel 每个要 7.1 µs, CUDA graph 里只要 0.64 µs(第 7 节)。

With both machines serial, 1.81 ms per verify remains. By category (MI355X ledger against the serialized GB300 profile): MoE experts + routing +782 µs, attention / indexer / KV +484 µs, DSpark draft +211 µs, all-reduce + MoE finalize +211 µs, dense GEMM + act. quant +167 µs, other glue +53 µs; GPU idle +183 µs; MI355X is ahead on mHC (−372 µs). Behind many of these rows is one cost: a dependent kernel takes at least 1.63 µs in a HIP graph against 0.64 µs in a CUDA graph, and a verify cycle with its draft runs 1,519 kernels on MI355X today against 1,406 on GB300.

两台机器都串行时, 每个 verify 周期还差 1.81 ms。 按类别(MI355X 账本对比串行化后的 GB300 profile): MoE 专家与路由 +782 µs, attention / indexer / KV +484 µs, DSpark draft +211 µs, all-reduce 与 MoE finalize +211 µs, dense GEMM 与激活量化 +167 µs, 其他胶水 kernel +53 µs;GPU 空闲 +183 µs;MI355X 领先的是mHC(−372 µs)。 很多行背后是同一个开销: 一个相互依赖的 kernel 在 HIP graph 里至少要 1.63 µs, 在 CUDA graph 里只要 0.64 µs, 而一个 verify 周期连同 draft, 今天在 MI355X 上要执行 1,519 个 kernel, GB300 上是 1,406 个。

What was measurement rather than machine?哪些是测量问题, 不是机器问题?

Three things, none of them a GPU property. The published target was two weeks stale: current main is 29% faster per verify than the blog's commit (1,190.1 against 874 tokens/s). A 30% launch-to-launch swing on GB300 was CPU placement (Docker without SYS_NICE drops SGLang's NUMA binding, and simulated acceptance adds host work). And natural acceptance on a random prompt measures which repetition loop each machine's numerics fall into (5.79 against 3.62 tokens per verify), while on real text the machines agree; compare cost per verify across machines.

三件事, 都不是 GPU 的属性。 公开的目标已经过时两周: 当前 main 每个 verify 比博客所用的 commit 快 29%(1,190.1 对 874 tokens/s)。 GB300 上同一配置不同启动之间 30% 的波动来自 CPU 放置(Docker 没有 SYS_NICE 时 SGLang 的 NUMA 绑定失效, 而模拟接受又增加了 host 工作)。 随机 prompt 上的真实接受长度, 测的是各机器的数值误差把生成带进了哪种重复循环(每次 verify 5.79 对 3.62 个 token), 在真实文本上两台机器一致;跨机器应当比较每次 verify 的耗时。

Where to start从哪里开始

  • Concurrency, 2.7 ms per verify. Ask the ROCm runtime for graph branches that cost what they cost on CUDA (7.1 against 0.64 µs per kernel in section 7's microbenchmark). Until then, make the side work cheap enough that overlap no longer matters: fold the attention preparation and mHC statistics into neighbouring kernels (1.77 ms of GB300's overlap, with the draft streams), and remove the separate shared-expert pass (0.96 ms). Both machines log that SGLang disables shared-expert fusion under EP4 because a rank holds only a slice of the routed experts; MoE TP4 lifts that restriction and deserves an A/B.
  • Kernel count. Each kernel removed from the verify graph saves at least 1.63 µs on MI355X, 2.5× what it saves on GB300, so fusion pays more here than it did for the CUDA work being ported. At 1,519 kernels per cycle the floor alone is 2.5 ms on MI355X against 0.9 ms on GB300.
  • Kernel time: MoE (+782 µs) and attention (+484 µs) first. Every other category is within 211 µs, and MI355X is already faster on mHC (−372 µs): GB300's mHC is slower on its own and hidden behind a side stream, so mHC is not what MI355X needs to borrow.
  • Host idle, 183 µs. kevin-mii/sglang#12 keeps speculative overlap scheduling on the forward stream and takes the cycle from 9.15 to 8.63 ms on this base.
  • Keep the prefill-graph advantage, and fix SGLANG_SET_CPU_AFFINITY for correctness; it does not move BS=1 speed.
  • 并发, 每个 verify 2.7 ms。请 ROCm runtime 让 graph 分支的开销与 CUDA 持平(第 7 节 microbenchmark 里每个 kernel 7.1 µs 对 0.64 µs)。 在那之前, 把侧 stream 上的工作做得足够便宜, 让重叠变得无关紧要: 把 attention 准备和 mHC 统计并进相邻的 kernel(连同 draft stream, 占 GB300 重叠收益的 1.77 ms), 并去掉单独的 shared expert 一趟(0.96 ms)。 两台机器的日志都显示, EP4 下每个 rank 只持有部分 routed expert, SGLang 因此关闭了 shared expert 融合;MoE TP4 没有这个限制, 值得做一次 A/B。
  • kernel 数量。从 verify graph 里每去掉一个 kernel, MI355X 上至少省 1.63 µs, 是 GB300 上的 2.5 倍, 所以融合在这里的回报比在被移植的 CUDA 工作里更高。 每个周期 1,519 个 kernel, 仅这个下限在 MI355X 上就是 2.5 ms, GB300 上是 0.9 ms。
  • kernel 耗时: 先看 MoE(+782 µs)和 attention(+484 µs)。其他每一类都在 211 µs 以内, 而 mHC 上 MI355X 已经更快(−372 µs): GB300 的 mHC 单独运行更慢, 只是被藏在了侧 stream 后面, 所以 mHC 不是 MI355X 需要向 GB300 借鉴的地方。
  • host 空闲, 183 µs。kevin-mii/sglang#12 把 speculative overlap 调度留在 forward stream 上, 在这个 base 上把周期从 9.15 ms 降到 8.63 ms。
  • 保留 prefill graph 的优势;为正确性修好 SGLANG_SET_CPU_AFFINITY, 它不影响 BS=1 的速度。

1What was held fixed1固定了什么

A cross-platform number means something only when everything that is not the platform is identical, and the metric does not depend on something the platform changes by accident. Both machines therefore ran the same checkpoint, the same 4,096 token ids, the same client code and the same timing formula:

跨平台的数字要有意义, 前提是平台以外的一切完全相同, 而且指标不依赖平台会顺带改变的东西。 所以两台机器用的是同一个 checkpoint、同样的 4,096 个 token id、同一份客户端代码和同一个计时公式:

  • checkpoint deepseek-ai/DeepSeek-V4.1-Flash@dba1be0a;
  • BBuf's random prompt: 4,096 ids drawn with seed 42, no chat template (prompt.json sha256 2331eb91…);
  • BBuf's client: flush the cache before every request, freeze GC once, temperature 0, ignore_eos, 1,024 output tokens, one discarded warm-up and six timed rounds per server;
  • a fresh server for every launch and at least two launches per configuration;
  • TP4 / EP4 with each platform's cookbook cell: Low-Latency (DSpark block 5, decode graph cap 64, memory fraction 0.8) or High-Throughput (DSpark off).
  • checkpoint deepseek-ai/DeepSeek-V4.1-Flash@dba1be0a;
  • BBuf 的随机 prompt: 用 seed 42 抽取的 4,096 个 id, 不套 chat 模板(prompt.json sha256 2331eb91…);
  • BBuf 的客户端: 每个请求前清缓存, 只 freeze 一次 GC, temperature 0, ignore_eos, 输出 1,024 个 token, 每个 server 丢弃一次热身、计时六轮;
  • 每次启动都是全新的 server, 每个配置至少启动两次;
  • TP4 / EP4, 用各平台自己的 cookbook cell: Low-Latency(DSpark block 5, decode graph 上限 64, 显存比例 0.8)或 High-Throughput(关闭 DSpark)。

Three modes per platform: off (plain decode), sim (static verify, acceptance forced to 5 or 6 tokens with mean 5.5, the published protocol) and real (natural acceptance).

每个平台三种模式: off(plain decode)、sim(static verify, 每步强制接受 5 或 6 个 token, 均值 5.5, 即公开的测试流程)和 real(真实接受)。

Not removed, and recorded instead: the hardware (four GB300 in one NVL72 tray with two Grace CPUs and NVLink-C2C, SM clock 2,070 MHz, 1,400 W limit; four MI355X on one socket of an eight-GPU node, 1,400 W), the stacks (CUDA 13.2 with SGLang main ffac53d779 in the dev-cu13 image; ROCm 7.2 with dsv41-amd-main@e2e824dc58, the branch of #39857, AITER acf8fdf9 and the int64 store-offset fix from kevin-mii/sglang#8), the auto-selected backends, and prefill: current main disables the prefill CUDA graph for DeepSeek-V4 on CUDA, while the MI355X cell captures breakable prefill graphs.

没有消除、改为记录下来的差异: 硬件(四张 GB300 在一个 NVL72 托盘里, 带两颗 Grace CPU 和 NVLink-C2C, SM 频率 2,070 MHz, 功耗上限 1,400 W;四张 MI355X 在一台八卡机器的同一个 socket 上, 1,400 W)、软件栈(CUDA 13.2, SGLang main ffac53d779, 镜像 dev-cu13;ROCm 7.2, dsv41-amd-main@e2e824dc58(即 #39857 的分支), AITER acf8fdf9, 外加 kevin-mii/sglang#8 的 int64 store 偏移修复)、自动选择的 backend, 以及 prefill: 当前 main 在 CUDA 上为 DeepSeek-V4 关掉了 prefill CUDA graph, 而 MI355X 的 cell 会捕获 breakable prefill graph。

2Calibration: reproduce the published GB300 numbers first2校准: 先复现公开的 GB300 数字

A gap measured against a misconfigured baseline is fiction, so the first GPU hours went to reproducing the published series with the author's own commit (835c3909), image (dev-dsv41), launchers and client.

和配置错误的 baseline 比出来的差距没有意义, 所以最先的 GPU 时间都花在复现公开数据上: 用作者自己的 commit(835c3909)、镜像(dev-dsv41)、启动脚本和客户端。

Plate ICalibration against the published GB300 series与公开 GB300 数据的校准
published (README)ours · unboundours · NUMA-bound02004006008001000tokens / s (BS=1, 4096 in / 1024 out)MoE TP4 · DSpark sim 5.5873.6870.6878.8MoE EP4 · DSpark sim 5.5854.6841.3848.5EP4 · DSpark off223.5223.9224.3
Two launches per bar, six rounds each. Unbound: −0.4%, −1.6%, +0.2% against the published medians; NUMA-bound: +0.6%, −0.7%, +0.4%.每根柱子两次启动、每次六轮。 不绑定时相对公开中位数为 −0.4%、−1.6%、+0.2%;NUMA 绑定时为 +0.6%、−0.7%、+0.4%。

Every configuration lands within 1.6% of its published median, and within 0.7% with NUMA binding. The harness is sound and the tray is not an outlier, which is what licenses the comparisons below.

每个配置都落在公开中位数的 1.6% 以内, 绑定 NUMA 后在 0.7% 以内。 说明测试环境可靠、这个托盘也不是异常机器, 后面的对比才站得住。

3The published target is two weeks stale3公开数据已经过时两周

On the same tray, current main (ffac53d779, 2026-09-23) against the author's 2026-09-11 commit: the plain decode step fell from 4.46 to 3.53 ms (−21%), and the DSpark verify cycle with the same EP4 layout and simulated acceptance from 6.48 to 4.61 ms (−29%). The blog's 874 tokens/s is no longer the number to beat: main runs 1,190.1 tokens/s under the identical protocol.

在同一个托盘上, 当前 main(ffac53d779, 2026-09-23)对比作者 2026-09-11 的 commit: plain decode 每步从 4.46 ms 降到 3.53 ms(−21%), 相同 EP4 布局、模拟接受下的 DSpark verify 周期从 6.48 ms 降到 4.61 ms(−29%)。 博客里的 874 tokens/s 已经不是要追的目标: 在完全相同的流程下, main 已经跑到 1,190.1 tokens/s。

Plate IICost per step across versions and platforms不同版本与平台的每步耗时
0246810milliseconds per step (lower is faster; NUMA-bound where available)Plain decode stepGB300 · 835c39094.46 msGB300 · main3.53 msMI355X · e2e824dc586.65 msDSpark verify cycleGB300 · 835c3909 (EP4)6.48 msGB300 · main4.61 msMI355X · e2e824dc589.15 ms
NUMA-bound GB300 arms; MI355X at SGLang's default placement. Old-commit rows use the author's launchers; main rows use the cookbook cells.GB300 取 NUMA 绑定的数据;MI355X 为 SGLang 默认放置。 旧 commit 用作者的启动脚本, main 用 cookbook cell。

The V4.1 work that reached main in between includes the integration series #39646, #39648, #39653 and #38798 (which carried the BS=1 decode, verify and communication kernels developed on the side branch), #39704 (mHC, metadata and router overhead), #39957 (inverse-RoPE + WO-A + MXFP8 fusion) and #40431 (FP4 indexer tile skipping). This page does not attribute the gain to individual PRs; the point is methodological: compare against current main measured on hardware you control, not against a blog figure.

这期间进入 main 的 V4.1 工作包括集成系列 #39646、#39648、#39653 和 #38798(带进了在侧分支上开发的 BS=1 decode、verify 和通信 kernel)、#39704(mHC、元数据和 router 开销)、#39957(inverse-RoPE + WO-A + MXFP8 融合)和 #40431(FP4 indexer 跳过不可见 tile)。 本文不把收益拆到单个 PR 上;要说明的是方法: 要和自己能控制的硬件上测出的当前 main 比, 而不是和博客里的数字比。

4A 30% swing that was not the GPU4与 GPU 无关的 30% 波动

The first GB300 pass had one configuration that would not repeat: the simulated-acceptance arm read 1,184.2 and then 958.2 tokens/s on two launches of an identical container. A control with three launches per side confirmed it: 899.7 / 995.0 / 908.8 tokens/s unbound against 1,192.9 / 1,191.3 / 1,183.0 with SGLang's NUMA binding (+30.5%). The same switch moved real acceptance by −0.3% and plain decode by +0.4%.

GB300 第一轮测试里有一个配置无法复现: 模拟接受的 arm 在两次启动完全相同的容器时, 分别得到 1,184.2 和 958.2 tokens/s。 每边各启动三次的对照证实了这一点: 不绑定时是 899.7 / 995.0 / 908.8 tokens/s, 打开 SGLang 的 NUMA 绑定后是 1,192.9 / 1,191.3 / 1,183.0(+30.5%)。 同一个开关对真实接受只改变了 −0.3%, 对 plain decode 只改变了 +0.4%。

Plate IIIGB300 sim rounds, unbound vs NUMA-bound; TTFTGB300 sim 各轮(不绑定对比 NUMA 绑定)与 TTFT
800900100011001200tokens / s per timed round (sim 5.5, main, 5 launches per lane)unboundmedian 933NUMA-boundmedian 1188TTFT, 4096-token prefill (ms)off412208real425206sim366212unboundbound
Each row of dots is one launch. GPU0 SM clock stayed at 2070 MHz in all six control launches; its median power was 401–425 W unbound and 461–474 W bound, i.e. the GPU was waiting, not throttling.每一行点是一次启动。 六次对照启动里 GPU0 的 SM 频率都是 2070 MHz;功耗中位数不绑定时为 401–425 W, 绑定后为 461–474 W, 说明 GPU 是在等待, 而不是降频。

Why would CPU placement move a GPU benchmark? The four TP ranks meet at an all-reduce in every layer, so whenever work launched from the host is on the critical path, the cycle runs at the pace of the slowest rank's CPU. Unbound, the Linux scheduler put most of TP2 and TP3 on the Grace that is not attached to their GPUs (16%–35% of their threads local, 9%–55% of their pages local). Bound, every rank ran 100% local with 93%–95% local pages. SGLang binds each scheduler to its GPU's node only if the process may call set_mempolicy; in Docker that needs --cap-add SYS_NICE, and without it SGLang logs a warning and runs unbound.

CPU 放置为什么会影响 GPU 测试?四个 TP rank 每一层都要在 all-reduce 处会合, 所以只要由 host 发起的工作处在关键路径上, 整个周期就会按最慢那个 rank 的 CPU 节奏走。 不绑定时, Linux 调度器把 TP2 和 TP3 的大部分线程放到了没有连接它们 GPU 的那颗 Grace 上(只有 16%–35% 的线程、9%–55% 的内存页在本地)。 绑定后, 每个 rank 的线程 100% 在本地, 内存页 93%–95% 在本地。 SGLang 只有在进程有权限调用 set_mempolicy 时才会把每个 scheduler 绑到它 GPU 所在的节点;在 Docker 里这需要 --cap-add SYS_NICE, 否则 SGLang 只打一条警告, 然后不绑定运行。

Why only sim, and why prefill? Real acceptance is resolved inside the captured verify graph. Simulated acceptance overrides it afterwards: apply_dflash_simulated_acceptance draws the forced length on the CPU, then launches fill_ and copy_ kernels and recomputes the sequence lengths eagerly, so every cycle gains a host-launched tail. Prefill is host-launched too, because main turns the breakable prefill graph off for DeepSeek-V4 on CUDA; its 4,096-token TTFT dropped from 412 to 208 ms with binding.

为什么只有 sim 受影响, 为什么 prefill 也受影响?真实接受在已捕获的 verify graph 内部完成。 模拟接受则在之后覆盖结果: apply_dflash_simulated_acceptance 先在 CPU 上抽出强制接受的长度, 再发射 fill_ 和 copy_ kernel, 并以 eager 方式重新计算序列长度, 所以每个周期都多出一段由 host 发起的尾巴。 prefill 同样由 host 逐个发射 kernel, 因为 main 在 CUDA 上为 DeepSeek-V4 关掉了 breakable prefill graph;绑定后 4,096 token 的 TTFT 从 412 ms 降到 208 ms。

Plate IVWhere the four schedulers run四个 scheduler 在哪里运行
GB300 tray · 2 × Grace (72 cores each) · NVLink-C2CCPU 0–71 · node 0GPU0 · TP0GPU1 · TP1CPU 72–143 · node 1GPU2 · TP2GPU3 · TP3unbound: TP2/TP3 threads on their GPU's node 16%–35%, their pages 9%–55%bound (--cap-add SYS_NICE): every rank 100% local threads, 93%–95% local pagesMI355X node · 2 sockets · GPUs 4–7 (our TP4) on node 1node 0 · CPU 0–63, 128–191GPUs 0–3 (unused)node 1 · CPU 64–127, 192–255GPUs 4–7GPU4 · TP0CPU 0–31GPU5 · TP1CPU 32–63GPU6 · TP2CPU 64–95GPU7 · TP3CPU 96–127SGLANG_SET_CPU_AFFINITY=1: rank r gets physical cores [32r, 32r+32)→ TP0/TP1 are pinned to node 0, remote from GPUs 4–7SGLang's own NUMA binding is CUDA/XPU-only; numactl is not installedlocal (same NUMA node as the GPU)remoteEvery layer ends in a TP all-reduce, so a host-launched step runs at the pace of the slowest rank's CPU.
Left: GB300 tray, placement measured from /proc (threads' last CPU and numa_maps). Right: the MI355X node as the container configures it, measured the same way.左: GB300 托盘, 放置情况从 /proc 读取(线程最近运行的 CPU 和 numa_maps)。 右: 容器默认配置下的 MI355X 节点, 用同样的方法测得。
machine · placement机器 · 放置TP0TP1TP2TP3
GB300 · unboundGB300 · 不绑定threads线程 76–90%
pages内存页 91–93%
threads线程 83–89%
pages内存页 90–93%
threads线程 22–24%
pages内存页 46–55%
threads线程 16–35%
pages内存页 9–11%
GB300 · NUMA-boundGB300 · NUMA 绑定threads线程 100%
pages内存页 93–94%
threads线程 100%
pages内存页 93–94%
threads线程 100%
pages内存页 95%
threads线程 100%
pages内存页 95%
MI355X · split (default)MI355X · 切分(默认)CPU 0-31
pages内存页 12–15%
CPU 32-63
pages内存页 9–15%
CPU 64-95
pages内存页 56–65%
CPU 96-127
pages内存页 59–64%
MI355X · node 1MI355X · node 1CPU 64-127
pages内存页 64–65%
CPU 64-127
pages内存页 64–65%
CPU 64-127
pages内存页 64–65%
CPU 64-127
pages内存页 64–65%
MI355X · unpinnedMI355X · 不固定CPU 0-255
pages内存页 36–64%
CPU 0-255
pages内存页 33–64%
CPU 0-255
pages内存页 21–64%
CPU 0-255
pages内存页 24–48%

Local = on the NUMA node of the rank's GPU (GB300: node 0 for TP0/TP1, node 1 for TP2/TP3; MI355X: node 1 for all four). GB300 threads are counted by the CPU each thread last ran on; page shares come from numa_maps and include the memory-mapped model files.

本地指位于该 rank 所用 GPU 的 NUMA 节点(GB300: TP0/TP1 为 node 0, TP2/TP3 为 node 1;MI355X: 四个 rank 都是 node 1)。 GB300 的线程按每个线程最近一次运行的 CPU 统计;内存页比例来自 numa_maps, 包含 mmap 进来的模型文件。

The same check on MI355X found two problems. SGLang's NUMA binding is CUDA/XPU-only (it returns early on ROCm), and the container sets SGLANG_SET_CPU_AFFINITY=1, whose arithmetic gives rank r physical cores 32r to 32r+31 regardless of which socket its GPU is on. Our TP4 runs on HIP devices 4–7, all on node 1, so TP0 and TP1 were pinned to node 0 with their memory there too; every number in our recent MI355X PRs was measured with two remote ranks.

在 MI355X 上做同样的检查, 发现了两个问题。 SGLang 的 NUMA 绑定只支持 CUDA/XPU(在 ROCm 上直接返回);而容器设置了 SGLANG_SET_CPU_AFFINITY=1, 它的算法不管 GPU 挂在哪个 socket 上, 都把物理核 32r 到 32r+31 分给第 r 个 rank。 我们的 TP4 跑在 HIP 设备 4–7 上, 全部位于 node 1, 于是 TP0 和 TP1 被钉在 node 0, 内存也在那里;我们近期 MI355X PR 里的每个数字, 都是在两个 rank 位于远端的状态下测的。

We then ran the placement A/B on MI355X in one session, mirrored in time: SGLang's split affinity, all four schedulers on node 1, and no pinning. Cost per verify cycle with simulated acceptance: 9.14 ms split, 9.14 ms node 1, 9.18 ms unpinned (node 1 0.0%, unpinned +0.5% against split); with real acceptance 9.13, 9.12 and 9.13 ms (−0.1% and 0.0%). Plain decode, against the split runs of the first session: 6.65 ms per step split, 6.66 node 1, 6.67 unpinned. Placement moves the MI355X BS=1 cycle by at most 0.5%. The host is still on the critical path here (the verify-to-draft wait that kevin-mii/sglang#12 removes), but that share of the cycle does not depend on which socket runs it, and sim costs the same as real acceptance on this machine. The affinity arithmetic is still wrong and should be fixed, but the 2x does not come from it.

随后我们在 MI355X 上用同一个 session、按时间对称的顺序做了放置对照: SGLang 的切分亲和性、四个 scheduler 全部放在 node 1、完全不绑定。 模拟接受时每个 verify 周期分别是 9.14、9.14、9.18 ms(相对切分, node 1 0.0%, 不绑定 +0.5%);真实接受时分别是 9.13、9.12、9.13 ms(−0.1% 和 0.0%)。 plain decode 以第一个 session 的切分结果为基准: 切分每步 6.65 ms, node 1 为 6.66 ms, 不绑定为 6.67 ms。 放置方式对 MI355X BS=1 周期的影响最多只有 0.5%。 host 在这里仍处在关键路径上(就是 kevin-mii/sglang#12 消掉的那段 verify 到 draft 的等待), 但这部分耗时和它跑在哪个 socket 上无关;在这台机器上sim 和真实接受的每周期耗时也相同。 这段亲和性算法仍然是错的, 应该修, 但 2 倍差距不来自它。

5Natural acceptance measures the prompt's attractor5真实接受长度测的是 prompt 的吸引子

With real acceptance GB300 posts 1,241.0 tokens/s, above its own simulated run, which looks like a free win for real acceptance until you divide by the acceptance length. GB300 accepts a median 5.79 tokens per verify (range 3.98–5.95) because most of its greedy continuations fall into loops (one 16-gram repeats up to 497 times); MI355X accepts a median 3.62 (range 2.39–5.28) with far less repetition (at most 139). Neither machine is run-to-run deterministic at temperature 0 on this prompt: identical requests return different text.

真实接受下 GB300 跑到 1,241.0 tokens/s, 比它自己的模拟接受还高, 看起来像是真实接受白送的收益, 但除以接受长度就不是了。 GB300 每次 verify 接受 token 数的中位数是 5.79(范围 3.98–5.95), 因为它的 greedy 续写大多掉进了循环(某个 16-gram 最多重复 497 次);MI355X 的中位数只有 3.62(范围 2.39–5.28), 重复少得多(最多 139 次)。 在这个 prompt 上, 两台机器在 temperature 0 下都不是逐轮确定的: 相同的请求会返回不同的文本。

Plate VAcceptance per round, and the cost of a verify每轮接受长度与每次 verify 的耗时
23456sim target 5.5accepted tokens per verify (real acceptance, one dot per timed round)GB300MI355Xcost per verify cycle (ms)GB300 real4.61GB300 sim4.61MI355X real9.13MI355X sim9.15
Real-acceptance rounds from every launch on each machine. The verify cost is the same with real or simulated acceptance on both machines; only the accepted length differs.每台机器所有启动的真实接受轮次。 两台机器上, 真实接受和模拟接受的每次 verify 耗时相同, 不同的只是接受长度。

The cost of one verify does not care: GB300 spends 4.61 ms per cycle with real acceptance and 4.61 with simulated; MI355X 9.13 and 9.15. A random prompt has no meaning, so tiny numeric differences decide which degenerate continuation wins. Comparing tokens/s under natural acceptance on it would credit one machine for its rounding. Acceptance is worth comparing only on real text, where it reflects the model.

每次 verify 的耗时并不受影响: GB300 真实接受时每周期 4.61 ms, 模拟接受时 4.61 ms;MI355X 分别是 9.13 和 9.15 ms。 随机 prompt 本身没有意义, 微小的数值差异就能决定模型走向哪种退化的续写。 在它上面比较真实接受的 tokens/s, 等于把某台机器的舍入方式算成了它的功劳。 接受长度只有在真实文本上才值得比较, 那时它反映的是模型本身。

So we replayed our MI355X real-text contract on GB300: the same four chat-encoded 4,096-token prompts, two unscored warm-ups and then 24 samples of each, the same request body and client code, two launches at the default stream layout. GB300 accepts a median 3.17 tokens per verify, MI355X 3.23 over its two base-arm servers (−1.9%; per prompt, GB300/MI355X: 3.06/3.08, 2.97/3.14, 3.64/3.71, 3.36/3.35). On real text the accepted length is a property of the model and the two machines agree, so tokens/s on real text follows the cost per verify. With host and scheduler time included, a verify on this text takes 4.71 ms on GB300 and 9.16 ms on MI355X, and decode runs at 670 against 353 tokens/s.

所以我们把 MI355X 的真实文本契约拿到 GB300 上重放: 同样四个 chat 编码的 4,096 token prompt, 先两次不计分的热身, 再每个 prompt 采样 24 次, 请求体和客户端代码相同, 默认 stream 布局下启动两次。 GB300 每次 verify 接受 token 数的中位数是 3.17, MI355X 两个 base arm server 合计是 3.23(−1.9%;按 prompt, GB300/MI355X: 3.06/3.08, 2.97/3.14, 3.64/3.71, 3.36/3.35)。 在真实文本上, 接受长度是模型的属性, 两台机器一致, 所以真实文本上的 tokens/s 取决于每次 verify 的耗时。 算上 host 和 scheduler 的时间, 这段文本上每次 verify GB300 要 4.71 ms, MI355X 要 9.16 ms;decode 速度分别是 670 和 353 tokens/s。

6Where 4.6 ms and 9 ms go64.6 ms 和 9 ms 各花在哪里

The GB300 profiles capture TP0 under nsys with CUDA-graph node tracing and split every cycle into the categories of our MI355X ledger. In the default layout streams overlap, so a category's kernel time there is not what it costs alone. We therefore also captured GB300 with every side stream off (section 7), where kernel time is standalone cost (99 verify cycles, overlap 1.05). The MI355X graph executes its kernels back to back; its ledger comes from our earlier BS=1 campaign on the previous base (8.92 ms per cycle against 9.15 ms today, genuine acceptance, profiled family shares scaled to unprofiled phase totals).

GB300 的 profile 在 nsys 下以 CUDA graph 节点粒度采集 TP0, 再把每个周期按我们 MI355X 账本的类别拆开。 默认布局下多个 stream 会并发, 所以那时某一类的 kernel 时长并不是它单独运行的开销。 因此我们又在关闭所有侧 stream 的条件下采了一次 GB300(第 7 节), 这时 kernel 时长就是单独运行的开销(99 个 verify 周期, 重叠系数 1.05)。 MI355X 的 graph 串行执行 kernel;它的账本来自我们在上一个 base 上做的 BS=1 实验(每周期 8.92 ms, 今天是 9.15 ms;真实接受;按 profile 得到的各族占比, 缩放到未开 profiler 时实测的阶段总时长)。

Plate VIAnatomy of one DSpark verify cycle一个 DSpark verify 周期的构成
0123456789milliseconds per verify cycle (TP0)MI355X · serial8,921 µs, prior baseGB300 · one stream7,195 µs meanGB300 · default streams4,644 µs, wall shareMoE experts + routingdense GEMM + act. quantattention / indexer / KVmHCall-reduce + MoE finalizeEngramDSpark draftother glueGPU idleEach bar sums to its mean cycle. With one stream a category's share is its standalone kernel time; in the default layout overlapping kernels split their interval evenly.
Hatched: GPU idle. The one-stream bar is GB300 with every side stream off, so each segment is a standalone cost; the default bar splits the time of overlapping kernels evenly.斜线: GPU 空闲。 单 stream 那一条是关闭所有侧 stream 的 GB300, 每一段都是单独运行的开销;默认布局那一条把重叠 kernel 的时间平均分摊。
category类别MI355X µsGB300 one stream µsGB300 单 stream µsgap µs差距 µsGB300 default, wall µsGB300 默认布局, 墙钟 µskernels MI / GBkernel 数 MI / GB
MoE experts + routingMoE 专家与路由1,9631,180+782815288 / 240
dense GEMM + act. quantdense GEMM 与激活量化1,8931,726+1671,066344 / 407
attention / indexer / KVattention / indexer / KV1,164681+484584171 / 204
mHCmHC9381,311−372565161 / 240
all-reduce + MoE finalizeall-reduce 与 MoE finalize945734+211828127 / 87
EngramEngram2517+8185 / 6
DSpark draftDSpark draft790579+211418161 / 161
other glue其他胶水 kernel516463+53200106 / 257
GPU idleGPU 空闲687504+183149–
total合计8,9217,195+1,7264,6441363 / 1602

With GB300 serial, the gap column is a number per category. MoE is the largest (+782 µs), then attention (+484 µs); all-reduce with MoE finalize (+211), the draft (+211) and dense GEMM (+167) follow, and glue kernels are close (+53). MI355X is faster on mHC (−372 µs): GB300's mHC statistics are cheap only because a side stream hides them. The GPU-idle row (+183 µs against serial GB300, +538 against the default layout) is the host wait that our follow-up to #39857, kevin-mii/sglang#12, removes by keeping speculative overlap scheduling on the forward stream (9.15 to 8.63 ms per cycle on this base); #13 additionally builds the draft-block metadata inside the draft graph. The table's total gap is within 80 µs of the timed one (1.81 ms): the MI355X ledger is 8.92 ms from the previous base against 9.15 ms today, and the serial capture runs 7.18 ms under nsys against 7.34 ms untraced.

GB300 串行之后, 差距那一列每一类都是一个确定的数。 MoE 最大(+782 µs), 其次是 attention(+484 µs);all-reduce 加 MoE finalize(+211)、draft(+211)和 dense GEMM(+167)随后, 胶水 kernel 很接近(+53)。 mHC 上 MI355X 更快(−372 µs): GB300 的 mHC 统计之所以便宜, 只是因为被侧 stream 藏起来了。 GPU 空闲那一行(对比串行 GB300 为 +183 µs, 对比默认布局为 +538), 正是我们在 #39857 之上的后续改动 kevin-mii/sglang#12 要消掉的 host 等待: 它把 speculative overlap 调度留在 forward stream 上(在这个 base 上每周期从 9.15 ms 降到 8.63 ms);#13 则进一步把 draft-block metadata 放进 draft graph 里构建。 表中的总差距和计时得到的差距(1.81 ms)相差 80 µs 以内: MI355X 账本来自上一个 base, 8.92 ms, 今天是 9.15 ms;串行 capture 在 nsys 下是 7.18 ms, 不开 profiler 时是 7.34 ms。

We also tried to rebuild the MI355X ledger on the current base with a kernel trace instead of a torch profile. rocprofv3 1.1.0 aborts the first replay of a packet-captured HIP graph (HSA_STATUS_ERROR_INVALID_PACKET_FORMAT). With packet capture disabled the server runs, but 60% of traced kernels land between 4 and 6 µs whatever they do (median 5.2 µs), the kind of instrumentation floor that ruled out the torch profile, and the fused all-reduce reads 50 µs per call because it absorbs the rank skew that tracing adds. What the trace does settle is structure: the current base runs 1,519 kernels per verify with its draft, against 1,363 on the previous base and 1,406 on GB300, and its attention main pass is now partial_gluon with _combine, the kernels GB300 runs. The ledger above therefore stays on the previous base.

我们也尝试用 kernel trace 代替 torch profiler, 在当前 base 上重建 MI355X 账本。 rocprofv3 1.1.0 在 packet capture 的 HIP graph 第一次重放时就会中止(HSA_STATUS_ERROR_INVALID_PACKET_FORMAT)。 关掉 packet capture 后 server 能跑, 但 60% 的 kernel 不论做什么, trace 出来都落在 4 到 6 µs 之间(中位数 5.2 µs), 这正是当初否定 torch profiler 的那种计时下限;融合的 all-reduce 每次调用读出 50 µs, 因为它吸收了 trace 带来的 rank 间等待。 trace 能确定的是结构: 当前 base 每个 verify 连同 draft 要执行 1,519 个 kernel, 上一个 base 是 1,363 个, GB300 是 1,406 个;它的 attention 主计算已经换成 partial_gluon 加 _combine, 和 GB300 用的是同一组 kernel。 所以上面的账本仍然用上一个 base 的。

The GB300 plain-decode step in the same categories, with side streams on and off (200 steps, TP0):

GB300 plain decode 每一步按同样类别的拆分, 分别是打开和关闭侧 stream(200 步, TP0):

category (GB300 plain decode step)类别(GB300 plain decode 每步)one stream µs单 stream µsdefault, wall µs默认布局, 墙钟 µsdefault, kernel µs默认布局, kernel µskernelskernel 数
MoE experts + routingMoE 专家与路由8935441,076240
dense GEMM + act. quantdense GEMM 与激活量化1,6631,0911,732403
attention / indexer / KVattention / indexer / KV579570623197
mHCmHC1,0474561,032240
all-reduce + MoE finalizeall-reduce 与 MoE finalize62961865484
EngramEngram1515155
other glue其他胶水 kernel317148214205
GPU idleGPU 空闲518114––
total合计5,6613,5575,3461374

7What concurrency is worth, measured7并发值多少: 实测

A profile of the default layout can only bound the concurrency prize: GB300's kernels add up to 7.2 ms per verify but occupy 4.4 ms, so overlap hides at most 2.8 ms. To measure it, we turned GB300's side streams off and timed the same cells again, in one session, two launches per arm, with the order mirrored in time. SGLANG_OPT_USE_MULTI_STREAM_OVERLAP=0 removes the attention-preparation (KV, compressor, indexer), mHC-statistics, routed-quantization and DSpark-draft streams. The shared experts still run beside the routed experts on another stream that no switch controls, so the fully serial arm also loads a patch that clears it (gb300/scripts/wb/serial_moe). The default arm of this session reproduced the earlier NUMA-bound numbers (4.61 ms per verify, 3.54 ms per step).

默认布局下的 profile 只能给并发收益一个上限: GB300 的 kernel 时长加起来每个 verify 有 7.2 ms, 却只占 4.4 ms, 所以重叠最多藏掉 2.8 ms。 为了实测, 我们在 GB300 上关掉侧 stream, 在同一个 session 里重跑相同的 cell, 每个 arm 启动两次, 并按时间对称排序。 SGLANG_OPT_USE_MULTI_STREAM_OVERLAP=0 会去掉 attention 准备(KV、compressor、indexer)、mHC 统计、routed 量化和 DSpark draft 这几条 stream。 shared expert 仍然在另一条 stream 上与 routed expert 并行, 没有开关能关掉它, 所以完全串行的 arm 另外加载了一个补丁把这条 stream 清掉(gb300/scripts/wb/serial_moe)。 这个 session 的默认 arm 复现了之前 NUMA 绑定的数字(每个 verify 4.61 ms, 每步 3.54 ms)。

Plate VIIGB300 with its side streams turned off关闭侧 stream 后的 GB300
0246810milliseconds (lower is faster; GB300 session 3, NUMA-bound)DSpark verify cycle (sim 5.5)GB300 · default streams4.61 msGB300 · OPT_USE_MULTI_STREAM_OVERLAP=06.38 msGB300 · every side stream off7.34 msMI355X · e2e824dc589.15 msPlain decode stepGB300 · default streams3.54 msGB300 · OPT_USE_MULTI_STREAM_OVERLAP=04.72 msGB300 · every side stream off5.65 msMI355X · e2e824dc586.65 ms
Two launches per GB300 bar, six rounds each; the verify cycle uses simulated acceptance 5.5. MI355X rows are the default-placement arms of sessions 1 and 2.GB300 每根柱子两次启动、每次六轮;verify 周期用模拟接受 5.5。 MI355X 为 session 1 和 2 的默认放置 arm。

The verify cycle grows from 4.61 to 6.38 to 7.34 ms and the plain decode step from 3.54 to 4.72 to 5.65 ms. Concurrency is worth 2.73 ms per verify on GB300, 98% of the bound, and 2.11 ms per decode step. The overlapped kernels hardly slow each other: at BS=1 they are small and leave most of the GPU idle, so a side stream fills it almost for free. One qualification: with its side streams off SGLang also takes a different path that launches 196 more kernels per verify (mHC statistics and glue that the overlapped path fuses), so 2.73 ms is what the overlapped layout is worth to GB300 as SGLang implements it, a little more than overlap alone. Against GB300's serial 7.34 ms, MI355X's 9.15 ms leaves 1.81 ms per verify, and 1.01 ms per decode step.

verify 周期从 4.61 ms 升到 6.38 ms, 再到 7.34 ms;plain decode 每步从 3.54 ms 升到 4.72 ms, 再到 5.65 ms。 所以在 GB300 上, 并发每个 verify 值 2.73 ms, 是上限的 98%, 每个 decode 步值 2.11 ms。 重叠执行的 kernel 几乎不互相拖慢: BS=1 时它们都很小, GPU 大部分时间空着, 侧 stream 几乎是白白把它填满。 需要说明一点: 关闭侧 stream 后, SGLang 还会走另一条代码路径, 每个 verify 多发射 196 个 kernel(重叠路径里融合掉的 mHC 统计和胶水 kernel), 所以 2.73 ms 是按 SGLang 现有实现、重叠布局对 GB300 的价值, 比纯粹的重叠收益略多一点。 以 GB300 串行的 7.34 ms 为基准, MI355X 的 9.15 ms 每个 verify 还差 1.81 ms, 每个 decode 步差 1.01 ms。

Why not overlap the same way on MI355X? We ran the graph microbenchmark of our earlier MI355X BS=1 campaign unchanged on GB300 (graph_floor.py: Triton and PyTorch kernels captured into one graph and replayed; device time per kernel, median of 31 rounds). A dependent tiny kernel costs 0.64 µs in a CUDA graph and 1.63 µs in a HIP graph. With fork/join branches the CUDA cost stays at 0.64 µs per kernel, while HIP rises to 7.1 µs, because a HIP graph with parallel branches leaves the packet-capture path. Two independent branches of 64 kernels finish in 70 µs on CUDA, sooner than one chain of 128 (85 µs), and in 373 µs on HIP, later than one chain (217 µs). Larger kernels narrow the ratio: a BF16 GEMV costs 5.2 against 8.6 µs per call.

为什么不在 MI355X 上同样重叠?我们把之前 MI355X BS=1 实验里的 graph microbenchmark 原样放到 GB300 上跑(graph_floor.py: 把 Triton 和 PyTorch kernel 捕获进一个 graph 再重放, 统计每个 kernel 的设备时间, 取 31 轮中位数)。 一个相互依赖的微小 kernel 在 CUDA graph 里要 0.64 µs, 在 HIP graph 里要 1.63 µs。 加上 fork/join 分支后, CUDA 每个 kernel 仍是 0.64 µs, HIP 却涨到 7.1 µs, 因为带并行分支的 HIP graph 会退出 packet capture 路径。 两条各 64 个 kernel 的独立分支, CUDA 上 70 µs 就完成, 比一条 128 个 kernel 的链(85 µs)还快;HIP 上要 373 µs, 比一条链(217 µs)还慢。 kernel 越大差距越小: 一个 BF16 GEMV 每次调用 5.2 µs 对 8.6 µs。

Plate VIIIGraph cost per kernel, CUDA on GB300 and HIP on MI355Xgraph 中每个 kernel 的开销: GB300 上的 CUDA 与 MI355X 上的 HIP
GB300 · CUDA 13.0 graphMI355X · HIP 7.2 graph0123456789microseconds per kernel inside a replayed graph (median of 31 rounds)dependent tiny kernel (1 program)0.641.63PyTorch elementwise chain0.751.5416 KiB producer → consumer1.041.83BF16 GEMV 6×5120×1152, per call5.178.57two parallel branches, per kernel0.552.91fork/join ladder, per kernel0.647.12
One GPU per machine, the same script and kernels. Tiny kernels are one program of 64 elements; the ladder puts one kernel on the capture stream and two on side streams per rung, 128 rungs.每台机器一张 GPU, 脚本和 kernel 相同。 微小 kernel 是一个 64 元素的 program;阶梯结构每一级在捕获 stream 上放一个 kernel、在侧 stream 上放两个, 共 128 级。

So the stream structure that saves GB300 2.7 ms would cost MI355X time today, which is why the ROCm path keeps one stream. The same numbers make a runtime request concrete: branches as cheap as on CUDA, and a lower floor per dependent kernel. Until that lands, the lever on MI355X is to shrink the side work, not to overlap it.

所以, 同样的 stream 结构在 GB300 上省下 2.7 ms, 在今天的 MI355X 上反而会增加耗时, 这也是 ROCm 路径保持单 stream 的原因。 这些数字让 runtime 需求变得具体: 分支要和 CUDA 一样便宜, 每个依赖 kernel 的下限要更低。 在那之前, MI355X 上能用的办法是把侧 stream 上的工作做小, 而不是去重叠它。

8What transfers to MI355X8哪些可以用到 MI355X 上

8.1Concurrency is the largest single gap8.1并发是最大的一项差距

Stream concurrency is worth 2.73 ms of GB300's verify cycle, 60% of the gap. The route GB300 takes is closed on MI355X today: a HIP graph kernel behind a fork/join costs 7.1 µs against 0.64 µs on CUDA, and in our earlier measurement a wait pending in another hardware queue slowed every dispatch of the running graph by about 1.3 µs. Two ways forward, not mutually exclusive: a ROCm runtime that keeps packet capture for branched graphs (section 7 sizes the prize), and meanwhile shrinking the side work until overlap no longer matters. The A/B says where that work sits: 1.77 ms in the streams SGLANG_OPT_USE_MULTI_STREAM_OVERLAP controls (attention preparation, mHC statistics, routed quantization, draft) and 0.96 ms in the shared experts.

stream 并发值 GB300 verify 周期的 2.73 ms, 占差距的 60%。 GB300 走的这条路今天在 MI355X 上走不通: HIP graph 里跟在 fork/join 后面的 kernel 每个要 7.1 µs, CUDA 上是 0.64 µs;我们之前还测到, 另一个硬件队列里挂着的等待, 会让正在运行的 graph 每次派发慢约 1.3 µs。 有两条路, 可以同时走: 让 ROCm runtime 在带分支的 graph 上保留 packet capture(第 7 节量化了能拿到多少);在那之前, 把侧 stream 上的工作做小, 小到重叠不再重要。 A/B 说明了这些工作在哪里: 1.77 ms 在 SGLANG_OPT_USE_MULTI_STREAM_OVERLAP 控制的几条 stream 上(attention 准备、mHC 统计、routed 量化、draft), 0.96 ms 在 shared expert 上。

8.2Kernel count costs more on MI355X8.2kernel 数量在 MI355X 上更贵

A dependent kernel costs 1.63 µs in a HIP graph and 0.64 µs in a CUDA graph, so every kernel removed from the verify graph saves 2.5× more on MI355X, and a fusion that is marginal on GB300 can pay here. The current base runs 1,519 kernels per cycle, 113 more than GB300; at the floor alone that difference is worth 0.18 ms per verify.

一个相互依赖的 kernel 在 HIP graph 里要 1.63 µs, 在 CUDA graph 里要 0.64 µs, 所以从 verify graph 里每去掉一个 kernel, MI355X 上省下的是 GB300 上的 2.5 倍;在 GB300 上可有可无的融合, 在这里可能值得做。 当前 base 每个周期执行 1,519 个 kernel, 比 GB300 多 113 个;仅按下限算, 这个差值每个 verify 就值 0.18 ms。

8.3Kernel time, ranked by the serialized gap8.3按串行化差距排序的 kernel 耗时

MoE (+782 µs) is first, with about 7 kernels per layer on MI355X (two GEMMs plus routing, sorting, quantization and activation) against 6 on GB300. Attention (+484 µs) is second. On the previous base MI355X spent 25 µs per layer in its sparse attention chain (15.1 µs main pass, 4.5 µs split reduce with inverse RoPE, 5.7 µs KV norm/RoPE/store); GB300's attention kernels take 14.1 µs per layer on their own. The current base already runs GB300's attention kernels, so what remains there is time per kernel, not a missing kernel. All-reduce with MoE finalize (+211 µs), the draft (+211 µs) and dense GEMM (+167 µs) are each about 0.2 ms; together they weigh as much as attention, but no single kernel dominates them.

排第一的是 MoE(+782 µs), MI355X 每层约 7 个 kernel(两个 GEMM, 加上路由、排序、量化和激活), GB300 是 6 个。 其次是 attention(+484 µs)。 在上一个 base 上, MI355X 每层在稀疏 attention 链路上花 25 µs(主计算 15.1 µs, 带 inverse RoPE 的 split reduce 4.5 µs, KV norm/RoPE/store 5.7 µs);GB300 的 attention kernel 单独运行每层 14.1 µs。 当前 base 已经在用 GB300 的 attention kernel, 所以这里剩下的是每个 kernel 的耗时, 而不是缺了哪个 kernel。 all-reduce 加 MoE finalize(+211 µs)、draft(+211 µs)和 dense GEMM(+167 µs)各约 0.2 ms;加起来和 attention 一样重, 但没有哪个 kernel 占主导。

per layer每层GB300 one stream µsGB300 单 stream µsMI355X previous base µsMI355X 上一个 base µsMI355X current base, traced µsMI355X 当前 base, trace µskernels GB / MIkernel 数 GB / MI
MoE expert GEMM 1 (gate/up)MoE expert GEMM 1(gate/up)12.718.1≤ 15.41.0 / 1.0
MoE expert GEMM 2 (down)MoE expert GEMM 2(down)9.28.8≤ 9.21.0 / 1.0
MoE routing, sorting, quantizationMoE 路由、排序、量化10.517.9–3.0 / 4.2
MoE activationMoE 激活1.34.2–1.0 / 1.0
attention main passattention 主计算7.615.1≤ 9.91.0 / 1.0
attention combine / split reduceattention combine / split reduce2.74.5≤ 5.21.0 / 1.0

GB300: nsys kernel time with every side stream off (kernels that use programmatic dependent launch can include a wait on their predecessor). MI355X previous base: the scaled Kineto ledger, whose small-kernel rows are upper bounds. Current base: rocprofv3 medians with packet capture off, upper bounds by the calibration above.

GB300: 关闭所有侧 stream 时 nsys 测得的 kernel 时长(使用 programmatic dependent launch 的 kernel 可能包含等待前一个 kernel 的时间)。 MI355X 上一个 base: 缩放后的 Kineto 账本, 小 kernel 那几行是上界。 当前 base: 关闭 packet capture 时 rocprofv3 的中位数, 按上面的校准也是上界。

Kernel by kernel, the largest single gap is the first expert GEMM (18.1 µs per layer on the previous base against 12.7 µs), while the second is at parity (8.8 against 9.2). Around the GEMMs MI355X ran 4.2 routing, sorting and quantization kernels per layer against 3, for 17.9 against 10.5 µs, and its activation cost 4.2 against 1.3 µs. On these numbers most of the MoE gap sits in the small kernels rather than the GEMMs (the small-kernel rows are upper bounds), which is where kernel count and fusion (8.2) apply. The attention main pass was 15.1 µs per layer on the previous base against GB300's 7.6; the current base runs GB300's partial_gluon, at most 9.9 µs traced, so that row has probably shrunk, and the rebuilt ledger will say by how much.

逐个 kernel 看, 差距最大的是第一个 expert GEMM(上一个 base 每层 18.1 µs, GB300 12.7 µs), 第二个基本持平(8.8 对 9.2)。 GEMM 周围, MI355X 每层要跑 4.2 个路由、排序和量化 kernel, GB300 是 3 个, 耗时 17.9 对 10.5 µs;激活 4.2 对 1.3 µs。 按这些数字, MoE 的差距大半在小 kernel 上, 而不在 GEMM(小 kernel 那几行是上界), 这正是 kernel 数量和融合(8.2)起作用的地方。 attention 主计算在上一个 base 上每层 15.1 µs, GB300 是 7.6 µs;当前 base 已经换成 GB300 的 partial_gluon, trace 读数最多 9.9 µs, 所以这一行很可能已经缩小, 重建后的账本会给出具体数字。

8.4Keep what MI355X already does better8.4保留 MI355X 已经做得更好的部分

MI355X answers the 4,096-token prefill in 143 ms against GB300's 208 ms bound and 412 ms unbound, because the MI350X cell captures breakable prefill graphs while CUDA main disables them for this model. That advantage should survive any port of NVIDIA-side changes: a change that needs eager prefill to work is a regression on this machine even if its decode numbers improve. The same holds for mHC, which MI355X runs 372 µs per verify faster than GB300 does on its own.

MI355X 完成 4,096 token prefill 只要 143 ms, GB300 绑定时要 208 ms, 不绑定时 412 ms, 原因是 MI350X 的 cell 捕获了 breakable prefill graph, 而 CUDA 版 main 对这个模型关掉了它。 移植 NVIDIA 侧改动时要保住这个优势: 如果某个改动必须靠 eager prefill 才能工作, 那么即使 decode 数字变好, 在这台机器上也是一次回退。 mHC 也一样: MI355X 每个 verify 在 mHC 上比 GB300 单独运行快 372 µs。

8.5Placement and measurement rules8.5放置与测量规则

  • Record where every scheduler runs (Cpus_allowed_list, per-thread CPU, numa_maps) in every arm; bind to the GPU's NUMA node, and on MI355X fix SGLANG_SET_CPU_AFFINITY to read the node from the PCI device instead of multiplying the local rank.
  • Use simulated acceptance only as a controlled experiment and bind NUMA when you do: it adds host-launched work that real acceptance does not have.
  • Compare cost per verify across platforms, and acceptance only on real text.
  • At least three fresh launches per configuration: the first two launches of the sim arm differed by 24% and their pooled median still looked like a plausible number.
  • Calibrate against the published number with its own commit, then move to current main.
  • Measure what overlap is worth by turning it off. The kernel-sum bound (2.8 ms) happened to be close here (2.73 ms measured), but it is only a bound.
  • Calibrate a kernel trace against an untraced microbenchmark before trusting small-kernel durations; on ROCm 7.2 rocprofv3 needs graph packet capture disabled and still reports a floor near 5 µs.
  • 每个 arm 都记录每个 scheduler 在哪里运行(Cpus_allowed_list、每个线程的 CPU、numa_maps);绑到 GPU 所在的 NUMA 节点;在 MI355X 上把 SGLANG_SET_CPU_AFFINITY 改成从 PCI 设备读取节点, 而不是用本地 rank 去乘。
  • 模拟接受只在受控实验里用, 用的时候一定要绑 NUMA: 它会引入真实接受所没有的、由 host 发起的工作。
  • 跨平台比较每次 verify 的耗时;接受长度只在真实文本上比较。
  • 每个配置至少启动三次新的 server: sim arm 的前两次启动相差 24%, 合并后的中位数看起来却仍像一个合理的数字。
  • 先用公开数字自己的 commit 校准, 再换到当前 main。
  • 要知道重叠值多少, 就把它关掉来测。 kernel 时长之和给出的上限(2.8 ms)这次恰好接近实测(2.73 ms), 但它只是上限。
  • 在相信小 kernel 的时长之前, 先用不开 profiler 的 microbenchmark 校准 kernel trace;在 ROCm 7.2 上, rocprofv3 需要关掉 graph packet capture, 而且仍然有接近 5 µs 的计时下限。

9Next experiments9下一步实验

  • Fold side-stream work into its neighbours on MI355X, one change at a time, each measured with the BS=1 A/B used here: the mHC statistics, the attention preparation (KV norm/RoPE/store with the compressor and indexer projections), then the MoE routing, sorting and quantization chain.
  • A/B MoE TP4 (EP1) against EP4 on MI355X. SGLang fuses the shared expert only without expert parallelism, and on GB300's older commit MoE TP4 ran 3.6% faster than EP4 under the published protocol.
  • File the ROCm requests with this page's numbers: fork/join graph branches (7.1 against 0.64 µs per kernel), the per-kernel graph floor (1.63 against 0.64 µs), and rocprofv3 aborting packet-captured graph replays.
  • Rebuild the MI355X ledger on the current base once a trace runs without that floor, or with event-timed phases as the previous ledger did.
  • Land kevin-mii/sglang#12 and #13, then measure the GPU-idle row again.
  • 在 MI355X 上把侧 stream 的工作并进相邻的 kernel, 每次只改一处, 每次都用本文的 BS=1 A/B 测量: 先是 mHC 统计, 再是 attention 准备(KV norm/RoPE/store, 连同 compressor 和 indexer 的投影), 然后是 MoE 的路由、排序和量化链路。
  • 在 MI355X 上对比 MoE TP4(EP1)和 EP4。 SGLang 只有在不用 expert parallelism 时才会融合 shared expert;在 GB300 的旧 commit 上, 按公开流程 MoE TP4 比 EP4 快 3.6%。
  • 把本文的数字附上, 向 ROCm 提需求: fork/join graph 分支(每个 kernel 7.1 µs 对 0.64 µs)、每个 kernel 的 graph 下限(1.63 µs 对 0.64 µs), 以及 rocprofv3 在 packet capture 的 graph 重放时中止。
  • 等 trace 没有这个计时下限了, 或者像上一份账本那样按阶段用 event 计时, 再在当前 base 上重建 MI355X 账本。
  • 合入 kevin-mii/sglang#12 和 #13, 然后重新测 GPU 空闲那一行。

AEvery armA所有 arm

Every group recomputed from its per-round measurements (warm-up excluded, launches pooled). Cycle cost for DSpark arms, step cost for plain decode.

每一组都从逐轮测量重新计算(去掉热身, 合并各次启动)。 DSpark arm 给出每周期耗时, plain decode 给出每步耗时。

armarmconfiguration配置launches启动tok/s median (range)tok/s 中位数(范围)ALms / cycle or stepms / 周期或步TTFT msper-launch tok/s各次启动 tok/s
4× GB300 · CUDA 13.2
s0-tp4BBuf 835c3909, MoE TP4 (EP1), sim 5.5, unbound作者 commit 835c3909, MoE TP4(EP1), sim 5.5, 不绑定2870.6 (844–883)5.516.33249878 · 864
s0-ep4BBuf 835c3909, MoE EP4, sim 5.5, unbound作者 commit 835c3909, MoE EP4, sim 5.5, 不绑定2841.3 (787–853)5.516.54247832 · 844
s0-nodsBBuf 835c3909, EP4, DSpark off, unbound作者 commit 835c3909, EP4, 关闭 DSpark, 不绑定2223.9 (215–224)–4.47323224 · 224
n0-tp4BBuf 835c3909, MoE TP4 (EP1), sim 5.5, NUMA-bound作者 commit 835c3909, MoE TP4(EP1), sim 5.5, NUMA 绑定2878.8 (863–892)5.516.29208876 · 881
n0-ep4BBuf 835c3909, MoE EP4, sim 5.5, NUMA-bound作者 commit 835c3909, MoE EP4, sim 5.5, NUMA 绑定2848.5 (836–862)5.516.48206849 · 848
n0-nodsBBuf 835c3909, EP4, DSpark off, NUMA-bound作者 commit 835c3909, EP4, 关闭 DSpark, NUMA 绑定2224.3 (221–225)–4.46205224 · 224
s1a-offmain ffac53d779, HT cell (DSpark off), unboundmain ffac53d779, HT cell(关闭 DSpark), 不绑定2281.8 (279–283)–3.55412281 · 282
s1a-simmain ffac53d779, LL cell, sim 5.5, unboundmain ffac53d779, LL cell, sim 5.5, 不绑定2968.3 (852–1,201)5.515.672431184 · 958
s1a-realmain ffac53d779, LL cell, real acceptance, unboundmain ffac53d779, LL cell, 真实接受, 不绑定21,245.0 (852–1,334)5.794.664251245 · 1245
n1a-offmain ffac53d779, HT cell (DSpark off), NUMA-boundmain ffac53d779, HT cell(关闭 DSpark), NUMA 绑定2282.9 (278–283)–3.53208283 · 283
n1a-simmain ffac53d779, LL cell, sim 5.5, NUMA-boundmain ffac53d779, LL cell, sim 5.5, NUMA 绑定21,190.1 (1,164–1,216)5.514.612131192 · 1188
n1a-realmain ffac53d779, LL cell, real acceptance, NUMA-boundmain ffac53d779, LL cell, 真实接受, NUMA 绑定21,241.0 (891–1,338)5.754.612061235 · 1269
e1-base-simmain ffac53d779, LL cell, sim 5.5, unbound (NUMA control)main ffac53d779, LL cell, sim 5.5, 不绑定(NUMA 对照)3908.8 (778–1,045)5.516.06366900 · 995 · 909
e1-nice-simmain ffac53d779, LL cell, sim 5.5, NUMA-bound (NUMA control)main ffac53d779, LL cell, sim 5.5, NUMA 绑定(NUMA 对照)31,185.6 (1,160–1,214)5.514.642121193 · 1191 · 1183
w3-base-simmain ffac53d779, LL cell, sim 5.5, NUMA-bound, default streams (session 3)main ffac53d779, LL cell, sim 5.5, NUMA 绑定, 默认 stream 布局(session 3)21,193.1 (1,163–1,205)5.514.612151183 · 1197
w3-opt0-simmain ffac53d779, LL cell, sim 5.5, NUMA-bound, SGLANG_OPT_USE_MULTI_STREAM_OVERLAP=0main ffac53d779, LL cell, sim 5.5, NUMA 绑定, SGLANG_OPT_USE_MULTI_STREAM_OVERLAP=02858.0 (822–884)5.516.38201858 · 859
w3-serial-simmain ffac53d779, LL cell, sim 5.5, NUMA-bound, every side stream offmain ffac53d779, LL cell, sim 5.5, NUMA 绑定, 关闭所有侧 stream2749.4 (730–767)5.517.34201747 · 752
w3-base-offmain ffac53d779, HT cell (DSpark off), NUMA-bound, default streams (session 3)main ffac53d779, HT cell(关闭 DSpark), NUMA 绑定, 默认 stream 布局(session 3)2282.8 (283–283)–3.54210283 · 283
w3-opt0-offmain ffac53d779, HT cell (DSpark off), NUMA-bound, SGLANG_OPT_USE_MULTI_STREAM_OVERLAP=0main ffac53d779, HT cell(关闭 DSpark), NUMA 绑定, SGLANG_OPT_USE_MULTI_STREAM_OVERLAP=02211.9 (209–212)–4.72196212 · 212
w3-serial-offmain ffac53d779, HT cell (DSpark off), NUMA-bound, every side stream offmain ffac53d779, HT cell(关闭 DSpark), NUMA 绑定, 关闭所有侧 stream2177.1 (173–177)–5.65197177 · 177
4× MI355X · ROCm 7.2
u-offe2e824dc58+#8, HT cell, SGLang split affinitye2e824dc58+#8, HT cell, SGLang 切分亲和性2150.3 (150–150)–6.65143150 · 150
u-reale2e824dc58+#8, LL cell, real, SGLang split affinitye2e824dc58+#8, LL cell, 真实接受, SGLang 切分亲和性2389.6 (269–579)3.549.14144390 · 371
u-sime2e824dc58+#8, LL cell, sim 5.5, SGLang split affinitye2e824dc58+#8, LL cell, sim 5.5, SGLang 切分亲和性3603.1 (590–610)5.519.15145602 · 603 · 603
n-offe2e824dc58+#8, HT cell, split affinity + main on node 1e2e824dc58+#8, HT cell, 切分亲和性 + 主进程在 node 12150.4 (150–150)–6.65143150 · 150
n-reale2e824dc58+#8, LL cell, real, split affinity + main on node 1e2e824dc58+#8, LL cell, 真实接受, 切分亲和性 + 主进程在 node 12415.4 (310–579)3.799.14145463 · 359
n-sime2e824dc58+#8, LL cell, sim 5.5, split affinity + main on node 1e2e824dc58+#8, LL cell, sim 5.5, 切分亲和性 + 主进程在 node 13601.6 (589–605)5.519.16145604 · 602 · 598
s2-split-simsession 2, sim 5.5, SGLang split affinitysession 2, sim 5.5, SGLang 切分亲和性2602.8 (588–610)5.519.14145601 · 605
s2-node1-simsession 2, sim 5.5, all schedulers on node 1session 2, sim 5.5, 所有 scheduler 在 node 12602.4 (591–608)5.519.14144603 · 602
s2-none-simsession 2, sim 5.5, no pinningsession 2, sim 5.5, 不固定2600.9 (588–607)5.519.18145602 · 600
s2-split-realsession 2, real, SGLang split affinitysession 2, 真实接受, SGLang 切分亲和性2433.9 (288–576)3.959.13144453 · 434
s2-node1-realsession 2, real, all schedulers on node 1session 2, 真实接受, 所有 scheduler 在 node 12442.6 (287–577)4.049.12144450 · 363
s2-none-realsession 2, real, no pinningsession 2, 真实接受, 不固定2316.8 (261–579)2.909.13144416 · 274
s2-node1-offsession 2, DSpark off, all schedulers on node 1session 2, 关闭 DSpark, 所有 scheduler 在 node 11150.1 (150–150)–6.66142150
s2-none-offsession 2, DSpark off, no pinningsession 2, 关闭 DSpark, 不固定1149.9 (150–150)–6.67143150

10Reproduce10复现

All inputs, per-round CSVs, placements, profiles, scripts and the analysis are in data/dsv41-gb300-vs-mi355x/, with hostnames and machine paths replaced by placeholders. The GB300 arms were driven by gb300/scripts/arm.py (one container per launch, idle check, nvidia-smi before and after); the MI355X arms by mi355x/scripts/run_arm*.sh under an exclusive eight-GPU lease with HBM canaries.

所有输入、逐轮 CSV、进程放置记录、profile、脚本和分析都在 data/dsv41-gb300-vs-mi355x/, 主机名和机器路径已替换成占位符。 GB300 的 arm 由 gb300/scripts/arm.py 驱动(每次启动一个容器, 先做空闲检查, 前后各记录一次 nvidia-smi);MI355X 的 arm 由 mi355x/scripts/run_arm*.sh 在独占八卡的租约下运行, 前后做 HBM canary。

One lease needs a note. After session 2 the canary on GPU 4 read 6.48 ms per copy against 7.15 ms before it, 10.4% faster, which crosses the lease tool's 10% threshold, so the tool flagged the session. We kept it: every process the monitor saw on the node's GPUs appeared and disappeared inside one of our server launches, the other three canaries moved by under 0.3%, a busier device can only lengthen a copy, and session 2's split-affinity cycles match session 1's within 0.2%. The lease records are in mi355x/leases/.

有一次租约需要说明。 session 2 结束后, GPU 4 上的 canary 每次拷贝耗时 6.48 ms, 开始前是 7.15 ms, 快了 10.4%, 超过了租约工具 10% 的阈值, 所以工具把这个 session 标记为可疑。 我们保留了它: 监控看到的每个使用这台机器 GPU 的进程, 都在我们某一次 server 启动的时间窗内出现和消失;另外三张卡的 canary 变化都在 0.3% 以内;设备被占用只会让拷贝变慢;session 2 切分亲和性的周期和 session 1 相差不到 0.2%。 租约记录在 mi355x/leases/。

Session 3 ran inside the GB300's persistent workbench container, which has no Docker CLI: gb300/scripts/wb/wb_arm.py starts each server as a fresh process there with the same cells, client, JIT caches and records as arm.py (and with CAP_SYS_NICE, so every arm is NUMA-bound); queue3.sh is the whole session, serial_moe/ the shared-expert patch and bs1_realtext.py the real-text client. On MI355X, graph_floor.py ran under a clean exclusive lease (session 3). The rocprofv3 calibration and traced server ran in sessions 3b and 3c, whose canaries moved by up to 11% in both directions after traced servers were killed; nothing from them is used as a timing, only the trace's kernel counts.

session 3 在 GB300 常驻的 workbench 容器里运行, 那里没有 Docker CLI: gb300/scripts/wb/wb_arm.py 在容器里把每个 server 作为全新进程启动, cell、客户端、JIT cache 和记录方式都与 arm.py 相同(容器有 CAP_SYS_NICE, 所以每个 arm 都绑定了 NUMA);queue3.sh 是整个 session, serial_moe/ 是 shared expert 的补丁, bs1_realtext.py 是真实文本客户端。 MI355X 上, graph_floor.py 在一个干净的独占租约里运行(session 3)。 rocprofv3 的校准和带 trace 的 server 在 session 3b 和 3c 里运行, 强杀带 trace 的 server 之后, 这两个租约的 canary 双向变化最多 11%;它们的结果都没有被当作计时使用, 只用了 trace 里的 kernel 数量。

# GB300, one launch of the Low-Latency cell with simulated acceptance, NUMA binding on
docker run -d --init --gpus all --ipc=host --shm-size 32g --network host --cap-add=SYS_NICE \
  -e SGLANG_RAGGED_VERIFY_MODE=static -e SGLANG_SIMULATE_ACC_LEN=5.5 -e SGLANG_SIMULATE_ACC_METHOD=match-expected \
  -v $MODEL_DIR:/models/DeepSeek-V4.1-Flash:ro lmsysorg/sglang@sha256:b8257f5c5f8f7c5a... \
  sglang serve --model-path /models/DeepSeek-V4.1-Flash --trust-remote-code --tp 4 --ep-size 4 \
    --mem-fraction-static 0.8 --speculative-algorithm DSPARK --speculative-dspark-block-size 5 \
    --cuda-graph-max-bs-decode 64 --disable-radix-cache --random-seed 42 --port 30000
python3 benchmark.py bench --prompt prompt.json --max-tokens 1024 --repeat 6 --out result \
  --url http://127.0.0.1:30000

# MI355X, the same client against the MI350X cell (HIP 4-7), placement mode split|node1|none
bash lease.sh session2 -- bash session2.sh      # calls run_arm2.sh LABEL MODE AFF per arm

# GB300 session 3 (inside the workbench): stream A/B, serial nsys, real text, graph floor
bash gb300/scripts/wb/queue3.sh

# analysis
python3 analysis/analyze.py && python3 analysis/compare_attribution.py && python3 analysis/realtext.py

Epilogue后记

The GB300 run was supposed to give us a target. It gave us a decomposition instead: of MI355X's 4.5 ms per verify, 2.7 ms is concurrency that CUDA graphs give GB300 almost for free and HIP graphs charge for, 1.3 ms is MoE and attention kernel time, and the rest is spread thin. Along the way it also gave us a reminder that a GPU benchmark can be a CPU benchmark in disguise, and a placement bug on our own machine that no GPU counter would have shown.

这次 GB300 实验本来是想拿到一个目标值, 结果得到的是一份拆解: MI355X 每个 verify 多出的 4.5 ms 里, 2.7 ms 是并发, CUDA graph 让 GB300 几乎白得, HIP graph 却要为它付出代价;1.3 ms 是 MoE 和 attention 的 kernel 耗时;其余分散在各处。 顺带还得到一个提醒, GPU 测试有时其实是在测 CPU;以及我们自己机器上的一个放置 bug, 任何 GPU 计数器都看不出来。