GH200 Is Launch-Bound
Abstract. The hypothesis that GH200 is compute-bound at W4A16 is wrong. Roofline operational intensity is 3.9 against a machine balance of 268 — deeply bandwidth-bound on paper. Measured decode is 94 tok/s, only 19% of the ~500 tok/s bandwidth ceiling. Telemetry: 27% compute, 13% HBM, 216 W of 700 W TDP. The missing 80% is kernel scheduling, CUDA-graph efficiency, and per-layer launch latency. The hardware is idling between launches.
Roofline
Hardware: GH200 480 GB (H100 die, 132 SMs, HBM3e ~4.0 TB/s achievable). Model: 64 layers, hidden 3584, GQA 28/4, head 128, SwiGLU intermediate 18944. Decode batch=1 is ~30.92 GFLOP/token. Operational intensity 3.9 FLOP/byte versus machine balance 268 — bandwidth-bound if kernels actually ran.
They do not. SM clock and memory clock sit at max while utilization does not. Power is 31% of TDP. The bottleneck moved off the roofline and into the launch path: graph capture across 64 layers, weight-swap vs two-graph dispatch, CPU-side scheduling between tokens.
Implication
Further quantization or tensor-core chasing will not close 81% of the gap. CUDA graphs, fused layer regions, and eliminating host round-trips will. Companion notes in the same architecture folder treat CUDA-graph weight swap, two-graph dispatch, and DeltaNet weight transplant as the actual levers — consistent with a launch-bound diagnosis, not a FLOP diagnosis.
Source report. Summarized from docs/architecture/GH200-COMPUTE-BANDWIDTH-ANALYSIS.md in the hive tree. This page is the paper. The report remains the primary.