Decoder-Only Transformer Optimization 101

By TongKe Xue Rohit Swamy Eric Gu et al. | September 15, 2026

How fast can one RTX 6000 Pro train a decoder-only Transformer to play Super Smash Bros. Melee? This report follows the path from a naive 163,645 tokens/s implementation through fused kernels, buffer reuse, CUDA-Oxide, and a smaller architecture that reaches 3,217,492 tokens/s while beating a level-9 CPU Fox.

Read the paper 24 pages • PDF

This browser cannot display the embedded PDF.

Open Decoder-Only Transformer Optimization 101 as a PDF.

Follow @lognprg Follow @bicro_