Decoder-Only Transformer Optimization 101
How fast can one RTX 6000 Pro train a decoder-only Transformer to play Super Smash Bros. Melee? This report follows the path from a naive 163,645 tokens/s implementation through fused kernels, buffer reuse, CUDA-Oxide, and a smaller architecture that reaches 3,217,492 tokens/s while beating a level-9 CPU Fox.