WebGPU transformer inference: 458× speedup by fusing 1,024 dispatches into one Second preprint applying kernel fusion, this time to autoregressive…