Stop being the product.
Become the owner.
or
sign uplog in

WebGPU transformer inference: 458× speedup by fusing 1,024…

WebGPU transformer inference: 458× speedup by fusing 1,024 dispatches into one

Second preprint applying kernel fusion, this time to autoregressive transformer decoding.

The finding: browser LLM engines waste 92% of their time on dispatch overhead. Fusing the full token×layer×operation loop into a single GPU dispatch eliminates it.

Parallel kernel (64 threads): 66-458× over unfused, beats PyTorch MPS 7.5-161× on same hardware.

Run it: gpubench.dev/transformer http://gpubench.dev/transformer
Preprint: doi.org/10.5281/zenodo.19344277 http://doi.org/10.5281/zenodo.19344277
Code: github.com/abgnydn/webgpu-transformer-fusion http://github.com/abgnydn/webgpu-transformer-fusion
Research: kernelfusion.dev http://kernelfusion.dev

Kernel fusion eliminates 92% GPU dispatch overhead — 458× faster transformer inference in the browser https://preview.redd.it/xyobajtxgdsg1.png?width=3456&format=png&auto=webp&s=9fcab29d6a19bbfc71bc6d8b38b79322ac5d58e3
#technology
earnings
4,000 mlx total
$0  total
engagement
4 views
0 reactions

0 comments