Gemma 4 E2B in Pure JAX on a Colab TPU: Google's 4-Bit Export Against an Exact Repack
This article provides a step by step guide to a Colab notebook that serves Gemma 4 E2B on a single TPU v5e chip with a pure-JAX engine and compares two 4-bit builds of the same model against the weights Google trained. Every number below was measured in the notebook on a Colab v5e-1 runtime, and the executed notebook is committed. Google ships E2B in 4 bits as gemma-4-E2B-it-qat-w4a16-ct . Its…
This article provides a comprehensive guide on how to run Gemma 4 E2B on a single TPU v5e chip using a pure-JAX engine on a Google Colab notebook. The notebook compares two 4-bit builds of the same model against the weights Google trained. The first build, provided by Google, is exported as gemma-4-E2B-it-qat-w4a16-ct, while the second build is a repack that stores the trained grid, resulting in a 342.6 times closer approximation to the original model in next-token predictions, and matching its top token 99.29% of the time.
The notebook also highlights the differences between the two 4-bit builds and provides step-by-step instructions on how to download the three checkpoints and run the notebook on a TPU runtime.
Brief written by urgent.news from Dev.to's own syndicated text. Machine-written — may contain errors; check the original before relying on it.