Gemma 4 QAT on One TPU v5e: What Runs and What Doesn't
This article provides a step by step guide to repacking Google's quantization-aware-trained (QAT) Gemma 4 weights for vLLM and serving them on one Google Cloud TPU v5e chip, with every build scored for classification, math, tool calling, throughput and long prompts. Every per-record output, log and script is committed. On one v5e chip the repacked QAT builds serve every Gemma 4 size from E2B to…
This article offers a step-by-step guide on repacking Google's quantization-aware-trained (QAT) Gemma 4 weights for serving on a single Google Cloud TPU v5e chip, testing their performance across various metrics. The QAT builds allow the Gemma 4 model, ranging in size from E2B to 26B parameters, to run on one v5e chip without issue.
The 4-bit repack stores the QAT grid values using int4 weights and activations, while the 8-bit repack employs int8 weights and activations, which show a slight speed advantage. Despite the 4-bit repack's slightly lower performance compared to the 8-bit repack, both models match the performance of Google's own 4-bit exports at the same speed.
The 12B model is the largest Gemma 4 size that can fit on the TPU v5e chip, with a weight size of 11.31 GiB and 675 output tokens per second at 16 requests. Overall, the repacking process enables Gemma 4 models to run efficiently on a single TPU v5e chip while maintaining competitive performance against Google's own 4-bit exports.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.