Urgent.News

What's breaking now, across thousands of outlets.

AI

Gemma 4 QAT on One TPU v5e: What Runs and What Doesn't

This article provides a step by step guide to repacking Google's quantization-aware-trained (QAT) Gemma 4 weights for vLLM and serving them on one Google Cloud TPU v5e chip, with every build scored for classification, math, tool calling, throughput and long prompts. Every per-record output, log and script is committed. On one v5e chip the repacked QAT builds serve every Gemma 4 size from E2B to…

This article offers a step-by-step guide on repacking Google's quantization-aware-trained (QAT) Gemma 4 weights for serving on a single Google Cloud TPU v5e chip, testing their performance across various metrics. The QAT builds allow the Gemma 4 model, ranging in size from E2B to 26B parameters, to run on one v5e chip without issue.

The 4-bit repack stores the QAT grid values using int4 weights and activations, while the 8-bit repack employs int8 weights and activations, which show a slight speed advantage. Despite the 4-bit repack's slightly lower performance compared to the 8-bit repack, both models match the performance of Google's own 4-bit exports at the same speed.

The 12B model is the largest Gemma 4 size that can fit on the TPU v5e chip, with a weight size of 11.31 GiB and 675 output tokens per second at 16 requests. Overall, the repacking process enables Gemma 4 models to run efficiently on a single TPU v5e chip while maintaining competitive performance against Google's own 4-bit exports.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Running DeepAgents in a Docker Sandbox, with no cloud keys

Agent frameworks are easy to pip install and surprisingly hard to run responsibly. The moment you give an agent a filesystem, a shell, and a network, you have handed arbitrary generated code the same…

  • DeepAgents framework allows creating agents with LangGraph library
  • Docker Sandbox Kit enables running DeepAgents without cloud credentials
  • Kit consists of four files and grants phase-scoped network policy

I built a PR reviewer that survived its own model being deprecated

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend What I Built groq-pr-reviewer-net — a .NET 10 CLI that reads your git diff and returns a structured code review in your…

  • Author created PR reviewer tool named groq-pr-reviewer-net
  • Tool assesses code for bugs, security, performance, readability
  • Powered by open-weight Groq model with model swap capability

More from Friday 2 October →