Urgent.News

What's breaking now, across thousands of outlets.

AI

Gemma 4 E2B on an AMD MI300X: Which Weight Format Should You Serve?

This article provides a step by step guide to serving ten weight formats of Gemma 4 E2B on one AMD Instinct MI300X through vLLM, with every build timed across a grid of request counts and prompt lengths on the same card, image and day. Every log, report and script is committed. On the MI300X, fp8 is the only format that keeps pace with bf16: 0.75x for a single request and up to 1.09x with 8 or 64…

This article presents a detailed guide on how to serve ten weight formats of Gemma 4 E2B on a single AMD Instinct MI300X using vLLM. Each build is timed across a grid of request counts and prompt lengths, all recorded in logs, reports, and scripts. The article reveals that fp8 is the only format that keeps pace with bf16 on the MI300X, achieving 0.75x speed for a single request and up to 1.09x with 8 or 64 requests simultaneously.

Int8 W8A8 runs slower, from 0.29x to 0.87x, while 4-bit W4A16 builds perform from 0.14x to 0.63x. The best performing format also comes with the most significant deviation from Google's trained weights, creating a trade-off between speed and fidelity.

The article poses the question of why compare formats on a 192 GB card, as Gemma 4 E2B only takes up 9.42 GiB of memory in bf16. The MI300X, with 192 GB of HBM memory, has enough space to accommodate a KV cache of 9,045,060 tokens, with the smallest build stretching it to 9,475,223 tokens. The key differentiating factor then becomes speed and precision.

The article enumerates the ten builds, all starting from Google's gemma-4-E2B-it-qat-q4_0-unquantized release. These builds involve making adjustments to linear layers, vocabulary tables, and per-layer embeddings in various formats: fp8, FP8 E4M3, FP8 activations per token, FP8 E4M3FNUZ, FP8 E4M3 int4, FP8 E4M3FNUZ int4, and q4w4a16. Each format stores Google's quantization-aware-trained (QAT) weights with different amounts of rounding.

The article emphasizes that the choice of format on the AMD MI300X hinges on the trade-off between speed and precision, with the card's native multiplication of some number formats and emulation of others playing a crucial role.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Aymeric Lim Is Building AI-Literacy for a New Generation of Health Care

Lim said the culture of learning at NUH helps them continue to improve and adapt as health care evolves.

  • Aymeric Lim is NUH CEO in Singapore, transforming hospital into AI-literate institution.
  • Lim's leadership focuses on future-readiness, new care models, investing in people.
  • Lim values teaching, learning, and humanitarian service in healthcare.

More from Thursday 8 October →