{
  "id": 12899372,
  "title": "Gemma 4 E2B on an AMD MI300X: Which Weight Format Should You Serve?",
  "url": "https://urgent.news/2026/10/08/gemma-4-e2b-on-an-amd-mi300x-which-weight-format-should-you-serve",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-08T15:59:16.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/gde/gemma-4-e2b-on-an-amd-mi300x-which-weight-format-should-you-serve-1a01"
  },
  "original_language": "en",
  "account": "This article presents a detailed guide on how to serve ten weight formats of Gemma 4 E2B on a single AMD Instinct MI300X using vLLM. Each build is timed across a grid of request counts and prompt lengths, all recorded in logs, reports, and scripts. The article reveals that fp8 is the only format that keeps pace with bf16 on the MI300X, achieving 0.75x speed for a single request and up to 1.09x with 8 or 64 requests simultaneously. Int8 W8A8 runs slower, from 0.29x to 0.87x, while 4-bit W4A16 builds perform from 0.14x to 0.63x. The best performing format also comes with the most significant deviation from Google's trained weights, creating a trade-off between speed and fidelity.\n\nThe article poses the question of why compare formats on a 192 GB card, as Gemma 4 E2B only takes up 9.42 GiB of memory in bf16. The MI300X, with 192 GB of HBM memory, has enough space to accommodate a KV cache of 9,045,060 tokens, with the smallest build stretching it to 9,475,223 tokens. The key differentiating factor then becomes speed and precision.\n\nThe article enumerates the ten builds, all starting from Google's gemma-4-E2B-it-qat-q4_0-unquantized release. These builds involve making adjustments to linear layers, vocabulary tables, and per-layer embeddings in various formats: fp8, FP8 E4M3, FP8 activations per token, FP8 E4M3FNUZ, FP8 E4M3 int4, FP8 E4M3FNUZ int4, and q4w4a16. Each format stores Google's quantization-aware-trained (QAT) weights with different amounts of rounding.\n\nThe article emphasizes that the choice of format on the AMD MI300X hinges on the trade-off between speed and precision, with the card's native multiplication of some number formats and emulation of others playing a crucial role.",
  "summary": "This article provides a step by step guide to serving ten weight formats of Gemma 4 E2B on one AMD Instinct MI300X through vLLM, with every build timed across a grid of request counts and prompt lengths on the same card, image and day. Every log, report and script is committed. On the MI300X, fp8 is the only format that keeps pace with bf16: 0.75x for a single request and up to 1.09x with 8 or 64…",
  "key_points": [
    "fp8 format matches bf16 speed on AMD MI300X, achieving 0.75x for single request",
    "Int8 W8A8 format significantly slower, from 0.29x to 0.87x",
    "fp8 format has largest deviation from Google's trained weights"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}