{
  "id": 13126079,
  "title": "Gemma 4 From E2B to 31B on an AMD MI300X: fp8 Overtakes bf16 From 12B Up",
  "url": "https://urgent.news/2026/10/09/gemma-4-from-e2b-to-31b-on-an-amd-mi300x-fp8-overtakes-bf16-from-12b",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-09T13:46:34.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/gde/gemma-4-from-e2b-to-31b-on-an-amd-mi300x-fp8-overtakes-bf16-from-12b-up-h4e"
  },
  "original_language": "en",
  "account": "This article provides a step-by-step guide to serving every Gemma 4 size, from E2B to 31B, on a single AMD Instinct MI300X using vLLM in four weight formats. The performance of each format varies depending on the model size. At E2B, bf16 is fastest for a single user, but fp8 only catches up with 8 or more requests in flight. From 12B and larger, fp8 consistently outperforms bf16, providing 1.12x to 1.41x speed improvement at 12B and 1.20x to 1.45x at 31B. The 4-bit builds are 0.14x to 0.69x faster than bf16 at every size. The article also explains the reasoning behind testing various formats on different model sizes, the memory and performance implications of different quantization levels, and the results of the benchmarks conducted.",
  "summary": "This article provides a step by step guide to serving every Gemma 4 size, E2B, E4B, 12B, 26B-A4B and 31B, on one AMD Instinct MI300X through vLLM in four weight formats, with each build timed across the same grid of request counts and prompt lengths on the same image. Every log, report and script is committed. The answer to \"which format should I serve?\" changes with the size of the model. At…",
  "key_points": [
    "Gemma 4 models range from E2B to 31B size",
    "fp8 format overtakes bf16 from 12B and larger",
    "4-bit models are 0.14x to 0.69x faster than bf16"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}