Urgent.News

What's breaking now, across thousands of outlets.

AI

Gemma 4 From E2B to 31B on an AMD MI300X: fp8 Overtakes bf16 From 12B Up

This article provides a step by step guide to serving every Gemma 4 size, E2B, E4B, 12B, 26B-A4B and 31B, on one AMD Instinct MI300X through vLLM in four weight formats, with each build timed across the same grid of request counts and prompt lengths on the same image. Every log, report and script is committed. The answer to "which format should I serve?" changes with the size of the model. At…

This article provides a step-by-step guide to serving every Gemma 4 size, from E2B to 31B, on a single AMD Instinct MI300X using vLLM in four weight formats. The performance of each format varies depending on the model size. At E2B, bf16 is fastest for a single user, but fp8 only catches up with 8 or more requests in flight. From 12B and larger, fp8 consistently outperforms bf16, providing 1.12x to 1.41x speed improvement at 12B and 1.20x to 1.45x at 31B.

The 4-bit builds are 0.14x to 0.69x faster than bf16 at every size. The article also explains the reasoning behind testing various formats on different model sizes, the memory and performance implications of different quantization levels, and the results of the benchmarks conducted.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Liquid AI's d1 Decision Models Went Open: Triage Support Tickets on a CPU With the 600M One (and Where It Fools You)

Attributed compile + one small real CPU run Primary sources (Liquid AI, 2026-10-07): Open d1: Edge decision models for text, vision, and audio (Liquid AI blog), Multimodal open d1 decision models for…

  • Liquid AI released d1 decision models (d1-3B and d1-omni-600M) as open source
  • d1-3B model scores 48.57 on Decision Index v0.2.1, best under 10B
  • d1-omni-600M model tested on 12 support tickets, flagged high-risk cases

How to test whether coding agents discover the skills you need

A coding skill can be clear, correct, and still fail to help if an agent never finds it. Testing a skill by naming it directly only answers whether the agent can use it after it has been loaded.

  • Conduct task-based assessments for coding agents to gauge skill discovery.
  • Analyze execution traces to measure skill loading without explicit guidance.
  • Track three metrics: skill activation, utilization, and application performance.

The AI Co-Pilot Dilemma: Does Code Generation Stifle Authentic Programming Practice?

AI coding assistants now write boilerplate, suggest entire functions, and explain error messages in seconds. For many developers, that speed feels like a gift.

  • AI coding assistants generate code rapidly, raising concerns about authentic practice.
  • Removing friction from programming may hinder learning if solutions aren't understood.
  • Beginners should write their own versions of code to develop deeper comprehension.

WIldGuard AI

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass What I Built WildGuard AI — AI-Powered Wildlife Exploration WildGuard AI is an AI-powered wildlife application…

  • WildGuard AI is an AI-powered wildlife identification app
  • Open-source AI models enable local inference and offline identification

More from Friday 9 October →