I Tested Q4_K_M vs MXFP4 on the Same Laptop — The Supposedly-Faster New Format Lost
I tested two local LLMs in two different quantization formats on the same laptop, on the same prompt, three trials each. The result is the opposite of what the marketing says: the supposedly-faster new format lost by 1.8x. Q4_K_M (the older integer-based format) hit 4.7 tokens/second and finished a 200-token generation in 44 seconds. MXFP4 (OpenAI's newer microscaling FP format) hit 2.6…
A reporter compared the performance of two local LLMs, Alibaba's Qwen3-14B in Q4_K_M format and OpenAI's gpt-oss-20B in MXFP4 format, on the same laptop with the same hardware and prompt. The supposedly faster MXFP4 format actually performed worse, generating text at half the speed of Q4_K_M. The report suggests that Q4_K_M is still the optimal choice for speed on M2 or M3 Macs, while MXFP4 may become competitive on newer M4-class Macs.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.