Urgent.News

What's breaking now, across thousands of outlets.

AI

LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation

Today, we release QAD Q4_0 GGUFs for four LFM2.5 models: LFM2.5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.6B. These updated 4-bit checkpoints maintain the models' performance while improving memory efficiency and speed. We compare the released QAD Q4_0 checkpoints against the original trained checkpoints using a benchmark suite covering reasoning, instruction-following, tool use, and agentic capabilities.

The BF16 GGUF serves as the in-format ceiling, and we also include a math evaluation: GSM8K for the smaller models, and AIME25 for the larger ones. Across all four models, the QAD checkpoints retain 97% of their BF16 baseline performance. The QAD Q4_0 checkpoints demonstrate higher decode throughput than their BF16 counterparts, ranging from 4-33% for the smaller models, and 3-14% for the larger ones.

They also match the performance of Unsloth's UD-Q4_K_XL checkpoint for the 230M and 1.2B models. To use the QAD GGUFs, integrate them into llama.cpp or any runtime that supports GGUF Q4_0 artifacts. The QAD checkpoints are available on Hugging Face for citation.

Written by urgent.news from Hugging Face's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at huggingface.co →

More in AI

More from Wednesday 19 August →