Uzu implementation note
In a recent development, Uzu has unveiled its speculative decoding implementation tailored for Qwen3.6 27B. The company plans to expand support to Qwen3.8 27B and Muse Glimmer in the near future. On Apple M5-series chips, Uzu outperforms MTPLX by nearly double and llama.cpp by over triple, particularly excelling in mathematical reasoning and coding tasks.
The model's architecture is specifically designed for Apple M5 chips to maximize GPU Neural Accelerator performance. Unlike other speculative decoding methods that generate small draft chains of 3-4 tokens, Uzu employs aggressive speculative budgets of 16-32 tokens, allowing for extensive utilization of hardware arithmetic throughput.
The company's hybrid draft model, DFlash-Weaver, integrates a compact autoregressive transformer to consolidate top-k predictions from a parallel DFlash drafter into a tree structure of coherent continuations. For Qwen3.6, which is a hybrid architecture featuring Gated DeltaNet layers, Uzu has implemented rollback-free tree verification kernels using Metal.
The company utilizes an adaptive tree construction algorithm that generates lengthy chains when the drafter's predictions are confident (in domains like agentic coding) or broad branching trees when not confident (in tasks like creative writing). Verification is performed using a communication-free procedure based on Gumbel couplings, which, while having a theoretically suboptimal acceptance rate, facilitates high-quality draft trees through aggressive oversampling and heuristic-based pruning.
In contrast to previous GPU generations, M5 chips achieve peak arithmetic throughput in int8 precision. Consequently, Uzu employs 8-bit dynamic activation quantization during inference, in addition to 4-bit asymmetric (Mirai-M) and 8-bit symmetric (Mirai-L) weight quantization. Block-diagonal random Hadamard transforms are applied to activations before quantization to eliminate outliers and ensure near-lossless activation quantization.
Uzu's quantization pipeline comprises two stages: initial post-training quantization (PTQ) using a modified version of YAQA, a leading second-order quantization algorithm, followed by a concise quantization-aware distillation stage optimizing only the scales and zeropoints of quantization groups. The checkpoints are calibrated using a subset of the OpenHermes-2.5 dataset and trained on QAD sets, maintaining a optimal size-accuracy Pareto frontier, often matching or surpassing Unsloth checkpoints at comparable bits per weight rates.
The firm measures autoregressive output and input speeds, as well as resident memory consumption, using a fixed 1,355-token prompt. Performance is evaluated as a mean of three consecutive runs on random samples from MT-Bench (conversational prompts, 80 samples), MATH-500 (mathematical reasoning, 128 samples), and WeChat (reserved memory consumption).
The quantization quality is assessed by measuring the KL divergence relative to the bfloat16 teacher on a proprietary data mixture comprising 45% public agentic interaction logs, 30% public SFT data, and 25% private chat logs.
Written by urgent.news from Daily Dose of DS's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.