Urgent.News

What's breaking now, across thousands of outlets.

AI

DFlash-2: Benchmarking Z-Lab's Successor to DFlash for Accuracy and Throughput Gains

A while back, we covered DFlash, a draft-token prediction technique that uses a diffusion model. At the time, we tested it on Gemma-4-12b-it-QAT, and the native Assistant model came out ahead — DFlash wasn't able to show a clear advantage. Recently, though, a successor called "DFlash-2" surfaced, with a number of enhancements on top of the original design. As of August 2026, only a handful of…

DFlash-2 is a successor to the DFlash draft-token prediction technique developed by Z-Lab. It builds upon the original DFlash design with two key enhancements: a Lightweight Path Selector and Local Convolution layers. The Path Selector checks the natural ordering of predicted token sequences, filtering out any inconsistencies. The Local Convolution layer limits information exchange to each token's immediate neighbors, addressing the accuracy drop-off at the end of a predicted token block.

DFlash-2 models are available for Qwen3.8-27B-DFlash2 and Muse-Glimmer-30B-DFlash2, with GGUF builds for llama.cpp. Although DFlash-2 support hasn't yet been integrated into llama.cpp's master branch, a pull request has been submitted. To run DFlash-2 with Muse-Glimmer-30B, clone the llama.cpp repository, apply the pull request, and then execute the llama-server command with the appropriate arguments for the DFlash-2 model file and other specified parameters.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

LLMs Pass the Data-Science Quiz, Then Give Different Advice: A Kaggle Benchmark of 36 Measured Judgment Calls

This is a submission for the Kaggle Benchmarking Challenge . What I Benchmarked I spend a lot of time in Kaggle tabular competitions, and the decisions that cost me the most were never about model…

  • LLMs achieved 94-100% accuracy in recognizing measured answers
  • Open-ended performance varied, with 56-81% correct recommendations
  • Four topics showed lower accuracy, including T02, T03, T11, and T07

I wanted a Cursor-style agent that runs on my own model, so I built one

Disclosure: I'm Ibrahim, the solo developer of OpenPilot. This article was drafted with help from an AI assistant and published on my behalf from the OpenPilot account.

  • Ibrahim built OpenPilot, an open-source Cursor-style agent
  • OpenPilot runs on users' own models and supports various platforms
  • Users control the agent via tool cards with token usage tracked

More from Friday 9 October →