{
  "id": 598758,
  "title": "DFlash Changes What Tokens per Second Means",
  "url": "https://urgent.news/2026/08/11/dflash-changes-what-tokens-per-second-means",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-08-11T20:28:31.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/pich/dflash-changes-what-tokens-per-second-means-4493"
  },
  "original_language": "en",
  "account": "DFlash, a lightweight block-diffusion model, has transformed the way we measure tokens per second by introducing a new benchmark: predictability. By incorporating DFlash, developers can utilize a dense 30B model with a 262,144-token context on a 24GB GPU, achieving remarkable throughput improvements. On a coding task, DFlash can generate 84.64 tokens per second, while mixed workload tasks reach 38.34 tokens per second. This change in perspective shifts the focus from pure hardware benchmarks to a more comprehensive predictability benchmark. The key to DFlash's effectiveness lies in its parallel block-diffusion approach, where a lightweight model drafts candidate tokens while the expensive autoregressive model verifies them. This allows the drafter to propose multiple output tokens with each full forward pass, amortizing the cost of verification across accepted prefixes. While speculative decoding still relies on the target model's authority for generating accurate output, it provides a more efficient way to leverage the expensive verification work. The impact of DFlash is evident when comparing different workloads. For code generation, DFlash achieves a 4.50 times speedup compared to regular decoding, while mixed workloads see a drop to 38.34 tokens per second. However, these differences highlight the importance of workload-specific considerations when evaluating throughput. The acceptance rate, a common metric, can sometimes lead to misleading conclusions. A shorter block length may achieve higher acceptance rates but at the expense of throughput. Instead, average acceptance length (τ) provides a more informative metric, indicating the number of output tokens generated per expensive verification pass. By accounting for factors like drafter latency, verification cost, batch shape, and KV-cache position, τ offers a more accurate representation of the true throughput potential. In summary, DFlash has not just made the Glimmer Muse Glimmer 30B model faster; it has fundamentally changed our understanding of what \"tokens per second\" truly means. By focusing on predictability and providing a more nuanced performance metric, DFlash enables developers to optimize their models for specific workloads, unlocking new levels of efficiency in demanding AI applications.",
  "summary": "I spent a night trying to fit a dense 30B model, 256K context, vision, and speculative decoding onto one 24 GB GPU. The fastest quant lost. The quant with the lowest perplexity lost too. What won was the configuration that made the whole system useful, not any single number impressive. The final setup runs Meta Muse Glimmer 30B on an NVIDIA RTX PRO 4000 Blackwell SFF capped at 70 watts. It holds…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}