{
  "id": 12913621,
  "title": "9 Triton kernels in 9 weeks: what PyTorch taught me by winning",
  "url": "https://urgent.news/2026/10/08/9-triton-kernels-in-9-weeks-what-pytorch-taught-me-by-winning",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-10-08T17:32:02.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/richael_42/9-triton-kernels-in-9-weeks-what-pytorch-taught-me-by-winning-3jim"
  },
  "original_language": "en",
  "account": "Richael, a Year 11 student from Auckland, documented their experience of writing a Triton kernel each week for nine consecutive weeks and comparing its performance against PyTorch. Each kernel's results were published, even when PyTorch outperformed the Triton kernel. Here's a week-by-week breakdown of what Richael learned from this experiment.\n\nThe first week's kernel focused on fused softmax. Softmax operations typically involve three passes over memory. By fusing these passes into a single operation, the Triton kernel's performance improved significantly, from around 55 GB/s to 230 GB/s, a four-fold increase. Notably, this version of the kernel outperformed PyTorch's own implementation. Richael's key takeaway from this lesson was that on memory-bound operations, minimizing the number of passes is crucial for speeding up the computation.\n\nIn the seventh week, Richael tackled fused cross-entropy. This kernel was designed to perform both the forward and backward passes in a single operation, with gradients written directly over the logits buffer. This approach eliminated the need for a separate [N, V] tensor that PyTorch's implementation created. On a T4 GPU with a vocabulary size of 131,072 in fp16, this kernel achieved a 15.90 ms forward pass time compared to 24.03 ms for PyTorch (a 1.51x improvement), as well as reduced peak memory usage by 1.67x. Richael discovered the importance of comparing kernels against each other rather than against an unfused PyTorch baseline, as this helped him avoid inaccuracies in his performance assessments.\n\nWeek six introduced RoPE (relative positional encoding), a technique that demonstrated not only the elegance of the solution but also its computational efficiency. The backward pass could reuse the exact same kernel as the forward pass, with the only difference being a sign flip on the sine function. Richael appreciated this kernel for its simplicity and efficiency, even though it didn't offer a significant speedup compared to PyTorch.\n\nIn week three, Richael evaluated Flash attention, which performed correctly but was about 25 times slower than PyTorch on the author's T4 GPU. The implementation was not optimized for tensor cores, which could have led to better performance. The kernel's loss was attributed to the compiler's inability to generate tensor-core instructions for the given layout.\n\nMatmul, the fourth kernel, displayed an interesting scenario where the output matched between Triton and PyTorch, but the throughput remained flat around 1 TFLOPS while cuBLAS reached 38 TFLOPS. Upon examining the compiled PTX code, Richael found that there were no tensor-core instructions emitted for the Triton layout, suggesting that the compiler might not have optimized the tensor operations as expected.\n\nThe eighth kernel, fused linear + cross-entropy, aimed to reduce the memory footprint by chunking the projection of the language model head (lm_head) into the loss calculation. However, despite the reduction in activation memory by about 10x, this kernel was slower than PyTorch on all benchmark sizes. Richael's conclusion was that memory considerations might sometimes outweigh performance gains, as the benefit of reduced memory usage did not translate into better overall speed.\n\nFinally, in week nine, Richael explored fused SwiGLU MLP, which competed with torch.compile on time. The kernel's primary advantage lay in memory management, where it held only two large activation tensors during forward and backward passes, compared to four tensors held by eager execution. While this kernel showed a memory advantage, Richael emphasized the importance of considering the trade-offs when benchmarking fused kernels against fused PyTorch implementations.",
  "summary": "I'm Richael (dh8116), a Year 11 student in Auckland. Since late July I've written one Triton kernel a week and benchmarked each against PyTorch, publishing the numbers even when PyTorch won, which was often. Here's what nine weeks taught me, kernel by kernel. The ones that won Fused softmax (week 2). Naive softmax makes three passes over memory. Fusing them into one took it from ~55 GB/s to ~230…",
  "key_points": [
    "Richael documented Triton kernel development weekly for nine weeks",
    "Fused softmax kernel improved performance four-fold over PyTorch",
    "Fused cross-entropy kernel reduced memory usage 1.67x compared to PyTorch"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}