Urgent.News

What's breaking now, across thousands of outlets.

Tech

9 Triton kernels in 9 weeks: what PyTorch taught me by winning

I'm Richael (dh8116), a Year 11 student in Auckland. Since late July I've written one Triton kernel a week and benchmarked each against PyTorch, publishing the numbers even when PyTorch won, which was often. Here's what nine weeks taught me, kernel by kernel. The ones that won Fused softmax (week 2). Naive softmax makes three passes over memory. Fusing them into one took it from ~55 GB/s to ~230…

Richael, a Year 11 student from Auckland, documented their experience of writing a Triton kernel each week for nine consecutive weeks and comparing its performance against PyTorch. Each kernel's results were published, even when PyTorch outperformed the Triton kernel. Here's a week-by-week breakdown of what Richael learned from this experiment.

The first week's kernel focused on fused softmax. Softmax operations typically involve three passes over memory. By fusing these passes into a single operation, the Triton kernel's performance improved significantly, from around 55 GB/s to 230 GB/s, a four-fold increase. Notably, this version of the kernel outperformed PyTorch's own implementation. Richael's key takeaway from this lesson was that on memory-bound operations, minimizing the number of passes is crucial for speeding up the computation.

In the seventh week, Richael tackled fused cross-entropy. This kernel was designed to perform both the forward and backward passes in a single operation, with gradients written directly over the logits buffer. This approach eliminated the need for a separate [N, V] tensor that PyTorch's implementation created. On a T4 GPU with a vocabulary size of 131,072 in fp16, this kernel achieved a 15.90 ms forward pass time compared to 24.03 ms for PyTorch (a 1.51x improvement), as well as reduced peak memory usage by 1.67x.

Richael discovered the importance of comparing kernels against each other rather than against an unfused PyTorch baseline, as this helped him avoid inaccuracies in his performance assessments.

Week six introduced RoPE (relative positional encoding), a technique that demonstrated not only the elegance of the solution but also its computational efficiency. The backward pass could reuse the exact same kernel as the forward pass, with the only difference being a sign flip on the sine function. Richael appreciated this kernel for its simplicity and efficiency, even though it didn't offer a significant speedup compared to PyTorch.

In week three, Richael evaluated Flash attention, which performed correctly but was about 25 times slower than PyTorch on the author's T4 GPU. The implementation was not optimized for tensor cores, which could have led to better performance. The kernel's loss was attributed to the compiler's inability to generate tensor-core instructions for the given layout.

Matmul, the fourth kernel, displayed an interesting scenario where the output matched between Triton and PyTorch, but the throughput remained flat around 1 TFLOPS while cuBLAS reached 38 TFLOPS. Upon examining the compiled PTX code, Richael found that there were no tensor-core instructions emitted for the Triton layout, suggesting that the compiler might not have optimized the tensor operations as expected.

The eighth kernel, fused linear + cross-entropy, aimed to reduce the memory footprint by chunking the projection of the language model head (lm_head) into the loss calculation. However, despite the reduction in activation memory by about 10x, this kernel was slower than PyTorch on all benchmark sizes. Richael's conclusion was that memory considerations might sometimes outweigh performance gains, as the benefit of reduced memory usage did not translate into better overall speed.

Finally, in week nine, Richael explored fused SwiGLU MLP, which competed with torch.compile on time. The kernel's primary advantage lay in memory management, where it held only two large activation tensors during forward and backward passes, compared to four tensors held by eager execution. While this kernel showed a memory advantage, Richael emphasized the importance of considering the trade-offs when benchmarking fused kernels against fused PyTorch implementations.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

I Built an Outdoor Break for the Five Minutes You Already Have

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass What I Built A few people tried my first version of Out There and told me it assumed a park and gave them little…

  • App Out There helps users find moments of pause in their day.
  • Users specify time frame (5, 10, or 15 minutes) and familiar setting.
  • App generates observation question and three steps using open-weight model.

To Retry or Not to Retry: Exponential Backoff, Jitter and Idempotency Keys Done Right

Hầu như dev nào cũng từng viết một vòng for i in range(3) bọc quanh một HTTP call, rồi tự tin là hệ thống đã "resilient". Mình cũng vậy, cho đến một đêm payment gateway của đối tác chậm khoảng 2 giây.

  • Retry decisions should consider error temporariness, idempotency, and remaining retry budget.
  • Use exponential backoff with jitter and libraries like tenacity for retry logic management.

More from Thursday 8 October →