The Flag We Tuned Around Got Deleted
The single most important llama.cpp flag for my dual Tesla P40 setup was -sm row . It split every layer's tensors across both GPUs and it was worth nearly double the throughput of the alternative: 12-14 tokens/sec against about 7 for layer split. Every stack I built was tuned around it. In July 2026, upstream llama.cpp deleted it. Not deprecated. Deleted. This is the story of a performance rule…
For a dual GPU Tesla P40 setup, the most crucial llama.cpp flag was "-sm row". This flag divided every layer's tensors between both GPUs, resulting in nearly twice the throughput compared to the alternative "layer split". Every model stack I built was optimized around "-sm row". However, in July 2026, the upstream llama.cpp project deleted the flag entirely, deleting a significant performance rule that had been in place twice.
This story covers five acts: Act 1 - Row wins in February 2026, Act 2 - Regression incident in March 2026, Act 3 - Fast mode becomes the wrong mode in May 2026, Act 4 - Upstream deletes row in July 2026, and Act 5 - The win comes back from elsewhere in the present.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.