Memory ordering in CPUs
Memory ordering plays a crucial role in the performance and scalability of CPUs, particularly when comparing strongly ordered architectures like x86 or SPARC to weakly ordered ones such as ARM or RISC-V. Many believe that weakly ordered machines inherently offer better scalability, but this is a misconception based on the assumption that CPUs strictly adhere to their memory models for every memory access.
In reality, some CPU cores do follow these rules meticulously, but most implement memory ordering optimistically, making assumptions about data consistency and avoiding contention.
Optimistic memory ordering assumes that most loads access data that hasn't been modified recently, and most stores don't contend for the same cache line. Consequently, out-of-order CPUs can execute these operations in any order, as long as they commit after ensuring no conflicts. This "trust, but verify" approach allows CPUs to execute instructions assuming they are uncontended most of the time, keeping only enough metadata to detect potential memory ordering violations after the fact.
When conflicts are detected, offending instructions are rolled back, the program state is reset, and the instructions are retried.
The key difference between strongly and weakly ordered machines lies not in whether they execute memory operations strictly in order or not, but in how they handle contention. Strongly ordered machines are more likely to report conflicts and retry when encountering contention, whereas weakly ordered machines are generally more permissive. Weakly ordered machines do have some advantages in this regard, but these benefits may not translate to significant real-world improvements in multi-threaded workloads.
In summary, the main distinction between strongly and weakly ordered machines is their approach to handling contention. Strongly ordered machines are more cautious, requiring additional metadata to check for ordering violations, while weakly ordered machines deal mostly with relaxed loads and stores that only need to be ordered with respect to memory barriers.
This nuanced difference means that both approaches have their costs, and measuring the performance impact can be challenging. Ultimately, the effectiveness of memory ordering strategies depends on the specific workload and the hardware's design.
Written by urgent.news from Lobsters's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.