Urgent.News

What's breaking now, across thousands of outlets.

Tech

Performance impact of Alignment

Performance impact of alignment is a crucial aspect of SIMD vectorization, affecting both auto-vectorization and explicit vectorization (like Vector API). An address is said to be aligned if the remainder of the address divided by the alignment_size is zero. When no alignment_size is specified, it is usually assumed to be the size of the access.

Cacheline-alignment, typically a 64-byte alignment, is frequently mentioned in discussions of vectorizing code. The alignment of memory or specific memory addresses is also relevant. For instance, Java Objects have an 8-byte alignment.

Array elements are aligned by their element size. Therefore, accessing a single element in an array is always aligned. However, loading a vector of elements from an array might not be guaranteed to be aligned. For example, loading a 16-element vector from an array would require a 64-byte alignment, as each int is 4 bytes long.

When moving from scalar to vector code, additional work is required to ensure vector accesses are aligned. On some platforms, especially older ones, unaligned access may result in errors or incorrect execution. However, most modern CPUs allow unaligned access, often with performance only slightly slower than aligned access.

Split accesses, occurring when a memory access crosses a cacheline boundary, can slow down execution as they increase the number of memory accesses. The impact on performance depends on the fraction of split accesses and whether memory accesses are a bottleneck.

I conducted benchmarks to visualize the performance impact of alignment. On my AVX512 laptop with 64-byte (16 ints) vectors, aligned loads and stores (green) yielded the best performance, while misaligned accesses (red) resulted in the worst performance. The performance difference between the best and worst cases was approximately 50%. If only one could choose to align loads or stores, aligning stores would provide a 20% performance improvement.

The pattern holds true for different vector lengths and CPUs. Store-alignment generally leads to better performance.

Written by urgent.news from Lobsters's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at eme64.github.io →

More in Tech

More from Tuesday 11 August →