RoPE vs sinusoidal positional encoding, measured: 55 logits of drift against 0.0005
Take two token embeddings. Keep them exactly five positions apart and slide the pair down a 2048-token sequence. Under sinusoidal absolute position encoding the attention score for that one unchanging pair runs from -33.91 to +21.61 — a 55.5-logit spread that changes sign 157 times on the way. Under RoPE the same pair scores -0.6102 every single time , to within 5.4e-4. That is the whole argument…
RoPE and sinusoidal positional encoding were evaluated with 55 logits of drift against 0.0005. Two token embeddings were kept five positions apart and slid down a 2048-token sequence. Sinusoidal absolute position encoding produced a 55.5-logit spread that changed sign 157 times. Rotary embeddings kept the same pair scoring -0.6102 every time, with only a 5.4e-4 difference.
The code, output, and details of the calculation were provided, showing no positional encoding at all produced a score of -5.8183. The sinusoidal encoding showed fluctuating logit values, while RoPE remained consistent at -0.6102. The gap-0 rows served as a sanity check, with RoPE matching the no-positional-encoding score at various positions. The worst same-distance spread across various gaps was 7.458e-04 absolute and 2.658e-03 relative.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.