Urgent.News

What's breaking now, across thousands of outlets.

More in AI

They Put 7 Attention Mechanisms on a Latin Square. Then Removed Them One by One.

Since GPT, nearly every Transformer repeats the same attention mechanism at every layer. Forty-eight identical blocks, differing only in learned weights. Nobody tested that.

  • Seven attention mechanisms arranged in a Latin square
  • Removing mechanisms one by one had negligible effect on performance
  • At least one mechanism from a different family crucial for maintaining performance

More from Friday 18 September →