Byte Language Models: Scaling, Emergent Abstractions, and Information Allocation
The paper challenges the assumption that language models need explicit tokenizers to be efficient demonstrating that standard flat Transformers can process raw byte sequences and actually outperform traditional subword models as parameter sizes scale. The prevailing thought in the field has been that processing raw bytes is computationally inefficient because the sequences are substantially…
We haven't written up this one. Lobsters has the full story — the link below goes straight to it.