tokenizers v1: encode, decode and scaling, measured
Tokenization, the process of converting text into a list of integers a machine learning model can understand, has been a crucial step in many ML workflows. Historically, it has not been the bottleneck, as computation-intensive modeling tasks dominate. However, as models become faster and workloads scale, tokenization has emerged as a key factor in accelerating or slowing down ML work.
To address this, the upcoming version 1 of tokenizers places a strong emphasis on performance, aiming to make tokenization light and scalable with your workflow.
In tokenizers v1, significant improvements have been made to the underlying BPE (Byte Pair Encoding) models, which are responsible for the majority of tokenization work. These enhancements include:
1. Fixed regular expressions for pre-token splitting: Instead of using a general-purpose regex engine at runtime, v1 employs a hand-written function optimized for the specific pattern used by each model. This allows for SIMD (Single Instruction, Multiple Data) instructions to process 64 bytes in a single operation, significantly speeding up the tokenization process.
2. Caching of pre-token IDs: Once a pre-token has been processed, v1 saves the resulting token IDs in a thread-local cache. This allows subsequent occurrences of the same pre-token to skip the merge process, reducing computational overhead. The effectiveness of this caching depends on the presence of repeated pre-tokens in the input data.
3. Optimization of the BPE merge loop: In the previous implementation, the merge loop allocated new memory for each pre-token and built a new priority queue, resulting in repeated allocations. v1 addresses this by reusing a scratch buffer owned by the caller, eliminating unnecessary memory allocations and improving overall performance.
These optimizations are not limited to BPE models; other model families such as WordPiece and Unigram also benefit from these changes. The tokenization pipeline remains unchanged, preserving the output, API, vocabulary, and merge ranks of v0.23. The focus has been on improving every aspect that can be optimized, leading to substantial performance gains in various scenarios, including single-threaded, multi-threaded, and scaling across threads.
Written by urgent.news from Hugging Face's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.