Urgent.News

What's breaking now, across thousands of outlets.

AI

tokenizers v1: encode, decode and scaling, measured

Tokenization, the process of converting text into a list of integers a machine learning model can understand, has been a crucial step in many ML workflows. Historically, it has not been the bottleneck, as computation-intensive modeling tasks dominate. However, as models become faster and workloads scale, tokenization has emerged as a key factor in accelerating or slowing down ML work.

To address this, the upcoming version 1 of tokenizers places a strong emphasis on performance, aiming to make tokenization light and scalable with your workflow.

In tokenizers v1, significant improvements have been made to the underlying BPE (Byte Pair Encoding) models, which are responsible for the majority of tokenization work. These enhancements include:

1. Fixed regular expressions for pre-token splitting: Instead of using a general-purpose regex engine at runtime, v1 employs a hand-written function optimized for the specific pattern used by each model. This allows for SIMD (Single Instruction, Multiple Data) instructions to process 64 bytes in a single operation, significantly speeding up the tokenization process.

2. Caching of pre-token IDs: Once a pre-token has been processed, v1 saves the resulting token IDs in a thread-local cache. This allows subsequent occurrences of the same pre-token to skip the merge process, reducing computational overhead. The effectiveness of this caching depends on the presence of repeated pre-tokens in the input data.

3. Optimization of the BPE merge loop: In the previous implementation, the merge loop allocated new memory for each pre-token and built a new priority queue, resulting in repeated allocations. v1 addresses this by reusing a scratch buffer owned by the caller, eliminating unnecessary memory allocations and improving overall performance.

These optimizations are not limited to BPE models; other model families such as WordPiece and Unigram also benefit from these changes. The tokenization pipeline remains unchanged, preserving the output, API, vocabulary, and merge ranks of v0.23. The focus has been on improving every aspect that can be optimized, leading to substantial performance gains in various scenarios, including single-threaded, multi-threaded, and scaling across threads.

Written by urgent.news from Hugging Face's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at huggingface.co →

More in AI

The AI Interview Paradox: Decoupling Skill Assessment from Tool Usage

Originally published on tamiz.pro . The modern engineering interview pipeline is suffering from a critical integrity failure.

  • AI usage common in engineering interviews, contradicting hiring practices.
  • Traditional interviews test memorization, not modern engineering skills.
  • Proposed three-tiered assessment evaluates tool integration, architecture, and verification.

More from Monday 21 September →