Urgent.News

What's breaking now, across thousands of outlets.

Tech

How Is Compression Prediction?

Over the past few weeks, repeated claims have surfaced on Hacker News: compression is prediction. Two 3Blue1Brown videos and an ngrok article have explored this concept through entropy and language models. Salvatore Sanfilippo questions how far this identification should be taken.

A probabilistic model assigns a conditional probability to each possible continuation, and an entropy coder converts this probability into bits for a sequence x1:n. The ideal payload length is the cumulative logarithmic loss of the model, up to coding overhead. Improving prediction under log-loss and reducing the encoded payload are equivalent optimization problems.

This correspondence has roots in classical information theory. Shannon linked probability to optimal code length, adaptive statistical compressors used conditional estimates, and the link between learning and compression has been developed through minimum description length, MacKay's treatment of information theory, and work like the Hutter Prize. Recent language-model results simply extend this older correspondence to a new scale.

While the equivalence is valid, the author wonders about its beginning and end points. Compression involves encoding data under an agreed model, but the problem begins before the model is applied and does not always end when the shortest bitstream is produced. Both encoder and decoder must agree on the objects being represented, the possibilities remaining, how the model is made available, and what the decoder must do with the representation.

Before introducing a sequential model, compression can be defined using a finite family of admissible objects, giving a counting lower bound. A code can then be interpreted probabilistically, and a distribution over serialized objects can be factored into next-symbol conditionals. This reinterpretation does not choose the object family, pay for unavailable information, or enforce operations like random access.

The core question is not whether prediction and compression are mathematically equivalent but what must be fixed before the equivalence applies, which aspect of the representation's bit count is meaningful, and what remains outside that measurement.

This article assumes familiarity with undergraduate mathematics and elementary proof-style arguments, but prior background in information theory is not required. The ngrok article differentiates between minification and "true" compression. Minification removes non-executable parts of a source file, creating a shorter version that cannot reconstruct the original.

Whether this operation is lossless depends on what the representation must preserve. If the object is the original source bytes, minification is lossy. However, if the object is the program's behavior, a semantics-preserving minifier yields a lossless result. The distinction precedes any probability model, as the encoder and decoder must first agree on what counts as the object and how to define equivalence among decoded outputs. Only then does description length become meaningful.

Written by urgent.news from Lobsters's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at lukefleed.xyz →

More in Tech

More from Saturday 15 August →