Write Like It's 1866: LLMs Relearn Telegraphese
In a groundbreaking study conducted in 1866, researchers explored the potential of LLMs, or large language models, to process and comprehend compressed records, akin to how telegraph messages were transmitted. The experiment utilized a 50-passage benchmark, with each passage consisting of around 1,300 questions. Two primary roles were assigned to the models: readers and writers.
The readers were to answer questions based on the plaintext records, while the writers were tasked with answering questions based on compressed records.
The researchers found that the models exhibited remarkable proficiency in comprehending the compressed records, with recovery ratios ranging from 0.99 to 1.10. This means that the models' answers using the compressed records were as accurate, if not better, than their answers based on the plaintext records. Notably, the GLM-5.3-Flash model, which serves as both the writer and reader in this study, achieved an impressive 48.4% savings in token usage when responding to questions based on compressed records.
This indicates significant efficiency gains when utilizing compressed records compared to plaintext.
The study also revealed that the compression technique, referred to as "cablese," did not negatively impact the models' performance. In fact, the compressed records were read slightly better than the regular responses, most likely due to compression's ability to suppress copy-the-record phrasing, which tends to negatively impact grading.
Every model tested, including four families that had never encountered an example of the compressed records, demonstrated the capability to answer questions from compressed records as well as or better than from plaintext records. This signifies that the ability to process compressed records was already ingrained in the models' training data, inherited from a century and a half of communication under metered bandwidth.
The researchers concluded that this technique of relearning telegraphese, or processing compressed records, is not a novel construct invented by the researchers. Instead, it is a capability that exists within the models' training data, which they inherited from their predecessors who communicated under the constraints of limited bandwidth.
The study emphasizes the importance of understanding the per-model spread in compression efficiency, as it varies from 25% to 49% under identical instructions. However, the researchers also noted that this efficiency spread is knowable, allowing users to make informed decisions when employing this technique.
Written by urgent.news from Hacker News's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.