Urgent.News

What's breaking now, across thousands of outlets.

Tech

How Many Tokens Is That Elasticsearch Hit? A Reproducible RAG Compression Benchmark

This is a follow-up to my earlier post introducing jtoken . That one was the pitch. This one is the actual measurement — the benchmark I run before claiming any savings number, and how you can run it on your own payloads. The question When your RAG pipeline pulls documents into a prompt, how much of the context window is you asked for this vs syntax overhead ? And when someone (including me)…

This post provides a reproducible benchmark for evaluating the token savings achieved through the jtoken compression library, when used within a Retrieval-Augmented Generation (RAG) pipeline. The key points are:

1. The benchmark uses real tokenizers, not characters, to accurately measure the impact of compression on model inputs. tiktoken, the library used, is easy to install and uses the same tokenization as GPT-4o-class models.

2. The methodology involves generating 50 realistic documents in various formats (ES e-commerce hits, Mongo extended-JSON activity docs, deeply-nested SaaS API events) and measuring the token count for both pretty JSON and jtoken representations. The script verifies that the compression process does not lose any data by round-tripping the encoded data back to its original format.

3. The benchmark results show that jtoken achieves significant token savings compared to pretty JSON, particularly for MongoDB documents (19.1% reduction) and nested API events (13.2% reduction). However, the savings are more modest for highly-prose-heavy documents (11.4% reduction).

4. The post explains why jtoken performs particularly well on repetitive machine-generated JSON, where the same values are repeated across many documents. It also highlights that the library is most effective when you need the data back, as the compression provides a more compact representation for programmatic consumption by LLMs.

5. The benchmark includes a case where a bug was discovered in jtoken 0.3.5, where encoded dates were not properly decoded. This issue was fixed in jtoken 0.3.6, and the round-trip verification process also revealed a packaging bug in conda-forge's Windows CI.

6. The author encourages readers to run the benchmark on their own data to get a personalized metric. They provide a simple command to clone the jtoken repository and run the benchmark script on their payloads. The repo also includes a text file (llms.txt) for programmatic discovery of the library.

In summary, the post demonstrates that jtoken can provide substantial token savings in a RAG setup, particularly for structured, repetitive JSON data. The key takeaway is that automated verification through round-trip tests is essential to ensure data integrity and catch potential bugs in the compression process.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

Turning Noisy Press Pages Into Hindsight-Backed Competitor Trend Reports

Every Monday, my market-intelligence pipeline could correctly tell me that a competitor had announced a pricing change. What it could not tell me was whether that was actually new.

  • Market-intelligence pipeline integrates Hindsight as long-term memory layer
  • Pipeline distinguishes between new competitor changes and price trend repetitions
  • Hindsight enables competitor-specific history recall and market-level reasoning

We Didn’t Have a Product Problem. We Had a Memory Problem.

Building PRATHIDHWANI: A Memory-Powered Customer Feedback Intelligence System Using Hindsight A few months ago, I was sitting in a product discussion when someone asked a deceptively simple question…

  • Team discovered lack of clear memory of past decisions hindered addressing customer feedback
  • PRATHIDHWANI system connects customer feedback, product changes, and historical context
  • Dashboard provides visual overview of feedback trends and historical signals

I Built a Support Agent That Remembers Customers Between Chats

A support agent can give the right answer and still make a customer repeat themselves. I built this system around a simple idea: use long-term memory to carry useful context between conversations, but…

  • Memory layer uses Hindsight with unique bank per customer for retrieval.
  • Retrieval is tailored to current query, preventing random reiteration of stored information.

Cancelled customers who keep access: the Stripe webhook bug nobody reports

If you take subscriptions through Stripe, you keep two copies of "who is paying": Stripe's, and a column in your own database ( subscribed , is_pro , plan ...) that your app actually checks.

  • Stripe webhook bug allows canceled customers to retain access
  • Webhook only handles checkout.session.completed event, no revocation
  • Failure to replay failed events can result in permanent loss

More from Tuesday 29 September →