Urgent.News

What's breaking now, across thousands of outlets.

AI

The Semantic Cache That Made a Free LLM Quota Feel Infinite

A token allowance is usually treated as a spending budget, which is the wrong mental model for free tiers. The right model is a cache to be managed, because agent workloads repeat themselves far more than developers realize. A semantic cache that serves previous responses for rephrased requests can cut token consumption by roughly half in typical agent loops. This article documents a working…

Token allowances for free language model tiers are often viewed as a spending budget, but this approach misses a crucial optimization opportunity. Rather than treating the allowance as something to be depleted, it's more effective to think of it as a cache to be managed efficiently. A semantic cache that serves previously generated responses for rephrased queries can reduce token consumption by up to half in typical agent workflows.

This article explains how to implement a zero-dependency semantic cache and fine-tune its parameters for safe production use. MonkeyCode's free offering currently provides a 10-million-token allowance and a free server option, making effective quota management a practical concern. The percentages and latency figures presented below are illustrative measurements from a controlled prototype, not guarantees of specific outcomes.

In most agent workflows, developers assume prompts are mostly unique, leading to unnecessary token consumption and latency. A closer look at request logs reveals that about 35% of prompts are semantically near-duplicates of previous answers. Sending duplicate prompts incurs unnecessary token costs, increased latency, and the risk of inconsistent responses when the model reformulates answers differently.

Instead of expecting the model to become smarter, the solution lies in preventing duplicate requests from reaching it in the first place. A semantic cache placed in front of the endpoint can recognize and return cached responses for similar prompts, reducing the number of calls to the model. The cache does not replace the model; it simply filters out redundant requests.

To implement a semantic cache, the first step is selecting an appropriate similarity metric and threshold. A zero-dependency option is character n-gram Jaccard similarity, which tokenizes text into three-character overlapping shingles and compares their sets. While more advanced methods like embedding models exist, they introduce additional dependencies and API costs.

A threshold of 0.92 was chosen through calibration of one hundred logged prompts, balancing the need to catch duplicates versus the risk of returning incorrect cached responses. The choice of similarity metric depends on prompt length, with character n-grams suitable for short prompts and word-based methods or embeddings potentially better for longer text.

After a response is retrieved from the cache, it should be stored with a time-aware eviction policy to prevent memory leaks. The prototype uses SQLite with a created_at timestamp and a TTL (time-to-live) of one hour, striking a balance between freshness and hit rate. Shorter TTLs protect against stale answers, while longer TTLs improve hit rates but risk serving outdated information.

The schema stores the original prompt, response text, token cost of the original call, creation time, and hit counter. The hit counter enables a popularity-aware eviction strategy, where the cache removes the oldest entries with the lowest hit counts when it exceeds a size limit. This strategy focuses on prompts that actually recur, rather than one-off requests.

The provided Python code demonstrates a complete cache implementation, including initialization, tokenization, similarity scoring, retrieval, and insertion. It is designed to be dependency-free, allowing it to run on any fresh Python environment, including free servers without package installation constraints.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

When a user uploads a video showing a person holding a deadly weapon, an automated system immediately flags the content for review, with an estimated 98-percent likelihood that it’s unsafe. TikTok’s content moderation system uses both artificial intelligence and human reviewers to assess videos and other content uploaded to the platform. The system uses AI to scan uploaded content and flag it for review if it appears to violate TikTok’s community guidelines or terms of service. TikTok’s community guidelines prohibit the platform from being used to promote violence, harassment, hate speech or other types of problematic content. In some cases, AI will flag content that’s actually acceptable, a phenomenon known as a “false positive.” In other cases, AI will fail to flag content that’s actually dangerous or otherwise problematic, a “false negative.” TikTok did not provide information on the accuracy rate of its AI system. Once content is flagged, human moderators review it to determine whether it violates TikTok’s guidelines. The company said it has teams of moderators around the world, fluent in local languages and knowledgeable about local laws and culture, to help ensure that content complies with TikTok’s guidelines and local regulations. TikTok said it removed more than 89 million videos from its platform between April and June for violating its community guidelines or terms of service. The company said that’s about 1 percent of all videos uploaded to TikTok during that period. The company said it also relies on users to report content they believe violates TikTok’s guidelines. Users can report content by tapping a button at the bottom of a video or by reporting it through TikTok’s support page. TikTok said it’s investing heavily in its content moderation efforts, including through the use of AI. The company said it’s also working with outside experts, such as mental health professionals, to help inform its approach to content moderation. The company’s content moderation efforts come as social media companies face growing scrutiny over their handling of violent or otherwise problematic content. Social media companies have long been criticized for not doing enough to stop the spread of hate speech, harassment and other types of problematic content on their platforms. TikTok’s parent company, ByteDance, did not provide a breakdown of the number of moderators it employs. However, the company said it has a “large workforce” dedicated to content moderation. The company said it’s also working to increase transparency around its content moderation practices. TikTok said it’s publishing a transparency report that will provide information on the number of videos removed from its platform and the reasons they were taken down. TikTok’s transparency report will also provide information on the number of accounts suspended or permanently banned from the platform. The company said it’s also exploring other ways to increase transparency around its content moderation practices. The move comes as lawmakers and regulators around the world are calling for more transparency around social media companies’ content moderation practices. In May, U.S. President Joe Biden and other world leaders called for new regulations to address the spread of misinformation and other types of problematic content on social media. TikTok said it’s committed to working with lawmakers and regulators to address concerns around content moderation. The company said it’s also committed to providing users with more information on its content moderation practices. TikTok said it’s investing in tools and technologies to help ensure that its platform is a safe and positive experience for users. The company said it’s also working to provide users with more control over the content they see on its platform. TikTok said it’s committed to transparency and accountability in its content moderation practices. The company said it’s committed to working with outside experts to help inform its approach to content moderation. TikTok’s content moderation efforts are part of a broader effort by social media companies to address concerns around problematic content on their platforms. Social media companies have been criticized for not doing enough to stop the spread of hate speech, harassment and other types of problematic content on their platforms. TikTok’s parent company, ByteDance, has said it’s committed to working with lawmakers and regulators to address concerns around content moderation. TikTok said it’s committed to providing users with a safe and positive experience on its platform. The company said it’s committed to transparency and accountability in its content moderation practices. TikTok’s efforts to address concerns around content moderation come as social media companies face growing scrutiny over their handling of violent or otherwise problematic content. The company’s approach to content moderation is part of a broader effort by social media companies to address concerns around problematic content on their platforms. TikTok said it’s committed to working with outside experts to help inform its approach to content moderation. The company said it’s committed to providing users with more information on its content moderation practices. TikTok’s content moderation efforts are part of a broader effort by social media companies to address concerns around problematic content on their platforms. The company said it’s investing in tools and technologies to help ensure that its platform is a safe and positive experience for users. TikTok said it’s committed to transparency and accountability in its content moderation practices. The company said it’s committed to working with lawmakers and regulators to address concerns around content moderation. TikTok’s parent company, ByteDance, has said it’s committed to working with lawmakers and regulators to address concerns around content moderation. The company said it’s committed to providing users with a safe and positive experience on its platform. TikTok said it’s investing heavily in its content moderation efforts, including through the use of AI. The company said it’s also working with outside experts, such as mental health professionals, to help inform its approach to content moderation. TikTok said it’s committed to working with outside experts to help inform its approach to content moderation. The company said it’s committed to providing users with more information on its content moderation practices. TikTok’s efforts to address concerns around content moderation come as social media companies face growing scrutiny over their handling of violent or otherwise problematic content. The company said it’s committed to transparency and accountability in its content moderation practices. TikTok said it’s investing in tools and technologies to help ensure that its platform is a safe and positive experience for users. The company said it’s also working to provide users with more control over the content they see on its platform. TikTok’s parent company, ByteDance, did not provide a breakdown of the number of moderators it employs. However, the company said it has a “large workforce” dedicated to content moderation. The company said it’s also working to increase transparency around its content moderation practices. TikTok said it’s publishing a transparency report that will provide information on the number of videos removed from its platform and the reasons they were taken down. The company said it’s also exploring other ways to increase transparency around its content moderation practices. Tik

More from Sunday 23 August →