{
  "id": 13340777,
  "title": "Put a Meter on Your AI Feature Before Someone Else Runs Up the Bill",
  "url": "https://urgent.news/2026/10/10/put-a-meter-on-your-ai-feature-before-someone-else-runs-up-the-bill",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-10T06:16:00.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/posting-dude/put-a-meter-on-your-ai-feature-before-someone-else-runs-up-the-bill-df4"
  },
  "original_language": "en",
  "account": "When deploying an AI feature on a free plan, it's important to put a meter on usage before someone else racks up unexpected costs. Here are the steps to follow, in the order I'd implement them:\n\n1. Set a hard spend cap at the provider, lower than feels comfortable. Configure a monthly hard limit and a soft alert at around 50% of that limit. The cap should be a number that would annoy you, but not ruin you. When the cap is reached, the feature should shut down and you'll receive an email. Create a separate project/key for each environment to avoid surprises.\n\n2. Implement a per-user daily quota counted in tokens, not requests. One request with a large pasted PDF costs more than multiple smaller ones. Record user ID, input tokens, output tokens, model used, and timestamp for each call. Before each new call, check the total for that user in the last 24 hours. Free users get a small token budget, paid users get a larger one, and nobody gets unlimited access.\n\n3. Cap the input before it reaches the model. Truncate or reject input that exceeds a size limit, and set a maximum number of output tokens for every call. If you allow users to send whatever they want, they might paste an entire document, causing the model to generate a lengthy response.\n\n4. Rate limit by IP and by account. Quota is measured per day, while rate limiting is done per minute. A script can make several calls per second, which is much faster than a human using the normal UI. Limit to a few calls per minute per account (and per IP for signed-out access) to prevent a loop from running for an extended period.\n\n5. Include a daily cost line in your morning check. Add the model spend from yesterday and the top 3 users by tokens to your morning glance. It takes only ten seconds. If any single user's cost is 10 times higher than everyone else's, investigate further.\n\n6. Create a clear message when someone hits their quota. Show something like, \"You've used today's free summaries, they reset at midnight UTC. Paid plans get more.\" This message serves as both an honest upgrade prompt and a clear indication that the free user has reached their limit.\n\nRemember, treat every model call like it costs real money, as it does, and it will be billed to you, not the user. Implement the meter first, then deploy the feature. This approach not only keeps your side project cheap but also works well with a general cost checklist for maintaining monthly bills under $20. If you've built an AI tool and want early users who will promote it positively, consider listing it on EarlyHunt when you're ready.",
  "summary": "The scary thing about an AI feature on a free plan is that the bill goes to you, not the user. One person with a loop script against your endpoint can burn through a month of budget while you sleep, and they're not even really a hacker, it's just cheaper than paying for the model themselves. None of what follows is clever. It's the boring stuff I'd ship before the AI feature goes live, in the…",
  "key_points": [
    "Set hard spend cap at provider, lower than comfortable",
    "Implement per-user daily token quota, check before each call",
    "Cap input size and output tokens, rate limit by IP/account"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}