{
  "id": 12675722,
  "title": "I asked an LLM which listings to skip. It kept getting the numbers wrong.",
  "url": "https://urgent.news/2026/10/07/i-asked-an-llm-which-listings-to-skip-it-kept-getting-the-numbers",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-07T17:40:09.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/gk_019a1c50f0a200e67225e0/i-asked-an-llm-which-listings-to-skip-it-kept-getting-the-numbers-wrong-4p41"
  },
  "original_language": "en",
  "account": "I spent a considerable amount of time scanning pages to locate the single line that disqualifies them. A job posting where the budget is hidden at the bottom. Headphones that are brand new, yet the seller, three scrolls down, notes they are refurbished. So, I thought, easy, I'll paste the page into an AI language model (LLM) with my own rules and ask should I bother with this listing?\n\nAttempt 1: I simply asked the model, using a set of my own rules: the budget should be at least $500, with a maximum of 3 revisions allowed. I provided the job post, and the model responded with either a 'yes' or 'no', along with the reason. It worked most of the time, until it started getting the numbers wrong. It claimed a $300 job was acceptable, despite the real price being $129. The model confidently stated a product had a 30-day return window on a page that never mentioned returns. The issue wasn't that it was wrong occasionally; it was that it sounded absolutely certain when it was incorrect.\n\nAttempt 2: To address the problem, I split the work between the model and myself. The plain code reads the numbers (price, budget, days, counts) and compares them to my set limits. If any of these values fail to meet the criteria, the answer is 'no', and the model never receives the request. The model scans the same page separately, but it only considers a value valid if that specific text is present on the page. If the code and model disagree, the model presents both findings to me, and I decide which to follow. When the model encounters missing facts, it doesn't guess; instead, it asks me for the value instead of providing an answer. The model only judges the softer aspects, such as whether the scope is realistic for the budget or if there's a catch, only after the numerical values pass the initial checks. It has to quote the specific line it's basing its assessment on.\n\nSome issues I encountered included the List Price: $79 and Was $99 sitting right next to the actual price. The parser had to disregard any lines that were clearly not the price. The text 21 315 participants, where the space acted as a thousands separator, was being read as 21. $12 per user, billed monthly needed the 'per user' and 'billed' removed before it could determine the billing frequency. The model sometimes struggled, and when it was busy, it still checked the numbers and clearly stated that the AI part was missing, rather than simply failing.\n\nUltimately, I wrapped the program in a Chrome extension called Kriterio. Users can set their limits once, open a page, and the extension displays it in green, amber, or red, along with the reason. I primarily use it for job postings, shopping, and other applications where numeric accuracy is crucial. It's free to try, and feedback is welcome: https://chromewebstore.google.com/detail/phmlfimaombpolnlcmgohocekndiinpf https://getkriterio.com/ I'm curious to know how others handle this situation—are you allowing the model to decide anything numeric, or do you always check it in code?",
  "summary": "I waste a stupid amount of time reading pages just to find the one line that rules them out. A job post where the budget is hidden at the very bottom. \"Brand new\" headphones where the seller note, three scrolls down, says refurbished. A So I thought, easy, I'll paste the page into an LLM with my own rules and ask \"should I bother with this?\" Attempt 1: just ask the model Something like: My rules:…",
  "key_points": [
    "User created AI tool to check listing accuracy",
    "Tool splits tasks between AI and human for numeric validation",
    "Extension displays listings in green, amber, or red based on checks"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}