Urgent.News

What's breaking now, across thousands of outlets.

AI

I asked an LLM which listings to skip. It kept getting the numbers wrong.

I waste a stupid amount of time reading pages just to find the one line that rules them out. A job post where the budget is hidden at the very bottom. "Brand new" headphones where the seller note, three scrolls down, says refurbished. A So I thought, easy, I'll paste the page into an LLM with my own rules and ask "should I bother with this?" Attempt 1: just ask the model Something like: My rules:…

I spent a considerable amount of time scanning pages to locate the single line that disqualifies them. A job posting where the budget is hidden at the bottom. Headphones that are brand new, yet the seller, three scrolls down, notes they are refurbished. So, I thought, easy, I'll paste the page into an AI language model (LLM) with my own rules and ask should I bother with this listing?

Attempt 1: I simply asked the model, using a set of my own rules: the budget should be at least $500, with a maximum of 3 revisions allowed. I provided the job post, and the model responded with either a 'yes' or 'no', along with the reason. It worked most of the time, until it started getting the numbers wrong. It claimed a $300 job was acceptable, despite the real price being $129.

The model confidently stated a product had a 30-day return window on a page that never mentioned returns. The issue wasn't that it was wrong occasionally; it was that it sounded absolutely certain when it was incorrect.

Attempt 2: To address the problem, I split the work between the model and myself. The plain code reads the numbers (price, budget, days, counts) and compares them to my set limits. If any of these values fail to meet the criteria, the answer is 'no', and the model never receives the request. The model scans the same page separately, but it only considers a value valid if that specific text is present on the page.

If the code and model disagree, the model presents both findings to me, and I decide which to follow. When the model encounters missing facts, it doesn't guess; instead, it asks me for the value instead of providing an answer. The model only judges the softer aspects, such as whether the scope is realistic for the budget or if there's a catch, only after the numerical values pass the initial checks. It has to quote the specific line it's basing its assessment on.

Some issues I encountered included the List Price: $79 and Was $99 sitting right next to the actual price. The parser had to disregard any lines that were clearly not the price. The text 21 315 participants, where the space acted as a thousands separator, was being read as 21. $12 per user, billed monthly needed the 'per user' and 'billed' removed before it could determine the billing frequency.

The model sometimes struggled, and when it was busy, it still checked the numbers and clearly stated that the AI part was missing, rather than simply failing.

Ultimately, I wrapped the program in a Chrome extension called Kriterio. Users can set their limits once, open a page, and the extension displays it in green, amber, or red, along with the reason. I primarily use it for job postings, shopping, and other applications where numeric accuracy is crucial. It's free to try, and feedback is welcome: https://chromewebstore.google.com/detail/phmlfimaombpolnlcmgohocekndiinpf https://getkriterio.com/ I'm curious to know how others handle this situation—are you allowing the model to decide anything numeric, or do you always check it in code?

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

🌿 FrostBite: Offline AI Garden Scout

🌿 FrostBite: Offline Open-Weight Garden Scout This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass 📖 What I Built FrostBite is a lightweight, zero-cloud CLI tool…

  • FrostBite is a CLI tool for gardeners to disconnect from screens.
  • Tool runs offline using local USDA data and smollm2:1.7b model.
  • Benefits include privacy, accessibility, customization, and resilience.

Wednesday assorted links

1. Palo Alto Networks (cybersecurity firm, check out YTD). 2. OAI doing math again. Quasi-Riemann! And just one metric of import. 3. The AI agents pitching literary magazines. 4.

Mega AI Battle: Benchmarking 6 Top LLMs with Advanced Bangla Logic Riddles

Mega AI Battle: Benchmarking 6 Top LLMs with Advanced Bangla Logic Riddles 🎯 Hi everyone! I am thrilled to share my project for the Kaggle Benchmarking Challenge .

  • Google Gemini excels in reasoning, correctly solving the Circular Spatial Trap with 12 chairs
  • All models fail Linguistic Semantic Trap, mistaking multiplication for addition
  • Google Gemini leads with 5/5, others score 2-3/5 in advanced Bangla logic riddles

More from Wednesday 7 October →