{
  "id": 12356972,
  "title": "Can an AI catch the catch? I benchmarked 14 models on bounty fine print",
  "url": "https://urgent.news/2026/10/06/can-an-ai-catch-the-catch-i-benchmarked-14-models-on-bounty-fine-print",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-06T10:38:43.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/shenjun93/can-an-ai-catch-the-catch-i-benchmarked-14-models-on-bounty-fine-print-19hb"
  },
  "original_language": "en",
  "account": "A report was released on the benchmarking of 14 AI models to assess their ability to read and interpret the fine print of bounty listings. The author, a solo developer in Vietnam, engaged in a few weeks of online work hunting such as bug bounties, hackathons, open-source bounties, and community contests. The difficulty wasn't in the tasks themselves, but in deciphering the fine print of the listings.\n\nThe benchmark involved 40 short listings, each with a question and one correct answer from a fixed set. The cases were fictional but modeled on real-life listings encountered during the author's online work hunting. The four categories of cases included hidden gates, real deadlines, eligibility, and still open. The authors used a structured schema for the models to answer the questions and compared their answers to the key.\n\nOut of the 14 models tested, six scored 40/40 correct answers, indicating their ability to catch the catch in the bounty fine print. These models were Claude Opus 5.5, GPT-6.1 Sol, Gemini 3.8 Flash, Gemini 3.7 Flash, Gemini 3 Flash, and Gemma 4 31B (open weights). The remaining eight models had varying degrees of success, with some missing one or more cases.\n\nThe report highlights several key findings. Firstly, the benchmark separates the cheap tier of models, with six models achieving a perfect score of 40/40. Of the models that missed cases, four were from the cheap tier, including Haiku, nano, Flash-Lite, and Qwen 3 Next. Secondly, the word \"free\" proved to be the most effective trap, with six models falling for it by answering \"none\" instead of the correct answer. Thirdly, the case where the bounty issue is open with no assignee but a linked pull request is approved and the maintainer wrote \"merging after the release freeze\" proved to be the next hardest case, with five models answering incorrectly. Lastly, timezones and format failures can lead to missing deadlines, with Haiku and Flash-Lite missing the deadline by converting the date once instead of twice.",
  "summary": "This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked I'm a solo developer in Vietnam. For the last few weeks I've been hunting paid work online: bug bounties, hackathons, open-source bounties, community contests. The hardest part isn't the work. It's reading the listing. The headline says \"free\" , \"open\" , \"global\" , \"deadline Oct 11\" . Then the fine print says: the free…",
  "key_points": [
    "Six AI models scored 40/40 correct answers in the benchmark",
    "\"Free\" in bounty listings proved most effective trap for models",
    "Open bounties with linked pull requests hardest case for models"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}