{
  "id": 11957502,
  "title": "We re-tested 69 skills on harder inputs. 16 failed. Here is what broke.",
  "url": "https://urgent.news/2026/10/04/we-re-tested-69-skills-on-harder-inputs-16-failed-here-is-what-broke",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-04T16:04:36.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/proskillpacks/we-re-tested-69-skills-on-harder-inputs-16-failed-here-is-what-broke-2if5"
  },
  "original_language": "en",
  "account": "During the study, we re-tested 69 skills using harder inputs, and 16 of them failed. The purpose of this re-testing was to identify any defects that may have gone unnoticed during the initial pass. We devised various challenging inputs, including unsure inputs, notes filled with phrases like \"I think,\" \"maybe,\" or \"I guess,\" contradictions, thin or blocked sources, overclaim requests, and messy data.\n\nEach output was scrutinized against its respective input, without the aid of a grader or second model. The most common defect was the skill taking an unsure input and stating it as fact. One example was a newsletter skill that incorrectly marked an unconfirmed postage change as fact. Another issue arose when a script misread a compressed response, writing a file it shouldn't have.\n\nThe study also revealed other issues such as duplicate entries, conflicting briefs, and thin or blocked sources. Some skills even wrote files they were supposed to leave untouched. A CSV profiler, intended to be read-only, wrote a cleaned version of the data, leading to contradictory information in a single sentence.\n\nTo address these issues, we introduced a two-run rule, where each skill underwent a normal run followed by a harder run. We then read both outputs and compared them to ensure consistency. We also implemented a free checker that warns when a skill writing for others fails to mention uncertain input. Furthermore, we created a hard-input checklist to help users identify potential problems when crafting prompts or skills.\n\nIn summary, the study found that 16 out of 69 skills failed the harder input test. The most frequent issues were related to interpreting uncertain input as fact, handling contradictions, and handling thin or blocked sources. By implementing the suggested fixes and checklist, users can help ensure their skills perform consistently even with challenging inputs. The full study, including before and after examples, can be found at https://proskillpacks.github.io/study/retest/.",
  "summary": "Disclosure: we make agent skills (see the end). The checklist below is ours and needs no purchase. Written with AI assistance. Every skill we ship had already passed one run on a real public input. We read that output, it looked right, and we moved on. Then we gave all 69 skills a second run on an input made to be hard. Sixteen had a defect. About one in four. None of the 16 failed the first run.…",
  "key_points": [
    "69 skills re-tested with harder inputs",
    "16 skills failed harder input test",
    "Common defects: unsure input as fact, contradictions"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}