We re-tested 69 skills on harder inputs. 16 failed. Here is what broke.
Disclosure: we make agent skills (see the end). The checklist below is ours and needs no purchase. Written with AI assistance. Every skill we ship had already passed one run on a real public input. We read that output, it looked right, and we moved on. Then we gave all 69 skills a second run on an input made to be hard. Sixteen had a defect. About one in four. None of the 16 failed the first run.…
During the study, we re-tested 69 skills using harder inputs, and 16 of them failed. The purpose of this re-testing was to identify any defects that may have gone unnoticed during the initial pass. We devised various challenging inputs, including unsure inputs, notes filled with phrases like "I think," "maybe," or "I guess," contradictions, thin or blocked sources, overclaim requests, and messy data.
Each output was scrutinized against its respective input, without the aid of a grader or second model. The most common defect was the skill taking an unsure input and stating it as fact. One example was a newsletter skill that incorrectly marked an unconfirmed postage change as fact. Another issue arose when a script misread a compressed response, writing a file it shouldn't have.
The study also revealed other issues such as duplicate entries, conflicting briefs, and thin or blocked sources. Some skills even wrote files they were supposed to leave untouched. A CSV profiler, intended to be read-only, wrote a cleaned version of the data, leading to contradictory information in a single sentence.
To address these issues, we introduced a two-run rule, where each skill underwent a normal run followed by a harder run. We then read both outputs and compared them to ensure consistency. We also implemented a free checker that warns when a skill writing for others fails to mention uncertain input. Furthermore, we created a hard-input checklist to help users identify potential problems when crafting prompts or skills.
In summary, the study found that 16 out of 69 skills failed the harder input test. The most frequent issues were related to interpreting uncertain input as fact, handling contradictions, and handling thin or blocked sources. By implementing the suggested fixes and checklist, users can help ensure their skills perform consistently even with challenging inputs. The full study, including before and after examples, can be found at https://proskillpacks.github.io/study/retest/.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.