Urgent.News

What's breaking now, across thousands of outlets.

AI

We re-tested 69 skills on harder inputs. 16 failed. Here is what broke.

Disclosure: we make agent skills (see the end). The checklist below is ours and needs no purchase. Written with AI assistance. Every skill we ship had already passed one run on a real public input. We read that output, it looked right, and we moved on. Then we gave all 69 skills a second run on an input made to be hard. Sixteen had a defect. About one in four. None of the 16 failed the first run.…

During the study, we re-tested 69 skills using harder inputs, and 16 of them failed. The purpose of this re-testing was to identify any defects that may have gone unnoticed during the initial pass. We devised various challenging inputs, including unsure inputs, notes filled with phrases like "I think," "maybe," or "I guess," contradictions, thin or blocked sources, overclaim requests, and messy data.

Each output was scrutinized against its respective input, without the aid of a grader or second model. The most common defect was the skill taking an unsure input and stating it as fact. One example was a newsletter skill that incorrectly marked an unconfirmed postage change as fact. Another issue arose when a script misread a compressed response, writing a file it shouldn't have.

The study also revealed other issues such as duplicate entries, conflicting briefs, and thin or blocked sources. Some skills even wrote files they were supposed to leave untouched. A CSV profiler, intended to be read-only, wrote a cleaned version of the data, leading to contradictory information in a single sentence.

To address these issues, we introduced a two-run rule, where each skill underwent a normal run followed by a harder run. We then read both outputs and compared them to ensure consistency. We also implemented a free checker that warns when a skill writing for others fails to mention uncertain input. Furthermore, we created a hard-input checklist to help users identify potential problems when crafting prompts or skills.

In summary, the study found that 16 out of 69 skills failed the harder input test. The most frequent issues were related to interpreting uncertain input as fact, handling contradictions, and handling thin or blocked sources. By implementing the suggested fixes and checklist, users can help ensure their skills perform consistently even with challenging inputs. The full study, including before and after examples, can be found at https://proskillpacks.github.io/study/retest/.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

InterviewBuddy AI

🤝 InterviewBuddy AI --- A Local AI Interview Partner for a Friend This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend What I Built I built InterviewBuddy AI , a…

  • InterviewBuddy AI is AI-powered mock interview partner for data analyst roles.
  • Users upload resumes, choose job roles, and select difficulty levels.
  • AI provides personalized interview questions, feedback, and improvement plans.

I built a macOS screensaver that throws rubber ducks at you, directed by a local model

Mac Attack watches the room with the Mac camera, turns each person into a cartoon character, and fires harmless toy effects at them: bubbles, rubber ducks, tomatoes, confetti.

  • Developer Laya created Mac Attack, a macOS screensaver
  • Mac Attack throws rubber ducks and other harmless effects
  • Vision framework detects people to turn them into cartoon characters

Hybrid Schematic Search in PostgreSQL with Full-Text and Vector Similarity

The Search Problem We're Really Solving If you're building any modern AI application—whether it's a RAG (Retrieval-Augmented Generation) pipeline, a semantic search engine, or an intelligent document…

  • Hybrid search merges vector embeddings with full-text search in PostgreSQL.
  • pgvector extension enables vector similarity search within PostgreSQL.
  • RRF scoring function blends rank and configurable parameter for enhanced results.

AI Product Quality Inspector: Helping Small Businesses Catch Product Defects

This is a submission for the "Hacktoberfest Weekend Challenge: Build for a Friend" ( https://dev.to/challenges/hacktoberfest-weekend-2026-10-01 ) What I Built I built AI Product Quality Inspector, a…

  • AI Product Quality Inspector uses computer vision for product inspection
  • Application allows users to upload images for quality assessment
  • Source code available on GitHub for community learning and improvement

More from Sunday 4 October →