"Is There Any Point to This?" Removing the Jev Plugins I Added to Claude Code After One Week
At the end of my previous article , I wrote what I would do next. I would run the plugin that suggests skills with Jev in a setting where it adds nothing and only writes a log, and match the skills Jev picked against the skills Claude Code actually used. I ran it for a week starting September 21, 2026. Of the 539 times Jev picked a skill, that skill was used in the same session within 30 minutes…
At the end of my previous article, I announced the next steps I would take. I intended to use the Jev plugin in a setting where it provides no additional value and simply logs information, then compare the skills Jev suggested against those actually employed by Claude Code. Running this test for a week starting from September 21st, 2026, I found that the Jev plugin made 1,242 skill selections, with Jev picking a skill in 539 instances.
About 5% of those 539 instances resulted in the corresponding skill being invoked within 30 minutes of the suggestion, totaling 28 occurrences. Following these findings, I promptly removed the plugin without further investigation. During this week, Claude Code also utilized a different plugin that examines whether Claude Code declares a task complete without properly verifying it.
Over the same period, this second plugin caused Claude Code to halt 8 times. Upon close examination of two instances where Claude Code was stopped, it emerged that Claude Code had completed executing my environment's verification script just before the interruption occurred. Ultimately, my conclusion is as follows: constructing a judgment model like Jev and integrating it as a component into a pipeline with a fixed flow is the only feasible approach to assess its impact.
When retrofitted into Claude Code, a system that independently determines its next steps, even minor differences in component integration can lead to significant behavioral changes. Since Jev's skill predictions only overlapped slightly with Claude Code's own choices, the performance of the plugins could not be easily determined using surface-level statistics alone.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.