Urgent.News

What's breaking now, across thousands of outlets.

AI

The Model Didn’t Get Dumber. My Agent Skills Got Stale.

When Claude Opus 5 and GPT-5.6 arrived, I expected my coding agents to become noticeably better. Instead, some of my workflows felt worse. The agents seemed more eager, less predictable, and occasionally “dumber” than before. Naturally, I blamed the new models. Very scientific of me. Maybe it was a skill issue Then I watched Andrej Karpathy’s interview on the No Priors podcast. One idea stuck…

When Claude Opus 5 and GPT-5.6 were launched, I anticipated my coding agents to improve significantly. Regrettably, some of my workflows felt diminished. The agents exhibited eagerness, unpredictability, and at times, an apparent decrease in competence. I initially blamed the new models, a common reaction. However, I later discovered the real issue may lie in how I instructed the agents.

A revelation came from Andrej Karpathy's interview on the No Priors podcast. The key idea was that the problem might not be a lack of capability, but rather how we guide the agent, the memory provided, and how we structure the workflow. This perspective made me question the compatibility of my custom skills with the newer models.

To investigate, I issued this prompt to my agent: Can you audit our custom skills against the current models? Identify stale prompts, conflicting instructions, outdated assumptions, and simplify or remove unnecessary elements. Subsequent testing revealed that some instructions were written for older models, redundant, unnecessary, or overly demanding for the newer models.

After refining these instructions and retesting, the outcomes felt considerably better. This experience aligns with official advice from both Anthropic and OpenAI, who suggest recalibrating instructions during model migrations. Anthropic's Claude Opus 5 manual advises removing verification instructions inherited from previous models as they can lead to unnecessary over-verification.

OpenAI's guidance for GPT-5.6 recommends eliminating repeated instructions, streamlining tool descriptions, and re-evaluating after each modification. Internal evaluations by OpenAI showed that leaner system prompts resulted in a 10-15% improvement in scores and reduced token usage. However, it is important to note that prompt performance does not always translate smoothly between models.

A 2024 ICLR study found only a weak correlation in performance across different prompt formats. In simpler terms, what worked well with yesterday's model might perplex tomorrow's. Therefore, it is crucial to maintain the instructions accompanying our custom skills. The new model-update checklist I propose is as follows: First, run existing skills against representative tasks.

Next, look for instructions created to work around the limitations of old models, remove duplicated or conflicting rules, and delete instructions that the new model naturally handles. Make one set of instructions at a time and retest. Retain only those changes that yield measurable improvement. This is akin to dependency maintenance, but for natural-language behavior.

The bottom line is that custom skills are not immutable documentation; they are integral to the agent system, which changes alongside the underlying model. Before concluding that a new model has deteriorated, it is essential to audit the instructions surrounding it. You may have upgraded the engine while retaining outdated owner's manuals, or in developer terms, you may have acquired technical debt rather than a more advanced engine.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

How Much VRAM Do You Really Need for Local LLMs?

The VRAM equation explained: quantized model sizes, real 2026 GPU options, and the 'can I run it?' answer for any local model. Can You Run It?

  • VRAM needed for 7B parameter model at 16-bit = 14GB
  • Quantization to 4-bit reduces 7B model VRAM to 7GB
  • 70B model Q4 precision needs ~40GB VRAM, plus additional memory

Unpopular Opinion: Why I’m an AI Skeptic

With all the hype in the past several years around AI (or more specifically GenAI), I'm not afraid to say – I'm an AI skeptic.

  • AI hype is driven by irrational enthusiasm, not solid evidence
  • Current AI lacks true self-learning and emotional capabilities
  • Enterprise benefits from AI are uncertain until technology matures

More from Sunday 16 August →