Urgent.News

What's breaking now, across thousands of outlets.

AI

How to test whether coding agents discover the skills you need

A coding skill can be clear, correct, and still fail to help if an agent never finds it. Testing a skill by naming it directly only answers whether the agent can use it after it has been loaded. It does not show whether an agent will discover it from the kind of request a teammate or customer actually makes. Expo’s October 8, 2026 evaluation write-up describes a practical way to test that gap:…

Determining whether coding agents acquire the necessary skills often proves elusive. A coding skill's clarity and correctness may not guarantee its usefulness if an agent fails to discover it. Testing a skill by directly naming it reveals whether an agent can utilize it post-loading but does not confirm whether it will identify the skill when confronted with a genuine request.

Expo's October 8, 2026 evaluation unveiled a pragmatic approach to gauge that critical gap: task-based assessments for coding agents, analysis of their execution traces, and measurement of whether they load pertinent skills without explicit guidance. The experiment's efficacy stems from its context-specific nature, involving a single plugin, model configuration, and ten task-based assignments.

Yet, it remains a methodological guide rather than a definitive performance benchmark. When maintaining coding instructions, the paramount question extends beyond "Is this skill invoked when called?" to "Will the agent correctly identify the appropriate entry point during regular operations, and what evidence substantiates its adherence to the guidance?"

Expo's evaluation framework tracked three distinct metrics: skill activation, skill utilization, and application performance. A singular skill activation does not guarantee that the guidance entered the agent's context; rather, it verifies that the guidance reached the agent's cognitive space. Absence of skill uptake implies the agent recognized the instruction but failed to implement it.

In some instances, a task can succeed even without a skill, as the agent may already possess the requisite knowledge. This underscores the inadequacy of "the final answer appearing correct" as evidence of skill discovery; it fails to ascertain whether the agent unearthed your instructions, complied with them, or executed the task via pre-existing expertise.

To discern a skill's contribution post-loading, refer to our guide on gauging plugin efficacy in Claude Code evaluations. This addresses a distinct query: whether a skill alters task outcomes when juxtaposed with a run devoid of the skill. The discovery assessment probes whether a typical task directive steers the agent towards the skill in question.

Commence with commonplace task directives, steering clear of mentioning the skill name, command, or targeted lookup within the prompt. Sample prompts may include "Add tabs and a settings screen to this app" or "Debug why this development build consistently crashes." Refrain from specifying the skill in such prompts to eschew tests of obedience to explicit instructions.

Prior to each execution, document the relevant skills and the observable evidence indicating their utilization. Expo's evaluation harness annotated scenarios with expected skills, evidence of guidance application, and product behaviors for scrutiny. This practice mitigates the risk of post-execution score adjustments. Varied task structures are crucial, encompassing new project creation, modifications to an existing project, and debugging tasks.

A modest but diverse task set can unveil whether discovery functions optimally for various prompt patterns. Consistency is key; maintain task stability across runs to compare skill catalog alterations, not fluctuating prompts. Subsequently, scrutinize the trace, noting the loaded skills, their sequence, and the agent's progression from general to specialized instructions.

Concurrently, examine edited files or application behavior for explicit recommendations embedded in the skill. A skill-trigger confirmation alone does not assure that the advice influenced the implementation. Maintain a comprehensive collection of instructions to furnish a clear starting point for agents. When a collection harbors numerous specialized instructions, agents must choose among them while concurrently addressing the task at hand.

Expo introduced 'expo-overview' as an entry point, guiding broad prompts to more specialized instructions in three illustrative scenarios. The entry point loaded first, followed by evidence of subsequent relevant skills. This design can be tested by providing a single discernible entry point, clearly stated in the title and introduction, accompanied by a succinct map linking common objectives to specialized instructions.

The entry point should aid in selection but should not overshadow the detailed skills. However, adducing an entry point does not absolve you from testing its effectiveness. Expo's findings indicated that despite the presence of the entry point, some downstream skills remained undetected. Trace analysis can pinpoint the juncture where selection halted; it may indicate insufficient router guidance or the agent relying on its innate knowledge instead.

Do not presume that adding a router resolves all discovery challenges. Expo determined that even with the entry point in place, certain relevant downstream skills still remained undetected. This underscores the necessity of targeted modifications rather than comprehensive skill revisions. Evaluate skills within the conditions under which they are typically utilized.

A skill's visibility may wane within a cluttered environment. Expo's Codex CLI catalog test exemplified how, with a shared token allowance for all skills, shortened descriptions and eventual skill disappearance occurred when the catalog exceeded that allowance. In an environment featuring 175 installed skills, Expo entries were present but only the initial 60 characters of each description were visible.

These observations, specific to Expo's tested version and configuration, underscore the importance of replicating your team's environment during evaluations. Incorporate the other plugins and instructions commonly installed by your team, then verify the agent's ability to perceive them before commencing a task.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

The AI Co-Pilot Dilemma: Does Code Generation Stifle Authentic Programming Practice?

AI coding assistants now write boilerplate, suggest entire functions, and explain error messages in seconds. For many developers, that speed feels like a gift.

  • AI coding assistants generate code rapidly, raising concerns about authentic practice.
  • Removing friction from programming may hinder learning if solutions aren't understood.
  • Beginners should write their own versions of code to develop deeper comprehension.

Liquid AI's d1 Decision Models Went Open: Triage Support Tickets on a CPU With the 600M One (and Where It Fools You)

Attributed compile + one small real CPU run Primary sources (Liquid AI, 2026-10-07): Open d1: Edge decision models for text, vision, and audio (Liquid AI blog), Multimodal open d1 decision models for…

  • Liquid AI released d1 decision models (d1-3B and d1-omni-600M) as open source
  • d1-3B model scores 48.57 on Decision Index v0.2.1, best under 10B
  • d1-omni-600M model tested on 12 support tickets, flagged high-risk cases

Gemma 4 From E2B to 31B on an AMD MI300X: fp8 Overtakes bf16 From 12B Up

This article provides a step by step guide to serving every Gemma 4 size, E2B, E4B, 12B, 26B-A4B and 31B, on one AMD Instinct MI300X through vLLM in four weight formats, with each build timed across…

  • Gemma 4 models range from E2B to 31B size
  • fp8 format overtakes bf16 from 12B and larger
  • 4-bit models are 0.14x to 0.69x faster than bf16

WIldGuard AI

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass What I Built WildGuard AI — AI-Powered Wildlife Exploration WildGuard AI is an AI-powered wildlife application…

  • WildGuard AI is an AI-powered wildlife identification app
  • Open-source AI models enable local inference and offline identification

More from Friday 9 October →