Urgent.News

What's breaking now, across thousands of outlets.

AI

Evaluating the AI-Assisted Developer Experience

This article is a copy-edited transcript of the presentation given at DevFest Melbourne , October 3, 2026. Section titles have been added to help parse the written format. Speaker commentary is in italics. All images in this presentation were generated using Nano Banana Pro . Hi! I'm Katie, and this is "Evaluating the AI-Assisted Developer Experience". I really like talks where they give you the…

Hi! I'm Katie, and this is "Evaluating the AI-Assisted Developer Experience." The summary on the second slide is: "We burn tokens so you don't have to." This presentation aims to demonstrate that point. To gauge the audience's familiarity with "Agent Skills," show of hands were taken. Many raised their hands, indicating prior knowledge or usage.

However, most lowered their hands when asked if they believe these skills actually improve their development experience. This raises an important question: are these skills actually beneficial?

Agent Skills, as defined on agentskills.io, are a lightweight, open format for extending AI agent capabilities with specialized knowledge and workflows. They are essentially a folder containing a SKILL.md file, and optionally, other resources like scripts, references, assets, or other files that aid an AI agent in completing tasks.

These skills can be created by individuals to assist with common workflows or imported from existing sources. Google has even created a collection of 150+ skills, available for download from their GitHub repository.

The question at the heart of this project is whether these Google Skills actually help developers. To answer this, a team has been evaluating the usefulness of these skills, providing data to skill authors and leadership. The goal is to prove that installing a skill will result in more accurate results, use fewer tokens, and provide faster responses.

Several evaluation frameworks exist, and choosing the right one depends on the ecosystem. DeepEval and Harbor are two examples mentioned in the presentation. Ultimately, the evaluation aims to provide empirical data comparable to how software would be tested before release.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Monday 5 October →