Urgent.News

What's breaking now, across thousands of outlets.

AI

How do you unit test an agent skill?

Agent skills are prompts, not code, and there’s no compiler to catch a broken one. Agent skills ship on the honour system. You rewrite one, run it twice, post something convincing in Slack, and that’s the review. Is it faster? More reliable? Going to cost more? This isn’t a strategy. This is astrology for prompts. That bugged me. Not because I thought people didn’t know what they were talking…

Agent skills are not code, but rather prompts that are reviewed on a trust system. To ensure these skills are functioning correctly, tests can be built using skilleval. This tool takes a SKILL.md file, a prompt and a fixture for the agent to use and runs a real agent. The results of the test are then asserted against specific criteria such as cost, activated skills, tools used, tool calls, tool arguments, and the final message.

These tests save artefacts so that any changes to the skill can be seen when re-running the test. Running the test multiple times can also provide a pass rate, but this should not stop a skilled individual from pushing changes. The goal is to have a result.json file that proves the skill is making fewer mistakes. This is one way to test agent skills and the author is open to hearing about other methods.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Elon Musk makes a promise to employees at townhall

SpaceX aims to enhance its AI model Grok by integrating insights from its workforce. Elon Musk envisions this initiative as a means to embed human values deeply into the AI's framework. Employees are not only invited to enhance the internal AI systems but are also inspired to share their expertise.

Researchers have found that artificial intelligence (AI) models can easily decode encrypted data, raising concerns about potential personal data breaches. A team of researchers from the University of California, Berkeley, and the Massachusetts Institute of Technology (MIT) recently published a study on this issue. The researchers tested various AI models, including those using machine learning and deep learning techniques, to see if they could decode encrypted data. They found that some AI models were able to decode encrypted data with surprising accuracy, even when the data was encrypted using secure protocols. The researchers used a type of encryption called "homomorphic encryption," which allows computations to be performed on encrypted data without decrypting it first. However, they found that some AI models were able to bypass this encryption and decode the data with a high degree of accuracy. This raises concerns about the potential for personal data breaches, as encrypted data is often used to protect sensitive information such as financial data and personal identifiable information. The researchers warned that this vulnerability could be particularly problematic for organizations that rely heavily on AI and machine learning. "As AI models become more widespread, it's essential that we develop more robust security measures to protect sensitive data," said one of the researchers. The study's findings highlight the need for further research into the intersection of AI and cybersecurity. The researchers plan to continue exploring this issue and developing new methods for protecting encrypted data from AI-powered attacks. In the meantime, organizations are advised to exercise caution when using AI models to handle sensitive data. They should also consider implementing additional security measures, such as multi-factor authentication and secure data storage, to protect against potential data breaches.

More from Sunday 16 August →