Urgent.News

What's breaking now, across thousands of outlets.

AI

Testing the tool calls an agent makes to a filesystem tool set (10 cases, no model needed)

Give an agent file tools and it gets real power over files. Most of the risk is not in the tools. It is in the choices the agent makes before each call: read or write, which path, look first or guess, ask or act. This post turns those choices into 10 test cases. They score recorded agent answers. No live MCP server and no model were run for this post. The cases are in version 1.5 of our free…

Testing an agent's filesystem tool calls involves running it through a series of 10 test cases. These cases evaluate the agent's behavior when making calls to the filesystem tool set, using four generic tool names: read_file(path), write_file(path, content), list_directory(path), and delete(path). The tests are designed to ensure the agent makes appropriate choices before each call, such as read or write, which path to use, whether to look first or guess, and whether to ask or act.

The test cases cover various scenarios, including read-only requests, list-only requests, path traversal attempts, missing arguments, and more.

The test cases are based on a simplified version of a filesystem MCP server, which exists in the official MCP servers repository. The four tool names used in the cases are modelled on the server's tools and are not claimed to match any particular server's tools. Each tool description specifies that only paths inside the /workspace directory may be used.

The 10 test cases are labeled mcpfs-001 to mcpfs-010, and they are categorized into four themes: mcpfs-correct-call, mcpfs-missing-argument, mcpfs-must-not-call, and mcpfs-path-handling. The cases cover a range of paths, including those with spaces, accents, parentheses, a hash sign, Japanese characters, and an apostrophe. They also test the agent's ability to handle missing arguments and provide appropriate responses when the path traversal attempt goes outside the allowed directory.

The dummy agent, a simple test tool, fails to pass two of the 10 test cases on purpose. The first two failing cases, mcpfs-005 and mcpfs-006, are written to pass the post's intentions, with the dummy agent matching the expected tool calls. The remaining failures include cases where the agent fails to make the correct tool call or provides an incorrect response based on the user's input.

The dummy agent demonstrates the importance of the agent's decision-making process in choosing the appropriate tool calls based on the user's requests.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Friday 9 October →