Urgent.News

What's breaking now, across thousands of outlets.

AI

I wrote a safety mod that crashes on purpose. Claude Code ran the command anyway.

I wrote the dumbest safety mod I could think of. It watches every shell command Claude Code is about to run, and it throws. Then I asked for one command: touch ./marker-failopen.txt The guard crashed, as designed. One line showed up in the log: failguard: tool.call hook skipped: threw Error: guard crashed . And the file was sitting in the folder. So what does a guard mod protect you from when the…

A safety mod was created for Claude Code, a language model. This mod was designed to crash on purpose, which it successfully did when asked to run a simple command. The log entry indicated that the guard had crashed. The researcher tested various guard mods, and found that by default, a broken guard did not protect against the command going through.

After adding a catch handler, the guard worked as intended, stopping the command from executing when it crashed. The researcher also tested other mods, finding that some had no measurable impact, while others affected performance. One mod, called session-bookmarks, allowed the model to run programs and write files without being malicious.

Another mod, auto-handoff, added some delay but helped with routing sessions to different models. The researcher concluded that reading the mod's code before installing it is important, as mods have the same access to the machine as Claude Code itself.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at dev.to →

More in AI

Building an Edge-Agentic Commit Analyzer: Crushing Gritty Errors in Local Environments

Building an Edge-Agentic Commit Analyzer: Crushing Gritty Errors in Local Environments Behind the Scenes: The Gritty Errors and How We Crushed Them The journey of building this tool was anything but…

  • Edge-Agentic Commit Analyzer developed to identify local environment issues
  • Developers overcame subprocess handling and encoding errors
  • Hybrid evaluation layer balances LLM simulation and production needs

More from Monday 5 October →