The AI safety check that runs on a laptop and nearly matched a 35B model
Guardrail selection has typically meant choosing between a purpose-built classifier and an LLM acting as a judge. Decision models such The post The AI safety check that runs on a laptop and nearly matched a 35B model appeared first on The New Stack .
The AI safety check performed on a laptop, conducted by Red Hat's AI Safety team, nearly matched a 35-billion-parameter model, Qwen3.6-35B. The team evaluated nine guardrail configurations, testing prompt injection and content safety using NVIDIA's open-source NeMo Guardrails toolkit. Among the configurations, Red Hat's DeBERTa-based prompt-injection classifier proved to be the fastest, with a median latency of 54.1 milliseconds, compared to the other models.
Qwen3.6-35B, as an LLM judge, achieved the highest accuracy of 89.31%, while Red Hat's classifier secured second place at 89.01%. The content-safety benchmark shifted the leaderboard order, with Jev leading at 86.20%, followed by DiffusionGemma and Qwen at 85.53-85.47%, respectively. Red Hat's 125-million-parameter Granite Guardian classifier was sixth at 80.27%.
Decision models like Jev offer flexibility for classification without the cost of generating unused tokens, but they face competition from smaller predictive models in content safety. The accuracy of decision models can be influenced by the tailored risk definitions used in the policy, while latency varies based on deployment and the underlying hardware.
Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.