Urgent.News

What's breaking now, across thousands of outlets.

AI

Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks

Security benchmarks for LLM-based agents often report the attack success rate (ASR) as a measure of model robustness and use these scores to compare different models and defense mechanisms, assuming that they describe the security of the agent. In this paper, we explore whether it also influences the benchmark's measurement. To measure the effect of the benchmark representation, we introduce…

We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.

Read the original at arxiv.org →

More in AI

Running DeepAgents in a Docker Sandbox, with no cloud keys

Agent frameworks are easy to pip install and surprisingly hard to run responsibly. The moment you give an agent a filesystem, a shell, and a network, you have handed arbitrary generated code the same…

  • DeepAgents framework allows creating agents with LangGraph library
  • Docker Sandbox Kit enables running DeepAgents without cloud credentials
  • Kit consists of four files and grants phase-scoped network policy

Gemma 4 QAT on One TPU v5e: What Runs and What Doesn't

This article provides a step by step guide to repacking Google's quantization-aware-trained (QAT) Gemma 4 weights for vLLM and serving them on one Google Cloud TPU v5e chip, with every build scored…

  • Gemma 4 models run efficiently on a single TPU v5e chip
  • 4-bit repack stores QAT grid values using int4 weights and activations
  • 8-bit repack shows slight speed advantage over 4-bit repack

More from Friday 2 October →