{
  "id": 10793699,
  "title": "Quoting Anthropic Frontier Red Team",
  "url": "https://urgent.news/2026/09/29/quoting-anthropic-frontier-red-team",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-29T22:20:28.000Z",
  "source": {
    "name": "Simon Willison",
    "slug": "simon-willison",
    "url": "https://simonwillison.net/2026/Sep/29/anthropic-frontier-red-team/"
  },
  "original_language": "en",
  "account": null,
  "summary": "We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}