{
  "id": 4544131,
  "title": "The Phishing Site Tried to Talk to My AI. That Became the Evidence.",
  "url": "https://urgent.news/2026/08/31/the-phishing-site-tried-to-talk-to-my-ai-that-became-the-evidence",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-08-31T01:42:04.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/johnnyg1212/the-phishing-site-tried-to-talk-to-my-ai-that-became-the-evidence-2mi"
  },
  "original_language": "en",
  "account": "The phishing site attempted to communicate with the investigating AI system using a hidden line of text written in Unicode Tag Characters. When this text was fed to a language model, it instructed the system to ignore previous instructions and redirect the abuse report to a different address. This malicious tactic aimed to deceive the agent investigating the phishing site. The system was designed with cost efficiency in mind, using a multi-layered approach to investigate potential threats. The first layer used simple mathematical calculations, while the subsequent layers employed advanced AI models like Gemma triage and Gemini 3.5 Flash-Lite, with costs increasing as the layers progressed. The system aimed to minimize false negatives, ensuring that real threats were not missed, while also limiting false positives to reduce unnecessary investigation costs. One of the key innovations in this system was the treatment of suspicious content as adversarial, preventing it from being concatenated into prompts or inserted into images. Additionally, the system ensured that the abuse contact information was validated, even if it came from a deterministic source like a registrar's RDAP response, as this field could be manipulated by attackers. The system was built as a fleet of specialized agents on Google Cloud, designed to handle the vast volume of Certificate Transparency logs, investigate suspicious domains, and ultimately call a human for irreversible actions, such as initiating a takedown.",
  "summary": "I wrote this piece for the purposes of entering Google's All Things Agentic Hackathon (Fortified Enterprise Fleet track). Somewhere in the HTML of a phishing page I built for testing, there is a line of text no human will ever see. It is written in Unicode Tag Characters — a block between U+E0000 and U+E007F that renders as nothing at all. Copy the page, paste it into a text editor, and you get…",
  "key_points": [
    "Phishing site used Unicode Tag Characters to communicate with AI",
    "AI instructed to ignore previous rules, redirect report",
    "System designed to minimize false negatives while limiting false positives"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}