{
  "id": 163697,
  "title": "Pervasive Backdoor Vulnerabilities in Genomic Foundation Models",
  "url": "https://urgent.news/2026/08/04/pervasive-backdoor-vulnerabilities-in-genomic-foundation-models",
  "topic": "ai",
  "section": "AI",
  "published": "2026-08-04T00:00:00.000Z",
  "source": {
    "name": "bioRxiv",
    "slug": "biorxiv",
    "url": "https://www.biorxiv.org/content/10.64898/2026.07.30.741642v1?rss=1"
  },
  "original_language": "en",
  "account": "Genomic foundation models, designed to analyze and create DNA sequences, have a hidden weakness: they can be compromised through the introduction of malicious elements into their training data. Researchers have thoroughly tested this vulnerability across various model types, sizes, and tasks. By subtly altering just 5% of the data used to train these models, attackers can achieve almost guaranteed success in manipulating model outputs. This success rate remained remarkably consistent regardless of the model's size. However, the introduction of longer triggers and a higher rate of poisoning data significantly boosted the effectiveness of these attacks. Despite these vulnerabilities, the trained models generally performed well on their unaltered tasks, with only minor changes in performance. To combat this threat, researchers have developed a two-step defense system. This system uses a method that screens for changes in single nucleotides and validates results against a reference database. When tested on a variety of configurations, this combined approach proved to be highly effective, correctly identifying poisoned models in nearly every case. These discoveries highlight the significant risk posed by data poisoning in genomic foundation models. They underscore the urgent need for stricter controls over data provenance, more rigorous testing for adversarial attacks, and comprehensive security audits after models have been deployed.",
  "summary": "Genomic foundation models are increasingly used to interpret and design DNA sequences, yet their susceptibility to training-data manipulation remains poorly understood. Here we systematically evaluate backdoor poisoning across three model families, seven parameter scales ranging from 50 million to 7 billion, and 18 genomic classification tasks. We introduce two complementary 48-nucleotide…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}