Urgent.News

What's breaking now, across thousands of outlets.

AI

Pervasive Backdoor Vulnerabilities in Genomic Foundation Models

Genomic foundation models are increasingly used to interpret and design DNA sequences, yet their susceptibility to training-data manipulation remains poorly understood. Here we systematically evaluate backdoor poisoning across three model families, seven parameter scales ranging from 50 million to 7 billion, and 18 genomic classification tasks. We introduce two complementary 48-nucleotide…

Genomic foundation models, designed to analyze and create DNA sequences, have a hidden weakness: they can be compromised through the introduction of malicious elements into their training data. Researchers have thoroughly tested this vulnerability across various model types, sizes, and tasks. By subtly altering just 5% of the data used to train these models, attackers can achieve almost guaranteed success in manipulating model outputs.

This success rate remained remarkably consistent regardless of the model's size. However, the introduction of longer triggers and a higher rate of poisoning data significantly boosted the effectiveness of these attacks. Despite these vulnerabilities, the trained models generally performed well on their unaltered tasks, with only minor changes in performance.

To combat this threat, researchers have developed a two-step defense system. This system uses a method that screens for changes in single nucleotides and validates results against a reference database. When tested on a variety of configurations, this combined approach proved to be highly effective, correctly identifying poisoned models in nearly every case.

These discoveries highlight the significant risk posed by data poisoning in genomic foundation models. They underscore the urgent need for stricter controls over data provenance, more rigorous testing for adversarial attacks, and comprehensive security audits after models have been deployed.

Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at biorxiv.org →

More in AI

More from Tuesday 4 August →