{
  "id": 7433561,
  "title": "PHACTn enables training-free, context-independent inference of nucleotide variant tolerance across the genome",
  "url": "https://urgent.news/2026/09/14/phactn-enables-training-free-context-independent-inference-of",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-14T00:00:00.000Z",
  "source": {
    "name": "bioRxiv",
    "slug": "biorxiv",
    "url": "https://www.biorxiv.org/content/10.64898/2026.09.08.750126v1?rss=1"
  },
  "original_language": "en",
  "account": "Genome-wide prediction of single-nucleotide variant (SNV) tolerability poses a significant challenge in computational genomics, especially in non-coding regions where regulatory elements are complex and poorly understood. Traditional machine learning classifiers and genomic language models face limitations due to data circularity, demographic bias, resource-intensive computational demands, and limited biological interpretability. PHACTn (Phylogeny-Aware Computing of Tolerance for nucleotide variants) is introduced as a novel training-free, parameter-minimal method that infers nucleotide variant tolerability by leveraging the mammalian phylogenetic tree and explicitly modeling the independent evolutionary nature of substitutions relative to the query species. PHACTn requires only four interpretable parameters and does not necessitate training or GPU usage. Benchmarks against curated non-coding variants from ClinVar and disease-related non-coding variants from OMIM demonstrate superior performance compared to existing tools. Furthermore, PHACTn achieves state-of-the-art accuracy on variants within the range of alignment-based inference methods. These findings indicate that probabilistic phylogenetic modeling effectively captures evolutionary constraints that are inadequately addressed by large-scale sequence models. The method provides a robust, accessible, and biologically transparent alternative for genome-wide variant effect prediction.",
  "summary": "Accurate prediction of single-nucleotide variant (SNV) tolerability across the entire human genome remains a fundamental challenge in computational genomics, particularly for non-coding regions where the regulatory landscape is vast and poorly understood. Machine learning classifiers suffer from data circularity and demographic bias, while genomic language models demand massive computational…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}