{
  "id": 92022,
  "title": "What limits local ancestry inference at low divergence: a feasibility threshold, a metric that conceals failure, and a deficit of input more than architecture",
  "url": "https://urgent.news/2026/08/03/what-limits-local-ancestry-inference-at-low-divergence-a-feasibility",
  "topic": "ai",
  "section": "AI",
  "published": "2026-08-03T00:00:00.000Z",
  "source": {
    "name": "bioRxiv",
    "slug": "biorxiv",
    "url": "https://www.biorxiv.org/content/10.64898/2026.07.30.741148v1?rss=1"
  },
  "original_language": "en",
  "account": "Examining the limits of local ancestry inference at low divergence, researchers have identified three key factors. First, a feasibility floor exists - no method can exceed 0.575 accuracy at FST = 0.0022, and even the best-performing method at FST = 0.00042 only reaches 0.551. This means that fine-scale analysis for closely related populations, such as northern and southern Han, falls below the feasibility threshold. Second, while per-site accuracy may appear promising, it can conceal a failure in tract structure. The most accurate method per site generates 78.8 times more tracts than necessary, leading to an estimated admixture time that is 61.2 times older than it actually is. However, applying Viterbi decoding does not improve accuracy. Third, simulated data that performs well in a controlled environment does not always translate to real-world applications. The main limitation is not in the architectural design of the methods, but rather in the input data. Specifically, the presence of haplotype information in the released tools is crucial for achieving additional accuracy gains. While improvements in attention mechanisms, state-space layers, model capacity, objective functions, and self-supervised pretraining can result in at most a 0.006 accuracy improvement, they do not compensate for the missing haplotype data. The findings indicate that the statistic summarizing reference matching and the size of the reference panel are more important factors than the specific architectural design of the methods.",
  "summary": "Local ancestry inference assigns each position along an admixed chromosome to a source population, underpinning admixture mapping, ancestry-specific association testing and admixture dating. Validation is almost exclusively on continentally divergent sources (Hudson's FST {approx} 0.1) and coalescent simulations; we examine both restrictions. Across FST from 0.0022 to 0.243 we compare five…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}