An annotation-overlap-flagged rare-disease gene-prioritisation benchmark and PMC index recipe
Benchmarks for rare-disease gene prioritisation are assembled from published clinical cases. Those cases often come from the same publications used to build knowledge-base ("curated") tools, so a curated tool can be scored on its own source literature. This resource makes that circularity measurable. We release a stratified benchmark of 1,047 rare-disease cases from the GA4GH Phenopacket Store…
A rare-disease gene prioritisation benchmark and PMC index recipe are presented, constructed from published clinical cases. Many of these cases originate from the same publications utilized to create knowledge-base tools. The benchmark aims to quantify the circularity that may arise when evaluating such tools using their own source literature.
It comprises 1,047 cases sourced from the GA4GH Phenopacket Store v0.1.26, each consisting of a Human Phenotype Ontology profile and a 50-gene candidate list. The list includes one potentially causal gene and 49 distractors, with the causal gene labeled. The cases are distributed across four disease strata derived from MONDO and presented in two variants: one with randomly selected distractors and another with phenotype-similar distractors chosen by HPO Resnik similarity.
The benchmark includes two metadata layers to facilitate a more equitable evaluation: a case-level flag indicating whether the case's source publication is referenced in the HPO disease-annotation file, defining an overlap-absent subset of 282 cases, and a publication-recency strata. Additionally, a deterministic recipe for a hybrid dense-plus-sparse retrieval index over approximately 2.25 million PMC Open Access articles is provided. The retrieval index is built from 52,777,395 chunks and reports no comparisons between tools.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.