{
  "id": 11482696,
  "title": "How many samples does a false-positive budget need?",
  "url": "https://urgent.news/2026/10/02/how-many-samples-does-a-false-positive-budget-need",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-10-02T16:58:19.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/pm25coder/how-many-samples-does-a-false-positive-budget-need-1ni"
  },
  "original_language": "en",
  "account": "When setting up a false-positive budget, two key factors must be considered: the number of benign samples required and the number of attack samples needed to compare different rules. For a given threshold, the false-positive rate is modeled as a Beta distribution with parameters k+1 and n-k, where n is the total number of benign samples. Solving for the required number of benign samples (n) to meet a 95% confidence interval of 2% false-positive rate involves finding the smallest n such that P(Bin(n, 0.02) = k+1) = 0.95. For example, if k = 0, meaning the threshold is the highest benign score, solving the equation yields n = 148.3, requiring approximately 149 benign samples.\n\nThe budget achieved by setting k = 0 and n = 149 provides a median realized false-positive rate of 0.46%, with a 5th–95th percentile range of 0.03%–1.99%. However, this threshold may not meet the promise at all times, as setting the threshold at the empirical quantile (k = floor(0.02n)) guarantees the promise but never achieves it. In practice, a rule claiming to have met the promise through quantile thresholding is unreliable, as additional data will not fix the issue.\n\nComparing two detection rules involves analyzing the number of attacks required to detect a significant difference between them. The difference in recalls between a strict and a loose threshold follows a paired question model, where only the attacks where the two rules disagree carry information. With a discordance rate (π_d) and standard error (SE) of the difference, the paired interval provides a more accurate estimate than the two marginal intervals. To achieve 80% power for detecting a 2-point difference, roughly 1,000–2,000 paired attacks are needed. However, the tests often report no significant difference even when differences exist, leading to inaccurate conclusions.\n\nIn addition to attack-sample variance, calibration-draw variance must be considered when evaluating the false-positive budget. Each rule produces a unique threshold per calibration draw, and resampling the benign calibration set allows for measuring each rule's recall stability. Analyzing nine detectors on a public prompt-injection benchmark, the calibration-draw variance often surpasses the attack-sample variance, necessitating the sum of both variances for an accurate budget calculation. The ordinary bootstrap method is not consistent for a sample maximum, so a with-replacement resample is recommended to provide a more accurate estimate of calibration-draw variance. Finally, the resolution of a detection number depends on three factors: an adequate number of benign traces, enough attack samples to detect the intended difference, and sufficient calibration draws to determine whether the outcome is due to the rule or random chance.",
  "summary": "Pick a detector, run it over some benign traffic, and choose a threshold. Usually the threshold is an order statistic of the benign scores — the k-th largest, or a quantile — and you report something like false-positive rate ≤ 2% at 95% confidence . Two questions tend to get skipped in that sentence: how many benign samples does the promise actually need, and how many attacks does it take to…",
  "key_points": [
    "Approximately 149 benign samples needed for 95% confidence interval of 2% false-positive rate",
    "1,000–2,000 paired attacks required to detect 2-point difference with 80% power"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}