The "Lying Prose" Problem: Why We Built a Custom CI Linter to Stop Documentation from Contradicting Its Own Data
Have you ever visited a SaaS product page where the hero headline proclaims: "Get started with 10 team seats for just $29/month." ...and then you scroll down three viewports to the interactive pricing calculator, only to see: "Starter Tier: 5 seats included — $39/month." Which one is true? You have no idea. You instantly lose confidence in the product, wonder if the checkout page will surprise…
The article discusses the issue of documentation contradicting its own data, known as Prose Data Drift. This occurs when there is a discrepancy between the structured business datasets and the human-written explanations provided to readers. The author explains how this problem manifested at fortunesweave.online, a data-heavy SaaS product for Fire Emblem: Fortune's Weave.
The author describes how they built the site using TypeScript schemas and structured JSON tables, ensuring a single source of truth for all data. They also developed interactive tools that dynamically ingested these datasets without any hardcoded values. However, despite these efforts, the documentation still contained inaccuracies.
The main cause of Prose Data Drift, according to the author, is the divergence between structured business datasets and the human prose written to explain them. They refer to this phenomenon as "Prose Data Drift" and explain that it happens when an engineer or technical writer creates a data table, and the UI updates instantly if the data changes. However, the accompanying editorial context may still contain outdated or contradictory information.
The author highlights that existing linters and tests do not catch Prose Data Drift. TypeScript only checks for type correctness, not semantic truth, and E2E tests verify DOM rendering, not editorial consistency. They also note that natural language writing tends to hardcode numbers, making it difficult to avoid inconsistencies.
To address this issue, the author developed a custom static analysis script called audit-data-drift.mjs. This script runs during the CI prebuild step and analyzes the data files to extract numerical values and their corresponding entities. By treating prose as untrusted state, the script helps identify and prevent Prose Data Drift, ensuring that the documentation accurately reflects the underlying data.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.