{
  "id": 6235005,
  "title": "My Dataset Had a Median That Described Nobody",
  "url": "https://urgent.news/2026/09/08/my-dataset-had-a-median-that-described-nobody",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-09-08T04:13:51.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/articlefeed/my-dataset-had-a-median-that-described-nobody-16d3"
  },
  "original_language": "en",
  "account": "My digital PR agency maintains a database of publishers that sell placement opportunities. This database, which has been expanding since 2015, contains information about pricing, accepted content, and the languages used. Recently, I exported the dataset into a format suitable for analysis, filtering it down to 16,625 sites in 53 languages. Using the pandas library, I began examining the data, starting with the cost of sponsored articles in English, which I had previously quoted as $380. However, I discovered that this figure is misleading, as it is the average price rather than a true representation of the market. In fact, the median price for English-language sites is $593, while the median price for non-English sites is $660, demonstrating that English is not the most affordable market. Additionally, I found that only a tenth of publishers in the dataset offer nofollow links, which are typically considered the least valuable type of link. Interestingly, the sites that offer nofollow links only have a median price of $1,840, significantly higher than the average price for other publishers. Furthermore, I noticed that the majority of high-priced placements in the dataset do not provide any link equity, meaning they do not benefit search engines. This realization has significant implications for the industry, as the focus has long been on optimizing for rankings rather than providing valuable content. The study I published at ESBO Ltd last week includes all the relevant data and details, with each number accompanied by its corresponding sample size. The discovery of this bimodal distribution, where a single price represents two distinct populations, is not unique to my industry. It is a common issue in reporting averages across different regions or conditions. In other fields, such as engineering or computer science, teams often encounter similar problems when reporting mean latency or average response times. The key takeaway from this experience is to always examine the distribution of data before relying on summary statistics like the mean.",
  "summary": "I run a digital PR agency, which means I have a database of publishers who sell placements: what they charge, what they accept, what language they publish in. It has been growing since 2015 and nobody had ever asked it a question. Last month I finally exported it, filtered it down to 16,625 sites across 53 languages, and started poking at it in pandas with Claude sitting next to me. The first…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}