{
  "id": 13148052,
  "title": "My jobs search API returned the same posting five times — syndication broke my URL dedup",
  "url": "https://urgent.news/2026/10/09/my-jobs-search-api-returned-the-same-posting-five-times-syndication",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-10-09T15:47:45.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/imapphelp/my-jobs-search-api-returned-the-same-posting-five-times-syndication-broke-my-url-dedup-ab0"
  },
  "original_language": "en",
  "account": "A Job Postings Search API the reporter utilizes (https://x402.freeq.one/tools/jobs.html) fetches live listings from various job boards and delivers structured data, including title, company, location, application URL, salary, and description. During the initial testing phase, the API found 240 job postings, but only 178 were unique. The same job listing appeared five times, indicating syndication issues. Companies publish job openings through their Applicant Tracking Systems (ATS) like Greenhouse, Lever, and Workday, which then distribute the postings to numerous job boards and regional aggregators. Each aggregator modifies the URL with its tracking mechanism, path format, and posting date. In an attempt to clean up the duplicated listings, the reporter first tried deduplication based on the apply URL, but this approach failed as no two copies shared a URL. Next, the reporter attempted to match listings based on a combination of company and title, which improved the results but still missed numerous duplicates. Eventually, a successful method was discovered through a three-layer normalization process. In this process, the URL undergoes several changes, including stripping query parameters, fixing the scheme and trailing slash, converting the host to lowercase. This normalization is crucial for matching purposes but is kept as a record for provenance evidence. The title is also normalized by converting it to lowercase, removing parenthetical suffixes like (m/f/d) and (remote), eliminating location prefixes, and collapsing whitespace. For the company field, legal suffixes (such as Inc, GmbH, Ltd) are omitted, and whitespace is collapsed. After normalizing the company and title fields, listings are grouped together. The earliest posting date is designated as the creation date, and all apply URLs are ordered with the canonical URL appearing first. While this method is not flawless, it effectively reduces the noise in the dataset by approximately 25%. The reporter emphasizes that whenever a dataset is syndicated, the URL serves as an attribute of the copy rather than the entity itself. Therefore, to deduplicate, it is essential to focus on the normalized content fields while treating URLs as provenance information rather than unique identifiers.",
  "summary": "I run a Job Postings Search API ( https://x402.freeq.one/tools/jobs.html ) that queries live postings across boards and returns structured results — title, company, location, apply URL, salary, description. Building it taught me my favorite kind of lesson: the data lies in boring, predictable ways. First real test: \"backend engineer, Berlin\". 240 hits. 178 unique jobs. The same posting appeared…",
  "key_points": [
    "API fetches job listings from multiple boards, returns 240 total, 178 unique",
    "Same posting appears five times due to syndication, tracking mechanisms",
    "Three-layer normalization process successfully reduces duplicates by 25%"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}