Urgent.News

What's breaking now, across thousands of outlets.

Tech

Why Post-Fetch Filtering on LinkedIn Guest Pages Drops Your Dataset Yield

The Illusion of Server-Side Filters in Guest Sessions Data engineers integrating the linkedin-jobs-scraper often assume that selecting an input parameter means LinkedIn's database executes that query. It does not. Because this Actor operates without a login, cookies, or an API key, it relies entirely on LinkedIn's public-facing guest search pages. Over the last few years, LinkedIn has…

LinkedIn's guest search features have progressively reduced filtering options, forcing data engineers to implement post-fetch filtering on scraped job listings. By default, the Actor cannot apply user-defined filters like experienceLevel, jobType, industry, or function during the initial search; it only applies these filters after fetching all search results.

This means that even with a high maxItems parameter, the final yielded dataset may be significantly smaller than anticipated if many jobs fail post-fetch filtering. To identify if your pipeline is being starved of job listings due to overly restrictive filters, compare the count of clean items in your default dataset against your expected minimum yield after the run completes.

If the cleanItemCount falls below your business requirements, trigger an alert through the Apify API to notify your orchestrator. EnrichmentDepth settings dramatically affect performance and output structure. The 'standard' mode provides lightweight scraping with retained fields such as applyUrl, ATS platform names, and salary ranges, while 'full' mode offers deeper data but drops the applyUrl field and introduces additional network requests for company enrichment.

Changing from 'standard' to 'full' enrichment can break workflows relying on candidate application tracking, as the static apply link is no longer guaranteed. When configuring your scraper, ensure the enrichmentDepth aligns with your data needs – standard for application routing, full for comprehensive job data.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

Why is PageSpeed Insights worse than my local Lighthouse score?

You run Lighthouse in Chrome DevTools on a client URL and see Performance 94. You paste the same URL into PageSpeed Insights and the lab Performance score drops to the sixties.

  • Lighthouse and PageSpeed Insights run on different machines
  • PageSpeed Insights lab data often worse than local Lighthouse
  • Differences arise from throttling, cache, and extension states

Does Dependency Security Really Need the Cloud?

The question I kept asking myself a simple question: Why does checking a dependency for a known vulnerability need to send anything to the cloud?

  • Dependency Vulnerability Companion operates locally without internet
  • Parses project file, matches against local vulnerability database
  • Supports multiple dependency declaration formats for vulnerability checks

More from Friday 2 October →