What you can (and can't) scrape from LinkedIn without logging in — a head-to-head test of 45 runs
Most "no cookies" LinkedIn scrapers on the market return data that a logged-out visitor never sees: exact connection counts, every past job, people search with "15,666 results". That data has to come from somewhere — usually logged-in accounts on the seller's side. Sometimes that's fine for you. Sometimes it's a compliance problem you only discover later. I wanted a clear answer to a narrower…
LinkedIn scrapers that do not require a login typically return data that a visitor without logging in would not see. This includes exact connection counts, every past job, and extensive search results. However, the availability of this data depends on the specific scraper being used. A head-to-head test of 45 runs using the most popular LinkedIn scrapers and the reporter's own scraper found that logged-out visitors can generally access job listings, company pages, and posts without any issues.
However, when it comes to profiles, some depth of information is only available to logged-in users. The test used consistent inputs across all scrapers, including a Python developer job search in New York, public profiles, company pages, profile posts, and ads. The results showed that logged-out data is typically sufficient for most use cases, making it the safer option to build on.
However, for profiles, users may lose some depth of information without a login, and it is important to consider whether this level of detail is necessary before using a logged-out scraper. Additionally, the test revealed three bugs that were only found by comparing side-by-side results from different scrapers. These bugs included incorrect remote work filters, unverified email addresses being billed, and missing posts from company pages.
Overall, the findings suggest that for most jobs, companies, and posts, logged-out data is reliable and sufficient, but profiles may require additional login access for a more complete picture.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.