Scrape any website into JSON by just listing the fields you want
Writing a scraper usually means inspecting the page, finding CSS selectors, handling edge cases, and then fixing it all again when the site changes its layout. For a lot of jobs that's overkill: you just want these five facts from these pages . So I built an extractor that works the other way around. You describe the data, an LLM reads the page, and you get JSON back. The whole input {…
Extracting data from websites into JSON format has traditionally involved painstaking inspection of pages, identifying CSS selectors, managing edge cases, and then reworking the scraper when the site's layout changed. However, a new approach simplifies the process significantly. An extractor works in the opposite direction: you specify the desired data, an LLM analyzes the page, and you receive JSON output.
To use this tool, you provide a list of starting URLs and define the fields you want in the JSON output. For example, for the GitHub repository of the Crawlee web scraping library, you would specify the fields: name (string), description (string), license (string), and primary_language (string). The extractor returns a JSON object containing the URL, a success indicator, and the extracted data. In this case, the description field includes a detailed explanation of what Crawlee is and does.
The model is trained to only use information available on the page and return null if something is not present. Users can choose the data type for each field, such as string, number, integer, boolean, or even a full JSON Schema for nested results. Plain-language instructions can guide the model, and pricing can be set in USD, with free results for blocked pages or those requiring JavaScript.
This method is particularly effective for pages with varying content, changing layouts, or when only a few specific fields are needed from numerous sites. It is especially useful for product pages (name, price, availability), job postings (title, company, location, salary, remote), company websites (what they do, industry, headquarters, contact page), articles (headline, author, date, summary), and event pages (date, venue, price).
However, for scraping large numbers of pages from a site with a fixed layout, a traditional selector-based scraper may be more cost-effective.
To get started, visit the AI Web Data Extractor on Apify. New accounts receive free monthly credit, allowing you to try it out on your own URLs without cost. The tool can also be integrated into code or used by AI agents through the Apify MCP server. Feedback is welcome, particularly for examples where the tool produces incorrect results.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.