{
  "id": 2964771,
  "title": "OCR It – pull text out of un-copyable documents for your LLM",
  "url": "https://urgent.news/2026/08/24/ocr-it-pull-text-out-of-un-copyable-documents-for-your-llm",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-08-24T06:25:31.000Z",
  "source": {
    "name": "Hacker News",
    "slug": "hacker-news",
    "url": "https://github.com/thiagotigaz/ocr-it"
  },
  "original_language": "en",
  "account": "A Chrome extension simplifies extracting text from paginated documents trapped in viewers, such as scanned books, slide decks, and PDFs. Once installed, drag a capture region, assign a hotkey, and hit it on every page to screenshot the exact area, OCR it, and append the text to a running transcript. Alternatively, let the extension handle the entire process with a single hotkey: capture, turn the page, repeat until the document ends. The extension runs OCR locally using a bundled Tesseract build, requiring no API keys or network connections, ensuring data privacy and security. After each capture, the extension lists the pages with thumbnails, allowing users to identify any drift in the region. Text is editable in place, and bad reads can be re-run individually. The extension supports auto-page turning after capture and provides progress reports, ensuring reliable end-detection for unattended loops. Supported languages include English, Portuguese, Spanish, and other ~100 languages via vendor-provided models.",
  "summary": null,
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}