Urgent.News

What's breaking now, across thousands of outlets.

AI

ABBYY gives old-school OCR a job in the AI pipeline

FineParser runs in a CPU-powered container, preserving tables and layout before handing documents to an LLM

ABBYY gives old-school OCR a job in the AI pipeline

ABBYY has repackaged its FineReader OCR engine as a self-hosted tool called FineParser, designed to convert documents into structured text for use by AI systems. The tool operates within a Docker container on a CPU, without necessitating a GPU. FineParser's primary objective is to maintain the layout of a document while extracting its content, transforming images of documents in various languages into structured, formatted text suitable for processing with modern generative AI language models (LLMs).

While preserving document structure is its main purpose, FineParser can also assist in digitizing print text for modern content management systems (CMS), archival purposes, or as part of a production pipeline.

ABBYY refers to FineParser's approach as "deterministic AI," emphasizing that the tool extracts text and document structure rather than generating plausible renditions of them. The output of FineParser can be fed into a generative AI system, resulting in less predictable responses. The company also offers a programmable machine-learning framework called NeoML, which is free and open-source (FOSS) and available on GitHub.

Additionally, ABBYY publishes an OCR Software Development Kit (SDK) for companies seeking to integrate the FineReader engine into their own products.

FineParser itself is not open source, but ABBYY maintains a GitHub repository containing examples and community support. The self-hosted tool comes with a free tier allowing 1,000 pages per month for one year. Subscriptions to higher tiers connect to a license server for validation; for fully offline deployment, an Enterprise plan is required.

Preserving the document structure means recognizing columns in reading order, headings, and tables, including those without borders, instead of generating a jumble of extracted text. FineParser also handles handwriting and more than 200 languages. Developers can submit documents via FineParser's REST API, receiving the output as plain text, JSON, or DocLang, a compact format intended for LLM input.

ABBYY asserts that its decades-old OCR approach—running on a CPU, preserving the layout, and leaving text generation to other tools—still has a place in the AI pipeline. The company believes that the old-fashioned aspect can still be useful.

Written by urgent.news from The Register's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at theregister.com →

More in AI

More from Tuesday 22 September →