Urgent.News

What's breaking now, across thousands of outlets.

AI

TeleOCR: A 1.2B Vision-Language Model for Structured Document Parsing

Explore TeleOCR’s document-parsing benchmarks, distorted-page handling, table and formula extraction, limitations, and comparisons with other OCR VLMs.

TeleOCR: A 1.2B Vision-Language Model for Structured Document Parsing

TeleOCR is an open-source vision-language model with approximately 1.2 billion parameters, maintained by XingChen-AGI. It is designed for parsing both digital and camera-captured documents, particularly those with geometric distortion and structured content like tables and formulas. Maintained under the Apache-2.0 license, TeleOCR leverages the Transformers library and can be loaded via AutoProcessor and AutoModel.

One of the model's key features is its ability to handle deformation-aware document modeling, adaptive sampling, and content-structure decoupled learning. This allows it to accurately parse documents with perspective or curvature distortions, making it suitable for phone photos of pages. It has demonstrated strong performance on several document benchmarks, including an overall score of 96.87 on OmniDocBench v1.6 and 88.53 on Wild_OmniDocBench. The model also secured first place in the ICDAR 2026 Sci-ImageMiner Challenge.

TeleOCR is particularly effective in parsing distorted pages without the need for dewarping preprocessing or a dedicated rectification model. It excels at extracting text, tables, and reading order from mixed-layout documents. The reported read-order edit score of 0.122 indicates its effectiveness in preserving table structure alongside text extraction. However, it should be noted that the reported benchmark results may not guarantee performance on all languages, scan qualities, or production pipelines.

Although the model is more compact than larger general vision-language models (VLMs), its specific details such as input resolution limit, training dataset size, training steps, VRAM requirements, and inference speed are not specified in the supplied material. Therefore, it is recommended to run a representative test set before replacing a specialized OCR pipeline.

While TeleOCR performs well on various benchmarks, it is important to consider its limitations. The documentation does not provide information on maximum input resolution, supported image dimensions, maximum page count, context length, batch-size guidance, or inference speed. Additionally, it does not disclose VRAM requirements or hardware benchmarks.

The 1.2B parameter count alone is insufficient to calculate the actual inference memory needed, as it can be influenced by factors such as image preprocessing, precision, framework behavior, and generation length.

In summary, TeleOCR is a specialized vision-language model for structured document parsing, particularly suited for handling distorted pages and mixed-layout documents. Its compact size makes it a viable option for document parsing tasks, but users should be aware of the lack of detailed information regarding its operational parameters and limitations before deployment.

Written by urgent.news from HackerNoon's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at hackernoon.com →

More in AI

More from Saturday 3 October →