Sarvam AI updates vision model, doubles down on Indic-language push
Sarvam launched the original vision model in February as part of its efforts to build homegrown, "sovereign" AI systems tailored to India. The model was designed to read scanned documents and images and convert them into usable digital text, a task known as optical character recognition (OCR).
Artificial intelligence firm Sarvam AI unveiled an enhanced version of its vision AI model, Sarvam Vision 2.1, in an effort to enhance the performance of AI tools for Indian languages. The company aims to make its AI tools more efficient and affordable for businesses digitizing paperwork, such as forms, tables, and handwritten records, that often use India's diverse regional scripts.
Launched in February as part of Sarvam's focus on developing locally relevant AI systems, the original vision model performed well overall but faced challenges with complex forms, multi-page tables, occasional text hallucination, and high operational costs. The updated 2.1 version addresses these issues by incorporating real-world and artificial data, including handwritten forms in various Indian languages, to improve accuracy and reduce costs for businesses using it in production.
Sarvam Vision 2.1 leads or ranks near the top in several industry benchmarks, surpassing global rivals like Gemini 3.6 Flash and Claude Opus 5 in document-reading accuracy. This improvement is particularly notable in the Indian-language segment, where most international AI models still struggle. Sarvam AI also published Indic DiarBench, a benchmark designed to evaluate AI models' performance on complex multi-speaker conversations across India's 22 official languages, in collaboration with IIT Madras's AI4Bharat initiative.
Written by urgent.news from Economic Times Tech's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.