AI to decipher even skewed scans: Sarvam launches Vision 2.1—how it extracts data from messy documents in seconds..
Deciphering complex language fonts, archival documents, and handwritten notes has long been a significant challenge for AI. Addressing this issue, Sarvam AI has launched its upgraded vision-language model, Sarvam Vision 2.1. This model goes beyond merely reading documents; it transforms messy scans and legacy files into clean, structured data that businesses can directly integrate into their production workflows.
The company states that Vision 2.1 was developed to overcome the limitations of the previous version. Its focus extends beyond simple document reading to converting unstructured scans into clean, structured data, enabling businesses to utilize this information directly in their production workflows.
What can Sarvam Vision 2.1 do?
Vision 2.1 is specifically designed to handle documents containing a mix of standard text, tables, forms, and handwritten information.
Understanding complex tables
The model excels at processing nested, complex, and multi-page table structures. It has enhanced capabilities for interpreting and organizing information from tables that span multiple pages or deviate from standard formatting.
Extracting key information from forms
Vision 2.1 can extract key-value information—such as names, dates, and financial figures—from structured forms. This allows essential data within documents to be organized for direct use.
Recognizing handwriting in Indian languages
Additionally, the model possesses the unique ability to recognize handwritten notes and annotations in Indian languages. The company has paid special attention to regional variations in handwriting styles. Improved Accuracy for Irregular Scans
The model has undergone specialized training to minimize hallucinations and spatial inconsistencies often found in scans of dense or irregularly formatted documents. The objective is to achieve a more precise understanding of the information and layout within the documents.
New OCR Benchmark Covers 22 Indian Languages
Alongside Vision 2.1, Sarvam AI has introduced the new Sarvam India OCR benchmark. According to the company, this benchmark encompasses all 22 official Indian languages.
The benchmark comprises a total of 6,909 samples, consisting of 6,609 samples in regional languages and 300 English baseline samples. The source material includes newspapers, brochures, books, and historical records dating back to 1800.
How Did Vision 2.1 Perform on the Benchmark?
According to the company, Sarvam Vision 2.1 achieved an overall word accuracy of 87.39% on the new Sarvam Indic OCR benchmark. Meanwhile, the model scored 87.30% on the global olmOCR-bench. These benchmarks are used to evaluate the model's capabilities regarding document processing and OCR.
How Was the Model Trained?
Vision 2.1 was trained on a combination of synthetic and real-world document scans. Special emphasis was placed on diverse forms of regional handwriting and scenarios involving form extraction.
Subsequently, the model was fine-tuned using supervised learning. Additionally, Reinforcement Learning with Variable Rewards (RLVR) was employed to enhance the model's structural and textual accuracy.
How Will It Benefit Businesses?
The objective of Vision 2.1 is to make the process of extracting information from documents more automated and structured. It is particularly suited for business workflows that require data extraction from large volumes of scanned documents, forms, or tables.
According to the company, the focus is on converting unstructured scans into structured data that businesses can directly utilize in their production workflows.
Model available via Sarvam's APIs
Sarvam Vision 2.1 is made available through Sarvam's Document Intelligence APIs. These APIs include tools for converting multi-page files into structured text and automatically extracting data from forms and tables.
Disclaimer: This content has been sourced and edited from Amar Ujala. While we modified it for clarity and presentation, the original content belongs to its respective authors and website. We do not claim ownership of the content.

