The best OCR models in 2026: open-source and vision-language models compared on published benchmarks
For AI and engineering teams choosing an OCR model: the published scores as of September 2026, what each license allows, and what the scores miss.

Key takeaways
- On OmniDocBench v1.6, small open-source models built for documents hold the top three scores. TeleOCR scores 96.91, OvisOCR2 96.47 and PaddleOCR-VL-1.6 96.34, and each has about 1.2 billion parameters or fewer.
- On olmOCR-Bench, Datalab reports 85.8 for its Chandra OCR 2, and Ai2 reports 82.4 for its olmOCR 2. The two benchmarks use different pages and different scoring, so you can't compare a score on one with a score on the other.
- Read the license as well as the score. Chandra OCR 2's weights are the model files you download. They're free only for research, personal use and companies under $2 million in both revenue and funding. They also can't be used to compete with Datalab's API. HunyuanOCR's license grants no rights in the EU, the UK or South Korea, so you can't use it there.
- Scores fall on photographed pages. On Wild-OmniDocBench, which uses photos of OmniDocBench v1.5 pages, the first PaddleOCR-VL scored 91.93 on the clean pages and 72.19 on the photos.
- A model gives you the page as text. You still build field extraction and review yourself, or use a platform such as Docsumo that provides them.
On this page
On OmniDocBench v1.6, the best OCR model in September 2026 is a small open-source model built for documents. Large general-purpose AI models score lower there. TeleOCR scores 96.91, OvisOCR2 96.47 and PaddleOCR-VL-1.6 96.34. Each is released under Apache 2.0 and has about 1.2 billion parameters or fewer. Parameters measure a model's size, and that's small for an AI model. On olmOCR-Bench, Datalab reports 85.8 for its Chandra OCR 2. That's the highest reported score we found for a model you can download. Both benchmarks use mostly research papers, books and reports, so test any model on your own documents before you pick one.
The scores below come from each benchmark's own page or from the model card, the page where each maker describes its model. We checked them on September 29, 2026, and we didn't run the models ourselves.
The best OCR models in 2026, by published score#
The table lists 14 models in four groups. Most are vision-language models (VLMs), which read a page image and write out its text and layout. Size is the parameter count each maker states, except for Chandra OCR 2, whose count comes from Hugging Face. A model's weights are the trained files you download and run, and the license says how you may use them.
OmniDocBench scores come from its v1.6 leaderboard, updated on September 11, 2026. "Not published" means we found no score for that model on that benchmark. Datalab, which makes Chandra OCR 2, and the makers of dots.mocr ran olmOCR-Bench on their own models. Ai2 runs olmOCR-Bench and also makes olmOCR 2.
| Model | Size | License | OmniDocBench | olmOCR-Bench |
|---|---|---|---|---|
| Small document models (1.2B or fewer)5 tools | ||||
| TeleOCRAugust 2026 | 1.2B | Apache 2.0 | 96.91 | Not published |
| OvisOCR2July 2026 | 0.8B | Apache 2.0 | 96.47 | Not published |
| PaddleOCR-VL-1.6May 2026 | 0.9B | Apache 2.0 | 96.34 | Not published |
| MinerU2.5-ProMay 2026 | 1.2B | Apache 2.0 (weights) | 95.75 | Not published |
| GLM-OCRFebruary 2026 | 0.9B | MIT | 95.22 | Not published |
| Larger document models (3B to 7B)4 tools | ||||
| Chandra OCR 2March 2026 | 5.3B | Modified OpenRAIL-M | Not published | 85.8 |
| dots.mocrMarch 2026 | 3B | MIT | Not published | 83.9 |
| olmOCR 2October 2025 | 7B | Apache 2.0 | Not published | 82.4 |
| DeepSeek-OCR 2January 2026 | 3B | Apache 2.0 | 90.25 | Not published |
| General vision-language models3 tools | ||||
| Gemini 3 ProClosed, API only | Not published | Closed | 92.91 | Not published |
| Qwen3-VL-235BSeptember 2025 | 235B | Apache 2.0 | 89.78 | Not published |
| GPT-5.2Closed, API only | Not published | Closed | 86.59 | Not published |
| Classic OCR engines2 tools | ||||
| PaddleOCR PP-OCRv6June 2026 | 1.5M to 34.5M | Apache 2.0 | Not published | Not published |
| Tesseract 5.5.3July 2026 | n/a | Apache 2.0 | Not published | Not published |
Which OCR model fits your job#
Start from the job the model has to do. Then check the license and the hardware you have.
What do you need OCR for?
What the two benchmarks measure#
OmniDocBench comes from OpenDataLab, the team that also makes MinerU. Its leaderboard now scores models on 1,651 PDF pages. The pages are research papers, textbooks, exam papers, slides, magazines, newspapers, research and financial reports, and handwritten notes. The overall score averages three parts, for text, tables and math formulas. So a third of the score comes from math formulas, which invoices and bank statements rarely contain.
olmOCR-Bench comes from Ai2, which also makes olmOCR. It runs more than 7,000 pass-or-fail tests on 1,400 documents. These include math papers, old scanned textbooks and letters, dictionary pages and pages with tables. Some tests pass only when page headers and footers are left out of the text. That fits olmOCR's goal of clean, readable text. On a bank statement or an invoice, though, the top of the page often holds the account number or the invoice number. Check that a model trained to drop page headers keeps them. Several scores in Ai2's own table were reported by the model makers, and Ai2 marks them.
The authors of Wild-OmniDocBench point out that standard benchmarks use clean scans and digital PDFs. So they printed OmniDocBench v1.5 pages, bent them and photographed them. They also showed pages on a screen and photographed the screen. Then they scored models on the photos. The first PaddleOCR-VL fell from 91.93 to 72.19, and MinerU2.5 fell from 90.67 to 70.91. The drop was smaller for dots.ocr, at about 10 points, and for Qwen3-VL-235B, at about 9. The benchmark's table covers the first PaddleOCR-VL and MinerU2.5, not PaddleOCR-VL-1.6 or MinerU2.5-Pro. On the same photos, TeleOCR's model card reports 87.36 for the newer PaddleOCR-VL-1.6.
Our take. Use the leaderboards to make a short list of three or four models. Then test them on your worst files and choose the one that does best. Those files are phone photos, faxed forms, 40-page statements and handwritten amounts. A clean PDF only shows how a model handles the easy pages. The pages it gets wrong are the ones your team will fix by hand.
The models in detail#
The model profiles below follow the table's groups. Each fact comes from the model card or the code repository, checked on September 29, 2026.
Small document models
TeleOCR
Small document model- Size
- About 1.2B parameters
- License
- Apache 2.0
- Output
- Text, tables as OTSL tags that its sample code converts to HTML, formulas in LaTeX, and layout boxes
- Runs with
- Hugging Face Transformers; a community GGUF build runs on llama.cpp
OvisOCR2
Small document model- Size
- 0.8B parameters
- License
- Apache 2.0
- Output
- Markdown with text, formulas, tables and image regions
- Runs with
- vLLM
PaddleOCR-VL-1.6
Small document model- Size
- 0.9B parameters
- License
- Apache 2.0
- Output
- Markdown and JSON
- Runs with
- The PaddleOCR toolkit or a vLLM server on an NVIDIA GPU for whole pages. Its Hugging Face Transformers example reads single elements only
MinerU2.5-Pro
Small document model- Size
- 1.2B parameters
- License
- Apache 2.0 for the weights. The MinerU toolkit adds terms to Apache 2.0. Companies above 100 million monthly active users or $20 million in monthly revenue need a separate commercial license. Online services built on it must say they use MinerU
- Output
- JSON, converted to Markdown
- Runs with
- MinerU 4 runs on CPUs by default. Install
mineru[full]for the best speed on an NVIDIA GPU
GLM-OCR
Small document model- Size
- 0.9B parameters
- License
- MIT for the model; its SDK's layout model is Apache 2.0
- Output
- Markdown, or JSON in your schema
- Runs with
- vLLM, SGLang, Ollama, or MLX on Apple silicon; Z.ai also runs a hosted API
Larger document models
Chandra OCR 2
Larger document model- Size
- 5.3B parameters, per Hugging Face
- License
- Code under Apache 2.0. The weights use a modified OpenRAIL-M license. They're free for research, personal use and companies under $2 million in both revenue and funding. They can't be used to compete with Datalab's own API, and other commercial use needs a commercial license from Datalab
- Output
- Markdown, HTML or JSON, with layout
- Runs with
- vLLM or Transformers. Datalab measured 1.44 pages a second on one H100 80 GB GPU
dots.mocr
Larger document model- Size
- 3B parameters
- License
- MIT
- Output
- JSON with layout boxes and text, plus Markdown
- Runs with
- vLLM on a GPU; its repository also describes CPU inference
olmOCR 2
Larger document model- Size
- 7B parameters
- License
- Apache 2.0. Ai2's model card adds that the model is intended for research and educational use under Ai2's Responsible Use Guidelines. Read both before commercial use
- Output
- Markdown, with headers and footers removed
- Runs on
- An NVIDIA GPU with at least 12 GB of memory
Classic OCR engines
PaddleOCR (PP-OCRv6)
Classic OCR engine- Size
- 1.5M, 7.7M or 34.5M parameters
- License
- Apache 2.0
- Output
- Text with box coordinates
- Runs on
- CPUs, Apple silicon or GPUs
Tesseract 5.5.3
Classic OCR engine- License
- Apache 2.0
- Output
- Plain text, hOCR, searchable PDF, TSV, ALTO and PAGE
- Runs on
- CPUs
Other models on the boards
General vision-language models score below the five small document models in the table on OmniDocBench v1.6. The best of them is Ovis2.6-30B-A3B, at 93.70. The best closed one is Gemini 3 Pro, at 92.91, and GPT-5.2 scores 86.59. Qwen3-VL-235B scores 89.78.
The board also lists DeepSeek-OCR 2, released under Apache 2.0 in January 2026, at 90.25. The first version of Tencent's HunyuanOCR scores 89.95. Its license grants no rights in the EU, the UK or South Korea, so you can't use it there.
Both boards list an older version of Mistral OCR. The current Mistral OCR 4.1 came out in July 2026, and our Mistral OCR comparison covers what it returns and costs.
What a model returns, and what your process needs#
An OCR model turns a page image into text, Markdown or JSON. Take the bank statement below. It's Harbor Street Bakery's July 2026 statement from First Midwest Bank, and it lists 142 transactions. A good model can return every page as text, with the transactions as Markdown tables.
Your process needs more than that. It needs all 142 rows in one table, in order, with no row lost or repeated where one page ends and the next begins. Dates printed as 07-14 need to come out in one format, such as 2026-07-14. The balances need to add up. The opening 42,180.55, plus 88,412.10 in deposits, minus 91,686.44 in withdrawals, must equal the closing 38,906.21. And a person needs to check any value the model isn't sure of.
Some models cover part of that. GLM-OCR can fill a JSON schema you give it, and PaddleOCR can merge a table that runs across pages. Your team writes and maintains the code for the rest. That code covers field mapping, which matches page text to named fields. It also covers the balance checks, a review screen and the link to your other systems.
| Field | Extracted value | Confidence |
|---|---|---|
| Account holder | Harbor Street Bakery LLC | |
| Address | 118 Harbor St, Portland, ME | |
| Bank name | First Midwest Bank | |
| Account number | •••• 4821 | |
| Account type | Business checking | |
| Statement period | 2026-07-01 → 2026-07-31 |
| Field | Extracted value | Confidence |
|---|---|---|
| Opening balance | 42,180.55 | |
| Total deposits | 88,412.10 | |
| Total withdrawals | 91,686.44 | |
| Closing balance | 38,906.21 | |
| Transaction count | 142 |
| Field | Extracted value | Confidence |
|---|---|---|
| Date | 2026-07-14 | |
| Description | ACH DEPOSIT STRIPE PAYOUT | |
| Amount | +3,284.10 | |
| Debit / credit | Credit | |
| Running balance | 51,902.44 | |
| Category | Card processor payout |
Docsumo
Platform, not a modelOur product- Reads
- Any document type, handwritten text included. Docsumo's own pre-trained models cover 250+ document types
- Output
- Fields and tables through the API and webhooks. A table that runs across pages comes out as one table, with the headers mapped
- Checks
- On bank statements, each transaction is checked against the statement's running balance
- Review
- Fields the model is unsure about go to your own reviewers. You set the confidence threshold for each field. When a reviewer clicks a field, Docsumo highlights the line on the page it came from, and each correction improves the model
- Accuracy
- Docsumo reports 99% field-level accuracy on 250+ document types. It's a field-level figure, not a score on a public benchmark, so you can't compare it with the scores above
- Pricing
- A free 14-day trial for up to 1,000 pages; Business and Enterprise plans are priced on request
How to test an OCR model on your own documents#
Test your short list on your own pages before you choose. Our guide to measuring OCR accuracy covers the metrics.
- Start with your worst pagesPhone photos, skewed scans and faxed pages. On Wild-OmniDocBench, photos cut some 2025 models' scores by about 20 points.
- Include your longest documentsA 40-page statement tests reading order and tables that run across pages.
- Add handwritingFilled-in forms and handwritten amounts show whether a model reads the writing or invents text where it can't.
- Score the fields you useCount the fields a person had to fix, not the characters the model got right.
- Read the license for your useCheck revenue caps, commercial limits and the countries the license covers.
- Time it on your own hardwarePublished speeds come from other GPUs, batch sizes and settings.
If your pages are invoices, not research papers#
The benchmarks are a fair guide for pages like the ones they test, such as research papers and textbooks. For those pages, start at the top of the table and pick the license that fits. For plain printed text on CPUs, PaddleOCR or Tesseract is enough. Invoices and bank statements need more than OCR. They need named fields and a review step, and Docsumo provides both as an intelligent document processing platform, in the cloud.
Book a demo and bring your hardest pages, or start a free trial.
Frequently asked questions#
Which AI model has the best OCR?
On OmniDocBench v1.6, as updated on September 11, 2026, the top three are small open-source models. TeleOCR scores 96.91, OvisOCR2 96.47 and PaddleOCR-VL-1.6 96.34. The best closed model on that board, Gemini 3 Pro, scores 92.91. On olmOCR-Bench, Datalab reports 85.8 for its Chandra OCR 2, the highest reported score we found for a model you can download. Datalab's hosted API reports 86.7, and other hosted services report up to 88.0 from their own runs.
What is the most accurate OCR?
It depends on your documents. Public benchmarks test few business forms, and scores drop on photographed pages. Run a sample of your own files and count the fields a person had to fix. Our guide to measuring OCR accuracy shows how.
Which OCR model is the fastest?
You can't compare published speeds, because each maker uses its own hardware and settings. GLM-OCR's model card reports 1.86 PDF pages a second. Datalab reports 1.44 pages a second for Chandra OCR 2 on one H100 GPU. For servers without a GPU, look at PaddleOCR's PP-OCRv6 or Tesseract.
What is the best open-source OCR model in 2026?
For whole pages with tables and math formulas, look at the top of OmniDocBench v1.6. The leading open models there are TeleOCR, OvisOCR2, PaddleOCR-VL-1.6, MinerU2.5-Pro and GLM-OCR. Their weights, the model files you download, are released under Apache 2.0 or MIT. For plain printed text on a CPU, see our Tesseract guide.
Is an OCR model the same as OCR software?
No. A model turns a page image into text, Markdown or JSON. OCR software and document AI platforms add what comes after, such as field extraction and review. Compare OCR APIs and OCR software.
Sources
- OpenDataLab: OmniDocBench, leaderboard and README (v1.6 leaderboard, updated September 11, 2026)
- Ai2: olmOCR README and olmOCR-Bench results
- Ai2: olmOCR-Bench, document types and test rules
- Wild-OmniDocBench: real-world captured pages and key results
- Li et al., Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training (arXiv 2603.23885)
- TeleOCR model card (Hugging Face)
- OvisOCR2 model card (Hugging Face)
- PaddlePaddle: PaddleOCR-VL-1.6 model card (Hugging Face)
- PaddleOCR on GitHub: release notes for 3.4.0, 3.6.0 and 3.7.0 (PP-OCRv6)
- OpenDataLab: MinerU2.5-Pro-2605-1.2B model card (Hugging Face)
- OpenDataLab: MinerU on GitHub (MinerU 4 README)
- OpenDataLab: MinerU Open Source License
- Z.ai: GLM-OCR model card (Hugging Face)
- Z.ai: GLM-OCR on GitHub
- Datalab: Chandra OCR 2 model card, benchmarks and commercial usage (Hugging Face)
- Unsiloed: olmOCR-Bench result for Unsiloed Parser, self-reported (May 20, 2026)
- dots.mocr model card (Hugging Face)
- Ai2: olmOCR-2-7B-1025-FP8 model card (Hugging Face)
- DeepSeek: DeepSeek-OCR 2 model card (Hugging Face)
- Tencent: HunyuanOCR license (Tencent Hunyuan Community License Agreement)
- Qwen: Qwen3-VL-235B-A22B-Instruct model card (Hugging Face)
- Tesseract OCR on GitHub
- Mistral AI: OCR 4.1 model page
First published .