OCR for insurance documents: what it reads, where it stops and how to automate the rest

For underwriting, claims and compliance teams at carriers, agencies, MGAs and TPAs: what OCR does with insurance paperwork, where it stops and how to turn scans into checked data.

Illustration of an insurance document under a magnifying glass, with its claim amount extracted into a field to accept or reject

Key takeaways

  • OCR in insurance is the use of optical character recognition to turn scanned, faxed and photographed insurance documents into machine-readable text.
  • It reads ACORD applications, certificates of insurance, loss runs, claim forms, medical bills, member ID cards and policies, but it returns text, not fields.
  • To get usable data, insurers add AI extraction, validation and a review step: each value lands in its field, dates and limits are checked, and only unsure values reach a person.
  • The hard parts are poor scans, handwriting, tables that run across pages, layouts that vary by carrier and sensitive data.
  • Automating the whole loop turns document handling into exception handling: people see only the fields and files that fail a check.
On this page
  1. What is OCR in insurance?
  2. OCR vs OCR plus AI extraction
  3. How to extract data from insurance documents with OCR
  4. OCR use cases in insurance
  5. Challenges of OCR in insurance, and how to handle them
  6. How Docsumo handles insurance documents
  7. The bottom line
  8. Frequently asked questions

OCR in insurance is the use of optical character recognition (OCR) to turn scanned, faxed and photographed insurance documents, such as applications, certificates of insurance, loss runs, claim forms and policies, into machine-readable text. OCR reads the characters but not what they mean: it can't tell a policy number from a claim number or check a limit against a contract. So insurers, agencies and TPAs pair it with AI extraction, validation and a review step to get checked data into their policy and claims systems.

This guide covers the documents OCR reads, where it stops, how the process works, the main use cases and the problems to plan for.

What is OCR in insurance?#

OCR stands for optical character recognition: software that finds the text in an image of a page and turns it into characters a computer can search and process. In insurance, it's the first step in reading documents that arrive as scans, faxes, emailed PDFs and phone photos. The main document types:

  • ACORD applications

    The ACORD 125 commercial insurance application, plus line sections such as the ACORD 126 for general liability and the ACORD 140 for property.
  • Certificates of insurance

    Evidence of the coverage a vendor, tenant or borrower holds, to check against what a contract requires. US liability certificates usually come on the ACORD 25; Australia calls them certificates of currency.
  • Loss runs

    The insurer's report of claims on a policy: policy number, coverage period, number of claims, paid losses and the date of each loss. Underwriters read them to judge a risk's claims history.
  • Claim forms and loss notices

    ACORD property, auto and liability loss notices, plus the police reports, estimates and invoices behind a claim.
  • Health claims and ID cards

    CMS-1500 (professional) and UB-04 (institutional) claim forms, explanation of benefits (EOB) statements, and member ID cards with the member and group numbers. See the claim form guide.
  • Policies and declarations

    The declarations page names the insured, the insurer and the agent or broker, and lists each coverage with its limit and premium. Endorsements change the terms.

OCR vs OCR plus AI extraction#

OCR's output is text and where it sits on the page. It can't tell which of the dates on an ACORD 125 is the proposed effective date, or whether a limit on a certificate meets the contract. AI extraction uses layout and context to put each value in its field, then checks it.

OCR only

  • Returns a block of text, not the policy number or the limit
  • Getting fields takes a template per layout, which breaks when a form changes
  • Passes every value through unchecked
  • Leaves someone to key the data into your systems

OCR plus AI extraction

  • Puts each value in its field: insured, policy number, dates, limits, amounts
  • Works across carrier and broker layouts without templates
  • Checks values and flags uncertain ones for a person
  • Sends the data to policy and claims systems by API

For the claims side of this comparison, with the fields to pull by claim type, see OCR claims processing.

How to extract data from insurance documents with OCR#

Every document goes through the same stages, whether it's a submission, a certificate or a claim. Only the fields and the checks change.

  • Scans
  • Faxes
  • Emailed PDFs
  • Phone photos
OCR + AI extraction
  1. 01Clean up the image
  2. 02Read the text
  3. 03Extract fields and tables
  4. 04Check and review
Policy and claims systems
How an insurance document becomes checked data
  1. CapturePull email, portal uploads, faxes and mail scans into one queue.
  2. Clean upStraighten, de-noise and sharpen scans and photos so the characters are clear.
  3. Classify and readSort pages by document type, such as application, loss run or claim form, and run OCR, handwriting included.
  4. Extract and checkPull named fields and tables, then check them: dates in range, totals that add up, values that agree across documents.
  5. Review and deliverSend low-confidence fields to a person, then pass the data to your policy, claims or agency management system by API.

OCR use cases in insurance#

The same pipeline serves four jobs across the policy lifecycle.

  • Submissions and underwriting

    Read ACORD applications, loss runs and statements of values from broker submissions, so underwriters review risks instead of keying them. See ACORD form processing.
  • Claims intake

    Turn loss notices, claim forms, estimates and medical bills into claim data, and see what's missing before an adjuster opens the file.
  • Policy servicing

    Read declarations pages, endorsements and change requests to keep the policy record in step with the documents.
  • COI compliance

    Read certificates of insurance and check limits, dates and additional insured status against each contract. See COI tracking.

For COI compliance, OCR plus AI extraction puts each policy, limit and date on an ACORD 25 in its own field, ready to check against the contract:

Parties
FieldExtracted valueConfidence
Issue date2026-09-18
ACORD version2016/03
ProducerAllegheny Risk Partners LLC
InsuredKeystone HVAC LLC
Insurer A · NAICBrightwater Casualty Company · 21873
Insurer B · NAICBrightwater Specialty Insurance Co. · 34418
Insurer C · NAICLaurel Highlands Mutual Insurance Co. · 13052
An ACORD 25 certificate of liability insurance being read.

Challenges of OCR in insurance, and how to handle them#

Plan for these five before you pick a tool:

  • Scan qualityFaxes, skewed scans and phone photos blur characters. Clean up images before OCR and send low-confidence values to review.
  • HandwritingClaim forms and applications are often filled in by hand. Use a tool that reads handwritten text, as Docsumo does, and keep a person on the values it's unsure about.
  • Tables across pagesLoss runs, statements of values and itemized bills often run over several pages. Extraction should join each into one table with the headers mapped.
  • Layouts that varyLoss runs, declarations pages and medical bills differ from one carrier, broker or provider to the next, and ACORD forms change by edition. Avoid tools that need a template per layout.
  • Sensitive dataClaim and health files carry personal and medical information. Ask for SOC 2 Type 2, ISO/IEC 27001 and, for health data, a HIPAA business associate agreement (BAA). Docsumo is SOC 2 Type 2 audited, ISO/IEC 27001:2022 certified and HIPAA compliant, and signs a BAA (security).

How Docsumo handles insurance documents#

Docsumo is an intelligent document processing platform. It reads printed and handwritten text and handles any document type, from ACORD forms and certificates to loss runs and claim files. On the Business plan it classifies mixed packets and splits them into documents; on the Enterprise plan it checks values across documents and manages each file as a case.

Fields it's unsure about go to your own team, with confidence thresholds set per field and each value linked to its line on the page, and reviewers' corrections improve the model. Tables that run across pages are joined into one table, with the headers mapped. Data goes to your policy, claims or agency management system through an API and webhooks. Docsumo isn't a policy or claims system, and it runs in the cloud only.

  • 99%field-level accuracy across 250+ document types
  • 95%+of documents processed straight through, without manual review
  • 99%accuracy on insurance compliance documents at Arbor

Vertikal RMS uses Docsumo to verify certificates of insurance through ACORD capture, and Arbor uses it on insurance compliance documents.

The bottom line#

OCR turns insurance paperwork into text, not into checked data. Pair it with AI extraction, validation and a review step, and applications, loss runs, claims and certificates become fields your underwriting, claims and compliance teams can use, with people handling only the exceptions.

Book a demo with a few of your own insurance documents, or start a free trial.

Frequently asked questions#

What is OCR in insurance?

OCR in insurance is the use of optical character recognition to read insurance paperwork, such as applications, certificates, loss runs, claim forms and policies, and turn scans, faxes and photos into machine-readable text. To get usable data, teams pair it with AI extraction so the text becomes named, checked fields.

What does OCR stand for in insurance?

Usually optical character recognition. In US health insurance, OCR can also mean the HHS Office for Civil Rights, which enforces the HIPAA Privacy, Security and Breach Notification Rules. Actuaries also use OCR for the outstanding claims reserve, the provision for claims not yet settled.

Is OCR enough to automate insurance document processing?

No. OCR gives you text, not which value is the effective date or the per-occurrence limit, and it checks nothing. Intelligent document processing adds extraction, validation and review: see IDP vs OCR.

How is OCR used in insurance underwriting?

Underwriters use OCR and AI extraction to read submissions: ACORD applications, loss runs and statements of values in commercial lines. In life insurance, traditional underwriting also draws on medical records, exam and lab results, and financial and tax information, which extraction turns into data the underwriter can review.

Can OCR read health insurance cards and claim forms?

Yes. OCR reads member ID cards, CMS-1500 and UB-04 claim forms and explanation of benefits (EOB) statements, and AI extraction maps the member and group numbers, codes and charges to fields. See document processing for healthcare.

Can OCR extract data from an insurance policy or policy schedule?

Yes. OCR and AI extraction can read the declarations page, or policy schedule, which names the insured, the insurer and the agent or broker, describes the insured property or location, and lists each coverage with its limit and premium, plus the deductibles and endorsements.

See Docsumo read your own documents

Bring a few real samples. We'll show the fields extracted, the checks that ran and what a reviewer would see.