Insurance data extraction: the documents, the fields and how to automate it
For underwriting, claims and operations teams at carriers, MGAs, brokers and TPAs: which documents to extract data from, the fields in each, and how intelligent document processing automates the work.

Key takeaways
- Insurance data extraction turns applications, loss runs, claim documents and certificates into structured fields that underwriting, claims and policy systems can use.
- The documents fall into three workflows: submissions (ACORD 125, 126 and 140, statements of values, loss runs), claims (loss notices, estimates, medical bills) and policy servicing (declarations pages, endorsements, certificates).
- Intelligent document processing for insurance classifies each document, extracts its fields and tables, validates them and sends only uncertain values to a person.
- The hard parts are poor scans, handwriting, tables that run across pages and checks between documents, so test any tool on your own worst files.
On this page
- Which insurance documents need data extraction
- How intelligent document processing for insurance works
- What to extract from each insurance document
- Example: an ACORD 25, read field by field
- Manual vs automated document extraction in insurance
- What to test before you automate
- The bottom line
- Frequently asked questions
Insurance data extraction is reading the documents that carriers, MGAs, brokers and TPAs receive, such as applications, loss runs, claim notices and certificates, and turning them into structured fields for underwriting, claims and policy systems. Done by hand, it means rekeying. When it's automated, software classifies each document, extracts its fields, checks them and sends only the uncertain ones to a person.
This guide covers the documents by workflow, how intelligent document processing handles them, the fields worth extracting and what to test before you automate.
Which insurance documents need data extraction#
Most of the work falls into three workflows, each feeding a different system:
Submissions and underwriting
The ACORD 125 application with line sections such as the ACORD 126 and ACORD 140, plus statements of values and loss runs. The data drives clearance, appetite and rating.Claims
The first notice of loss, often on an ACORD property, auto or liability loss notice, then photos, estimates, police and adjuster reports, invoices and medical bills. The data opens the claim and supports reserves and payment.Policy servicing and compliance
Declarations pages, endorsements and change requests, plus the certificates of insurance (ACORD 25) and evidence of property insurance (ACORD 28) that vendors and borrowers send to prove coverage.
How intelligent document processing for insurance works#
In insurance, intelligent document processing (IDP) is AI-driven, automated document processing: it covers the whole path from ingestion to your systems, not just the reading. Documents arrive by broker email, portal uploads, scans and faxes, often several to a PDF.
- Broker email
- Portal uploads
- Scans and faxes
- 01Classify and split
- 02Extract fields and tables
- 03Validate
- 04Review exceptions
Classification splits a submission packet into its application, line sections and loss runs. Extraction pulls each field and table with a confidence score. Validation checks the values against your rules and the other documents in the file, and review sends a person only the uncertain fields and failed checks. The data then goes to your systems through an API and webhooks.
Claims processing works the same way, one document at a time over the life of the claim: the first notice of loss, then photos, estimates and adjuster reports, then invoices or medical bills before payment. Each is read as it arrives and added to the claim file, so adjusters review the file instead of keying it. More in insurance claims automation.
What to extract from each insurance document#
The fields worth extracting are the ones a decision or a check depends on:
| Document | Fields to extract | Used for |
|---|---|---|
| ACORD 125 application | Named insured, FEIN, NAICS or SIC code, lines requested, proposed dates, locations, prior carriers, losses | Clearance, appetite, setting up the quote |
| ACORD 126 and 140 line sections | Liability limits and hazard class codes (126); building values, construction, occupancy and protection (140) | Rating |
| Statement of values | For each location: address, the values insured (building, contents, business income) and construction details | Property exposure and rating |
| Loss runs | Policy number, coverage period, number of claims, paid losses, date of each loss | Pricing; checking the losses the application reports |
| Loss notices (ACORD 1, 2 and 3) | Policy number, insured, date, time, place and cause of loss, description, contacts | Opening the claim; coverage check |
| Estimates and invoices | Line items, labor, parts, tax, totals | Reserves and payment |
| Medical bills (CMS-1500, UB-04) | Patient, provider, dates of service, procedure and diagnosis codes, charges | Injury and workers' compensation claims |
| Declarations pages and endorsements | Insured, policy number, coverages, limits, deductibles, endorsement numbers and dates | Keeping the policy record current |
| Certificates and evidence (ACORD 25, 28) | Insurers, policy numbers, dates, limits, additional insured, certificate holder or mortgagee | Vendor and lender compliance |
Example: an ACORD 25, read field by field#
Here's extraction on one common insurance document, an ACORD 25 certificate of liability insurance: the parties, each policy and its insurer, the limits and the certificate holder, each in its own group. Loss runs, loss notices and applications work the same way; only the fields and the checks change. See ACORD form extraction for the rest of the ACORD family.
| Field | Extracted value | Confidence |
|---|---|---|
| Issue date | 2026-09-18 | |
| ACORD version | 2016/03 | |
| Producer | Allegheny Risk Partners LLC | |
| Insured | Keystone HVAC LLC | |
| Insurer A · NAIC | Brightwater Casualty Company · 21873 | |
| Insurer B · NAIC | Brightwater Specialty Insurance Co. · 34418 | |
| Insurer C · NAIC | Laurel Highlands Mutual Insurance Co. · 13052 |
| Field | Extracted value | Confidence |
|---|---|---|
| General liability | GL 7204418 · A · occurrence | |
| Auto liability | CA 7204419 · A · any auto | |
| Umbrella | UMB 3310582 · B · occurrence | |
| Workers comp | WC 5519027 · C · per statute | |
| GL, auto, umbrella term | 2026-07-01 → 2027-07-01 | |
| Workers comp term | 2026-03-15 → 2027-03-15 | |
| Additional insured | General liability, auto | |
| Waiver of subrogation | General liability, auto, WC |
| Field | Extracted value | Confidence |
|---|---|---|
| Each occurrence | 1,000,000 | |
| General aggregate | 2,000,000 | |
| Products-comp/op aggregate | 2,000,000 | |
| Auto combined single limit | 1,000,000 | |
| Umbrella each occurrence | 2,000,000 | |
| E.L. each accident | 1,000,000 |
| Field | Extracted value | Confidence |
|---|---|---|
| Certificate holder | Riverside Plaza LLC | |
| Holder address | c/o Harbor Point Properties, 500 Penn Ave, Suite 1200, Pittsburgh, PA 15222 | |
| Certificate number | CL2609180417 | |
| Description of operations | HVAC maintenance and repair at 1 Riverside Plaza … | |
| Authorized representative | Signed |
Manual vs automated document extraction in insurance#
By hand, someone opens each document and keys it into the policy or claims system. Automated, software does the reading and people handle the exceptions.
Manual keying
- Each application, loss run and claim form is rekeyed by hand
- Tables that run across pages are retyped row by row
- Errors surface later, at quote, renewal or payment
- Documents wait in a queue until someone opens them
Automated extraction
- Documents are classified and read as they arrive, scans and handwriting included
- Tables that run across pages come back as one table, headers mapped
- Every field gets a confidence score, and only unsure ones go to a person
- Data reaches policy, claims and rating systems by API and webhooks
- 99%field-level accuracy across 250+ document types
- 95%+of documents processed straight through, without manual review
- 3,000+hours a month saved at Arbor on insurance compliance documents
Those are Docsumo's figures. Arbor also reached 99% accuracy on insurance compliance documents, and Vertikal RMS uses Docsumo for COI verification through ACORD capture.
What to test before you automate#
Clean sample files make every tool look good. Test these five on your own documents:
- Scan qualityFaxed loss runs, skewed scans and phone photos. Count the fields that come back wrong rather than flagged for review.
- HandwritingLoss notices and applications are often filled in by hand. Docsumo reads handwritten text; check that any tool you test does too.
- Tables across pagesLoss runs, statements of values and itemized estimates should come back as one table, not one per page.
- Validation rulesLoss dates inside the policy period, totals that add up, and values that agree across a submission. Docsumo checks values across documents with cross-document validation, on its Enterprise plan.
- SecurityClaim files hold personal and medical data. Ask for SOC 2 Type 2, ISO/IEC 27001 and a HIPAA business associate agreement (BAA). Docsumo is SOC 2 Type 2 audited and ISO/IEC 27001:2022 certified, and offers a BAA (security).
The bottom line#
Insurance data extraction pays off most on the documents that arrive every day in many layouts: submissions, claim files and certificates. Decide which fields your decisions depend on, check them as they're read, and let people handle only what fails. Start with one workflow and test on your own worst files.
Book a demo with a few of your own insurance documents, or start a free trial.
Frequently asked questions#
What is insurance data extraction?
It's turning the documents that carriers, MGAs, brokers and TPAs receive, such as applications, loss runs, claim forms, certificates and policies, into structured fields in underwriting, claims and policy systems. Automated tools classify each document, extract its fields, check them and send uncertain values to a person.
What is intelligent document processing for insurance?
It's AI software that handles insurance document processing end to end: it sorts a packet into its documents, extracts each field and table, checks the values against rules and the other documents in the file, and routes exceptions to a person. See intelligent document processing for insurance.
How is intelligent document processing used in claims processing?
It reads each claim document as it arrives, checks it and adds the data to the claim file, so adjusters handle only the exceptions. The first notice of loss gives the policy number, the insured, the date, time, place and cause of the loss; estimates and invoices add line items and totals; medical bills add providers, dates of service, codes and charges. See OCR claims processing for the fields by claim type.
Can OCR alone extract data from insurance documents?
No. OCR turns a scan into text, but it doesn't know which number is the policy number or the per-occurrence limit, and it checks nothing. Extraction adds the fields, validation and review: see OCR in insurance.
How accurate is automated insurance data extraction?
It depends on the documents and the tool, so test on your own files. Docsumo reports 99% field-level accuracy across 250+ document types, and Arbor reached 99% accuracy on insurance compliance documents. Fields the model is unsure about go to a person for review.
What claims data extraction software is best for insurance carriers?
It depends on the job: core claims systems, claims AI and document automation solve different problems. For document intake, look for classification of mixed claim packets, handwriting, tables that run across pages and a review queue, and test on your own claim files. Our comparison of claims automation software covers the options.