Data extraction API: how it works, how to choose one and how to use it

For developers and operations leads adding document extraction to their systems: what these APIs return, what to check before you pick one, and the 5 steps from API key to production.

Illustration of a developer writing API code that returns validated results

Key takeaways

  • A data extraction API takes a source, such as a PDF, a scan or a web page, and returns the data in it as structured output, usually JSON.
  • There are four kinds: OCR APIs return text, document extraction APIs return named fields and tables, web data extraction APIs return content from web pages, and database and application APIs return records a system already holds.
  • A good document extraction API returns a confidence score and a location for every value, so doubtful values can go to a person and any value can be traced to its source.
  • Document APIs are often asynchronous: you upload a file and get the result later, by webhook or from a status endpoint, so plan for retries, rate limits and review from the start.
  • Compare APIs on total cost per document, including the time people spend fixing values, not on the price per page.
On this page
  1. Types of data extraction API
  2. How a document extraction API call works
  3. The best data extraction APIs for documents
  4. Document extraction APIs compared
  5. How to choose a data extraction API
  6. How to use a data extraction API in 5 steps
  7. The bottom line
  8. Frequently asked questions

A data extraction API is a web service that takes a source, such as a PDF, a scanned form or a web page, and returns what's in it as structured data, usually JSON. For business documents, a document extraction API identifies the document type and returns named fields and tables, each with a confidence score and its place on the page, instead of a block of raw text.

This guide covers the four kinds of data extraction API, what happens in a document extraction call, the main options, how to choose one and how to put it into production.

Types of data extraction API#

Which kind you need depends on where the data lives.

  • OCR APIs

    Return the text on a page image, with each word's position and a confidence score. You write the rules that turn text into fields. See the best OCR APIs.
  • Document extraction APIs

    Return named fields and tables for a document type: an invoice's number and line items, the boxes on a tax or insurance form, the tables in a report or research paper. Smart, intelligent and form data extraction APIs are all this kind, and it's the one for business documents.
  • Web data extraction APIs

    Fetch web pages and return their content as clean data: an article's main text without menus and ads, or a product's name, price and reviews. Content extraction and scraping APIs usually mean this kind.
  • Database and application APIs

    Return records a system already holds, from a database or a software product's REST API. Nothing is read: the data is already structured.

How a document extraction API call works#

You send a file, the service reads and checks it, and the result comes back as JSON, often by webhook, because a long document can take a while.

  • PDFs and scans
  • Phone photos
  • Email attachments
Extraction API
  1. 01Classify the document
  2. 02Extract fields and tables
  3. 03Validate the values
  4. 04Review low-confidence fields
JSON to your system by webhook
What happens between the upload and the webhook

Field names differ by vendor, but the response usually carries the same things: the document type, each value, a confidence score and the page the value came from, often with its coordinates.

{
  "id": "doc_7Q2K9",
  "type": "invoice",
  "fields": {
    "invoice_number": { "value": "INV-20418", "confidence": 0.99, "page": 1 },
    "invoice_date": { "value": "2026-09-14", "confidence": 0.98, "page": 1 },
    "po_number": { "value": "PO-7731", "confidence": 0.64, "page": 1 },
    "total": { "value": 1260.00, "confidence": 0.99, "page": 1 }
  },
  "line_items": [
    { "description": "Work table", "qty": 2, "unit_price": 385.00, "amount": 770.00,
      "confidence": 0.97, "page": 1 },
    { "description": "Sheet pan rack", "qty": 2, "unit_price": 245.00, "amount": 490.00,
      "confidence": 0.96, "page": 1 }
  ]
}

The PO number at 0.64 is the value to send to a person. What you do with low-confidence values decides how much you can trust the rest.

The best data extraction APIs for documents#

Docsumo is our product, so we've put it first and said plainly what it doesn't do. The other three are the document AI services of the big cloud providers. Details come from each vendor's own website, checked in September 2026; links are in the sources.

1. Docsumo

IDP platformOur product
Docsumo is an intelligent document processing (IDP) platform with an API. You send a document; it extracts the fields and tables, checks them, and returns structured data through its API and webhooks. On the Business plan it also classifies and splits files that hold several documents. Fields below the confidence threshold you set go to your reviewers first, and their corrections improve the model. It supplies the software, not a review team, and it runs in the cloud only.
Reads
Any business document, printed or handwritten; pre-trained models for 250+ document types, including invoices, bank statements and tax forms
Output
JSON through the API and webhooks, or an Excel export
Testing
A test environment on the Business plan
Security
SOC 2 Type 2 audited, ISO/IEC 27001:2022 certified, and HIPAA and GDPR compliant; your documents aren't used to train shared or third-party models (security)
Pricing
A free 14-day trial for up to 1,000 pages, with the API and webhooks; Business and Enterprise plans are quoted
Best fitLending, insurance and AP teams that need checked fields, not raw text
  • 99%field-level accuracy across 250+ document types
  • 95%+of documents processed straight through, without manual review
  • 3,000+hours a month saved at Arbor

2. Amazon Textract

Cloud document AI API
AWS's machine learning service for text, handwriting, forms, tables and signatures, with separate APIs for invoices and receipts (AnalyzeExpense), IDs (AnalyzeID) and mortgage packages (Analyze Lending).
Reads
PNG, JPEG, TIFF and PDF; printed text in 6 languages, handwriting in English only
Output
JSON blocks with coordinates and a confidence score for each item
Pricing
Text $1.50, tables $15 and forms $50 per 1,000 pages (AWS's US West example); a 3-month free tier for new AWS customers
Best fitDevelopers on AWS who want OCR and form extraction as building blocks

3. Google Document AI

Cloud document AI API
Google Cloud's document service: Enterprise Document OCR for text in 200+ languages, Form Parser, a Gemini-based Layout Parser, a Custom Extractor, and pre-trained parsers for documents such as invoices, bank statements, pay slips, W-2s and US driver's licenses.
Reads
PDFs and images, including handwriting in some languages
Output
A JSON document object with text, layout, key-value pairs, tables and entities
Pricing
OCR $1.50 per 1,000 pages after the first 1,000; Form Parser and Custom Extractor $30 per 1,000 pages
Best fitDevelopers on Google Cloud who want OCR and custom extraction in one service

4. Azure Document Intelligence

Cloud document AI API
Microsoft's document extraction service, now called Azure Document Intelligence in Foundry Tools (formerly Form Recognizer), with read and layout models, prebuilt models for common documents, and custom extraction and classification models.
Reads
Printed text in 300+ languages and handwriting in 12; prebuilt models for invoices, receipts, IDs, bank statements, US pay stubs, US mortgage forms and US tax forms
Output
JSON through a REST API and SDKs for C#, Python, Java and JavaScript
Pricing
Read $1.50, prebuilt models $10 and custom extraction $30 per 1,000 pages; 500 free pages a month
Best fitTeams on Azure, including those that need to run it in containers

Document extraction APIs compared#

ToolReturnsRuns onPricing
IDP platforms1 tool
DocsumoOur productChecked fieldsTablesCloudFree trial; quoted plans
Cloud document AI APIs3 tools
Amazon TextractTextFormsTablesAWS$1.50 per 1,000 pages (text)
Google Document AITextLayoutEntitiesGoogle Cloud$1.50 per 1,000 pages (OCR)
Azure Document IntelligenceTextTablesKey-value pairsAzure, containers$1.50 per 1,000 pages (read)

How to choose a data extraction API#

  • Models for your documentsCheck it has pre-trained models for the documents you actually receive, and what happens with the ones it doesn't know.
  • Confidence and location on every valueWithout them you can't route doubtful values to a person or show a reviewer where a number came from.
  • A review stepDecide who fixes low-confidence values: a review queue in the product, or screens your team builds and maintains.
  • WebhooksResults pushed to you when a document is done, so you don't have to keep asking.
  • SecurityCertifications such as SOC 2 Type 2, how long files are kept, and whether your documents train shared models.
  • SDKs and a test environmentClient libraries for your language, and somewhere to try changes before production.
  • Pricing modelPer page, per document or per plan. Compare the total cost per document, including review time.

Price per page is only part of the cost. Put in your own volumes:

Total cost per document

Per-page prices look small until you add the time people spend fixing what the tool gets wrong.

From the tool's pricing page or your quote
Tool cost per month
$200
Review time per month
$1,167
Total per month
$1,367
Total cost per document
$0.27
How it's worked out
  • Tool cost = documents × pages × price per page.
  • Review time = documents × the share reviewed × minutes per review ÷ 60 × hourly cost.
  • Compare tools on the total per document, not the price per page: a cheaper tool that sends more documents to review can cost more.

How to use a data extraction API in 5 steps#

Write down what you need first: the document types, the fields, your monthly volume and how fast results must come back. Then:

  1. Get an API keySign up for a trial or a cloud account, create a key, and read the API docs for file size limits, rate limits and the current version. Keep the key on your server or in a secrets manager, never in browser code. Most APIs take it, or an OAuth 2.0 bearer token, in a request header over HTTPS.
  2. Send a test documentPost one real file with cURL or Postman before writing code: the file or a link to it, the document type if you know it, and your own reference ID so the result maps back to your record.
  3. Take results by webhookExtraction can take seconds or minutes, so many document APIs work asynchronously: they call a webhook URL you give them when a document is done, or you check a status endpoint until it is. Reply to the webhook at once with a 2xx code and process the payload in the background.
  4. Map the JSON to your systemMap the vendor's field names to yours, put dates and amounts in one format, and keep the page and position of each value so anyone can trace it to the source. See schema mapping.
  5. Handle errors and low confidenceRetry 429 (too many requests) and 5xx responses with growing waits, and honor any Retry-After header. Fix other 4xx errors, such as a bad file or an expired key, instead of retrying them. Send low-confidence values to a person, not into your system.

In the first weeks, compare the number of documents sent with the number of results received, and check a sample of values against the documents themselves.

“We love Docsumo for its ease of integration and the flexibility it provides with API and webhook callbacks.” Howard Leiner, CTO, Arbor

The bottom line#

Pick the kind of API by where your data lives: OCR APIs for text, web data extraction APIs for web pages, and document extraction APIs for named, checked fields from business documents. The big cloud services give developers building blocks; an IDP platform such as Docsumo adds the checks and the review step, so your team only handles the values it's unsure about.

Book a demo with a few of your own documents, or start a free trial and call the API on them.

Frequently asked questions#

What is a data extraction API?

A data extraction API is a web service that takes a source, such as a document, an image or a web page, and returns the data in it in a structured format, usually JSON. Document extraction APIs return named fields and tables, with a confidence score for each value.

What is a content extraction API and how does it work?

A content extraction API pulls the main content out of a web page or a file and returns it as clean, structured data. It fetches or receives the source, works out its layout, separates the content from everything around it, and returns text, fields or tables as JSON.

What is a web data extraction API?

A web data extraction API fetches web pages and returns their data in a structured form, such as an article's text or a product's name, price and reviews, so you don't have to write and maintain a scraper for each site. Product data extraction APIs are one kind. For PDFs and scans, use a document extraction API instead.

How do extraction APIs handle boilerplate removal and provenance?

Boilerplate removal drops what isn't content: menus, ads and footers on a web page, or the headers and footers repeated on every page of a document. Provenance means each value can be traced to its source: the URL and time it was fetched for web data, or the page and position on the page for a document field, usually with a confidence score.

How do you get data from an API?

Get an API key or token, send an HTTP request to the right endpoint with the key in a header, and read the JSON response. Document extraction APIs add a step: you upload the file, then receive the result by webhook or by checking a status endpoint.

What's the difference between web scraping and an API?

A scraper reads a web page's HTML the way a browser does, and breaks when the page layout changes. An API returns data in a documented, structured format, so the provider handles the parsing. See data extraction vs data scraping.

See Docsumo read your own documents

Bring a few real samples. We'll show the fields extracted, the checks that ran and what a reviewer would see.