Cost of bad data statistics: 82 figures on poor data quality, traced to the source
For writers, analysts and data leaders who need a number they can defend. Every figure is numbered, dated and linked to the organization that published it, and the famous old ones are traced to where they started.

Key takeaways
- The famous $12.9 million a year comes from 154 reference customers of data quality vendors, surveyed by Gartner in 2020. The $3.1 trillion credited to IBM traces back to a 2011 magazine article whose author called it his own extrapolation.
- Newer surveys measure other things. Data professionals said data quality issues affected 31% of revenue in 2023 (Monte Carlo), and a 2024 survey tied 6% of revenue to AI models built on bad data (Fivetran).
- Gartner predicts organizations will abandon 60% of AI projects that lack AI-ready data through 2026. At large financial institutions, data quality is the most cited obstacle to production AI, named by 84% (IIF and EY, 2026).
- Regulators put a price on it. Citi was fined $400 million in 2020, partly over data governance, and $135.6 million more in 2024 over data quality management.
- Documents are where much of it starts. 50% of US healthcare providers blame missing or inaccurate data for rising claim denials (Experian Health, 2025).
On this page
- The famous cost of bad data figures, traced
- What bad data costs organizations
- How much data is wrong
- Time spent finding and fixing bad data
- How far people trust their data
- Bad data and AI projects
- Fines and findings in financial services
- When bad data made headlines
- Bad data by industry
- Where bad data starts, and the 1-10-100 rule
- How to put a number on your own cost of bad data
- How we picked and checked these figures
- If your bad data starts in documents
- Frequently asked questions
The cost of bad data is usually quoted as $12.9 million a year for the average organization, a Gartner figure from 2020, or $3.1 trillion a year for the US economy, a number credited to IBM that we traced to one writer's rough extrapolation in 2011. Newer surveys measure it in other ways. In 2023, data professionals estimated that data quality issues affected 31% of their company's revenue, and a 2024 survey tied a loss of 6% of revenue to AI models built on bad data.
Below are 82 statistics on what poor data quality costs, grouped by where the damage shows up. Each one is numbered so you can link to it, and each links to the organization that published it. We checked every figure on the publisher's own page or report in September 2026.
The famous cost of bad data figures, traced#
Most articles on the cost of poor data quality repeat the same few numbers. Several are more than ten years old, and a few measured something narrower than the version that gets quoted. We followed each one back to where it started.
| Often quoted | What the source says | Cite instead |
|---|---|---|
| Poor data quality costs organizations $12.9 million a year (Gartner) | An estimate by 154 reference customers of 16 data quality vendors, gathered for Gartner's 2020 Magic Quadrant research. The companies were already buying data quality software, and the figure is what they believed bad data cost them | Keep it, but say it's from 2020 and who was asked |
| Bad data costs the US economy $3.1 trillion a year (IBM, 2016) | The earliest source we found is a 2011 SOA World Magazine article whose author extrapolated it from one healthcare estimate and added "no matter how far off my estimate is." An IBM infographic repeated it by 2013 with no method, and Thomas Redman's 2016 Harvard Business Review article made it famous as IBM's | No national estimate with a published method exists. Use a company-level figure and name who was surveyed |
| Poor data quality costs $15 million a year (Gartner), or $9.7 million, or $8.2 million | All older Gartner figures. $15 million is from Gartner's 2017 Data Quality Market Survey, published in 2018. A January 2017 article gave $9.7 million, and a 2009 study said more than $8 million | $12.9 million (Gartner, 2020), with who was asked |
| Bad data costs companies 15% to 25% of revenue (MIT Sloan) | A columnist's estimate in MIT Sloan Management Review in 2017, not MIT research. Thomas Redman built it mainly on an Experian survey in which respondents believed 23% of revenue was wasted, plus two consultants' figures | Cite it as Redman's estimate. For a survey figure, 31% of revenue affected (Monte Carlo, 2023) |
| Over a quarter of organizations lose more than $5 million a year to poor data quality, and 43% of COOs call data quality their top data priority (IBM) | The $5 million figure is from a Forrester data snapshot published in 2024, and it describes data and analytics employees who already see poor data quality as an obstacle, not organizations. We didn't find the 43% in IBM's 2025 studies of chief data officers or chief operating officers | Forrester, 2024, with that base. Or only 26% of chief data officers trust their data to support new AI revenue (IBM, 2025) |
| Bad data costs US businesses $600 billion a year (TDWI) | A 2002 TDWI estimate of postage, printing and staff time wasted on bad customer name and address records. TDWI said the full cost was much higher | Label it 2002 and mailing only, or leave it out |
| Only 3% of companies' data meets basic quality standards | 3% of 75 data quality assessments that managers made of their own departments' records, not 3% of companies | 47% of new records had at least one critical error (Nagle, Redman and Sammon, 2020) |
| Data scientists spend 80% of their time cleaning data | CrowdFlower's 2016 survey asked which task takes the most time. 60% named cleaning and organizing data and 19% collecting it, and a magazine write-up added the two together | Data professionals spend 37.75% of their time on preparation and cleansing (Anaconda, 2022) |
| Knowledge workers waste 50% of their time on bad data | Thomas Redman's estimate in Harvard Business Review in 2016, with no study behind it | Cite it as Redman's estimate |
| Poor data quality causes 40% of business initiatives to fail (Gartner) | A 2011 Gartner note says so in a key finding, but its analysis counts targeted benefits lost for any reason, partly from Standish Group project data, and its worked example puts data quality's share of that loss at 10% | Organizations will abandon 60% of AI projects without AI-ready data through 2026 (Gartner, 2025) |
| B2B data decays 2.1% a month, 22.5% a year, or 70.3% a year | The 22.5% traces to a HubSpot page that credits MarketingSherpa, with no report, year or method. The 70.3% isn't Gartner's. It comes from a consultant's 2009 survey of about 1,200 seminar attendees' business cards, counting any change, promotions included | At least 23% of an email list decays in a year (ZeroBounce, from 11 billion addresses checked in 2025) |
| Bad data costs companies 12% of revenue (Experian), or 30% (Ovum) | The 12% is what respondents to a 2013 Experian survey believed was wasted, mostly through bad contact data, and Experian's 2015 benchmark put it at 23%. We found no Ovum report, date or method behind the 30% | 62% of marketing organizations say they lose revenue directly to poor CRM data (Validity, 2026) |
| Sales reps waste 27.3% of their time, 546 hours a year, on bad data | A 2011 analysis by LeadJen of 12 of its own outbound lead generation campaigns, about inside sales reps calling lists nobody had cleaned. It's sometimes credited to Gartner | 46% of sales pros using AI agents say data quality issues hurt their sales (Salesforce, 2026) |
| 84% of CEOs worry about the quality of the data behind their decisions | KPMG's 2016 Global CEO Outlook. The next year's edition put the share at 56% | 67% of data professionals don't fully trust the data they use for decisions (Precisely and Drexel LeBow, 2024) |
Even Gartner's own number has moved. Its research has put the average annual cost of poor data quality anywhere from about $8 million to $15 million since 2009. Each figure came from a different group of organizations, so read them as separate estimates rather than a trend.
What bad data costs organizations#
1. Organizations estimate that poor data quality costs them $12.9 million a year on average, according to Gartner's 2020 survey of 154 reference customers of 16 data quality vendors. It's the most quoted cost figure, and it's a self-estimate by companies already buying data quality software (Gartner, 2020).
2. Gartner's 2017 Data Quality Market Survey put the average annual cost of poor data quality at $15 million per organization (Gartner, 2018). A Gartner article in January 2017 had cited $9.7 million (Gartner, 2017), and earlier surveys found $14.2 million in 2013 (Gartner, 2013) and more than $8 million in 2009 (Gartner, 2009).
3. Nearly 60% of organizations don't measure the annual financial cost of poor data quality, a Gartner survey found in 2017. Gartner's current guidance says 59% don't measure data quality at all (Gartner, 2018).
4. More than a quarter of data and analytics employees who see poor data quality as an obstacle to data literacy estimate their organization loses more than $5 million a year to it, and 7% put the loss at $25 million or more (Forrester, 2024).
5. Thomas Redman estimates that bad data costs most companies 15% to 25% of revenue, building on Experian research and the work of two consultants. He also estimates that about two-thirds of that cost can be found and removed for good (MIT Sloan Management Review, 2017).
6. Data professionals estimated that data quality issues affected 31% of revenue on average in 2023, up from 26% in 2022, in a survey of 200 people commissioned by Monte Carlo (Monte Carlo and Wakefield Research, 2023).
7. 95% of businesses have seen impacts from poor data quality. They include less reliable analytics (36%), a worse customer experience (32%) and damage to reputation and customer trust (32%) (Experian, 2021).
8. 62% of marketing organizations say they lose revenue directly because of poor CRM data quality, and nearly a third of teams spend six or more hours a week fixing and reconciling data, in a survey of 500 marketers (Validity, 2026).
9. Thomas Redman's rule of ten holds that a unit of work costs 10 times as much to finish when its data are flawed as when they're perfect. By that rule, 11 flawed records in a batch of 100 nearly double the cost of the batch (Harvard Business Review, 2017).
What it costs the US economy
10. The $3.1 trillion a year figure for the US economy, usually credited to IBM, first shows up in a 2011 SOA World Magazine article whose author extrapolated it from a healthcare estimate (Tibbetts, 2011). An IBM infographic repeated it by 2013 without a method (IBM, archived 2013), and it spread after Thomas Redman cited IBM in Harvard Business Review in 2016.
11. TDWI estimated in 2002 that poor customer data cost US businesses $611 billion a year, often rounded to $600 billion. The estimate covered only postage, printing and staff time wasted on bad name and address records (TDWI, 2002).
How much data is wrong#
12. 47% of newly created data records had at least one critical error, on average, across 75 data quality assessments collected over two years from a wide range of organizations. Only 3% of the assessments scored 97% or better, the bar the authors set for acceptable (Nagle, Redman and Sammon, Business Horizons, 2020).
13. Between 10% and 25% of customer and prospect records include critical data errors at any given time, from strong organizations to typical ones, according to a SiriusDecisions brief from 2008 or earlier (SiriusDecisions).
14. Organizations believe about a third of their customer and prospect data is inaccurate in some way, and only 50% think their CRM or ERP data is clean (Experian, 2021).
15. Data and analytics leaders estimate that 26% of their organization's data is untrustworthy, in Salesforce's 2025 survey of 7,652 people in 18 countries (Salesforce, 2025).
16. Poor data quality was the top challenge for 57% of data professionals in 2024, up from 41% in 2022 (dbt Labs, 2024). It was still the most reported problem in 2025, at 56% (dbt Labs, 2025).
17. 64% of data and analytics professionals called data quality their biggest data integrity challenge in 2024 (Precisely and Drexel LeBow, 2025 Outlook).
18. 81% of AI professionals say their company still has significant data quality issues, and 85% say leadership isn't addressing them, in a survey of 500 US data professionals commissioned by Qlik (Qlik and Wakefield Research, 2025).
19. At least 23% of an email list decays within a year, ZeroBounce estimates from the more than 11 billion addresses it checked in 2025, down from 28% in its previous report (ZeroBounce, 2026).
Time spent finding and fixing bad data#
20. Knowledge workers waste 50% of their time in what Thomas Redman calls hidden data factories, hunting for data, correcting errors and checking data they don't trust. It's his estimate, and the article gives no study behind it (Harvard Business Review, 2016).
21. Organizations say their data scientists spend 67% of their time preparing data rather than building AI models, in a 2024 survey of 550 respondents (Fivetran and Vanson Bourne, 2024).
22. Data professionals spent about 45% of their time loading and cleansing data in 2020 (Anaconda, 2020), and 37.75% on data preparation and cleansing in 2022 (Anaconda, 2022).
23. Data teams handled 67 data incidents a month on average in 2023, up from 59 in 2022 (Monte Carlo and Wakefield Research, 2023).
24. Resolving a data incident took 15 hours on average in 2023, up 166% in a year (Monte Carlo and Wakefield Research, 2023).
25. 70% of data leaders said data incidents take more than four hours to detect, in a 2024 survey of 200 US data engineers and leaders (Monte Carlo and Wakefield Research, 2024).
26. 74% of data professionals said business stakeholders found data issues first, all or most of the time, in 2023, up from 47% a year earlier (Monte Carlo and Wakefield Research, 2023).
How far people trust their data#
27. 67% of data and analytics professionals didn't completely trust the data their organization used for decisions in 2024, up from 55% a year earlier (Precisely and Drexel LeBow, 2025 Outlook).
28. A year later the same study surveyed only leaders, and 67% reported high trust, against 33% in the previous survey. The report says perspectives change when only leaders are asked, so the two editions don't make a trend (Precisely and Drexel LeBow, 2026).
29. 71% of organizations with data governance programs report high trust in their data, against 50% of those without one (Precisely and Drexel LeBow, 2026).
30. Only 35% of sales professionals completely trust the accuracy of their organization's data, in Salesforce's survey of 5,500 sales professionals in 27 countries (Salesforce, 2024).
31. Business leaders' confidence in the accuracy of their data is down 27% since 2023, and confidence in its relevance is down 18%, in Salesforce's survey of 552 US business leaders (Salesforce, 2025).
Bad data and AI projects#
AI puts a new price on data quality, because models and agents act on whatever they're given. Asked what holds AI back, most groups put data near the top of the list.
32. 63% of organizations don't have, or aren't sure they have, the right data management practices for AI, in Gartner's July 2024 survey of 1,203 data management leaders (Gartner, 2025).
33. Through 2026, organizations will abandon 60% of AI projects that aren't supported by AI-ready data, Gartner predicts (Gartner, 2025).
34. Organizations lose an average of 6% of global annual revenue, or $406 million, to underperforming AI models built on inaccurate or low-quality data, respondents to a 2024 survey estimated. Their organizations averaged $5.6 billion in revenue (Fivetran and Vanson Bourne, 2024).
35. Organizations whose AI initiatives succeed invest up to 4 times more, as a share of revenue, in foundations such as data quality and governance than those with poor AI outcomes. Gartner surveyed 353 data, analytics and AI leaders (Gartner, 2026).
36. 55% of organizations have avoided certain generative AI use cases because of data-related issues, in Deloitte's survey of 2,770 business and technology leaders (Deloitte, 2024).
37. 84% of data and analytics leaders say their data strategy needs a complete overhaul before their AI ambitions can succeed (Salesforce, 2025).
38. Only 26% of chief data officers are confident their data can support new AI-enabled revenue streams, in IBM's study of 1,700 senior data leaders (IBM Institute for Business Value, 2025).
39. 57% of data leaders see data reliability as a key barrier to moving AI projects from pilot to production (Informatica, 2026).
40. Only 12% of data and analytics professionals said their data was of sufficient quality and accessibility for AI in 2024 (Precisely and Drexel LeBow, 2025 Outlook).
41. 88% of data leaders said their data was ready for AI in 2025, yet 43% named data readiness as their biggest obstacle to AI (Precisely and Drexel LeBow, 2026).
42. Data quality is the most cited challenge to launching AI in production at large financial institutions, named by 84%, ahead of skilled staff at 77% (IIF and EY, 2026).
43. 38% of large enterprises name data quality and readiness as the top hurdle to scaling agentic AI, just ahead of integration with existing systems at 37% (UiPath, 2026).
44. 46% of sales professionals whose teams use AI agents say data quality issues hurt their sales, and they rank manual errors as the top data problem (Salesforce, 2026).
Fines and findings in financial services#
Banks report data to supervisors every day, and supervisors have started pricing the errors. Most of these cases involve reporting data, where one wrong field repeats across millions of records.
45. The OCC fined Citibank $400 million in October 2020 over long-standing weaknesses in risk management, compliance, data governance and internal controls (OCC, 2020).
46. In July 2024, Citi was fined another $135.6 million, $75 million by the OCC and $60.6 million by the Federal Reserve, for insufficient progress on the 2020 orders. The OCC cited a lack of processes to monitor how data quality problems affected regulatory reporting (OCC, 2024).
47. Only 2 of 31 global systemically important banks were fully compliant with all of the Basel Committee's principles for risk data aggregation and reporting, nearly ten years after the principles were published (Basel Committee on Banking Supervision, 2023).
48. None of the 25 significant euro-area banks in the ECB's 2016 review fully followed those principles. In 2024 the ECB called progress since then generally insufficient (European Central Bank, 2024).
49. An internal audit of Washington Federal's 2016 mortgage data found errors in 40 of 100 files, a 40% error rate. The CFPB traced the errors to a lack of appropriate staff, insufficient training and ineffective quality control (CFPB, 2020).
50. The CFPB ordered Freedom Mortgage to pay $3.95 million in 2024 because its 2020 mortgage data submission had widespread errors across numerous data fields. It was the company's second order over this kind of data (CFPB, 2024).
51. Goldman Sachs paid a $6 million SEC penalty in 2023 after more than 22,000 deficient blue sheet submissions over about ten years. Errors of 43 types left missing or inaccurate trade data for at least 163 million transactions (SEC, 2023).
52. Citadel Securities didn't report a 0 in one field for fully canceled orders, which made 31.2 billion of its order events inaccurate in the Consolidated Audit Trail. They were part of about 42.2 billion misreported events from 2020 to 2022, and FINRA fined the firm $1 million (FINRA, 2024).
When bad data made headlines#
Most bad data costs a little every day. Now and then a single wrong value costs a lot at once.
53. Unity Software told the SEC in May 2022 that problems with its Operate ad and monetization products, including the consequences of ingesting bad data from a large customer, would hurt its 2022 business by about $110 million (Unity quarterly report, 2022).
54. A Citigroup trader in London meant to sell $58 million of stocks in May 2022, made an input error, and created a basket worth $444 billion. Controls stopped most of it, but $1.4 billion of stocks were sold before the order was canceled. UK regulators fined the firm £61.6 million (FCA, 2024).
55. Samsung Securities put 2.81 billion of its own shares into 2,018 employees' accounts in April 2018, 1,000 shares for each share held, when it meant to pay a cash dividend of 1,000 won a share. Sixteen employees sold 5.01 million of them, and the stock fell as much as 11.7% that morning (Financial Services Commission, Korea, 2018).
56. A coding issue in an Equifax server changed how some credit scores were calculated for three weeks in 2022, and fewer than 300,000 consumers saw their scores shift by 25 points or more (Equifax, 2022).
57. Public Health England left 15,841 positive COVID-19 cases out of its daily figures between September 25 and October 2, 2020, after a technical issue in the process that loaded lab results into its dashboards (Public Health England, 2020).
58. JPMorgan's task force on its 2012 trading losses found that a new value-at-risk model ran on Excel spreadsheets filled in by copying and pasting from one sheet to another, with data uploaded manually without enough quality control. The errors understated risk, and the report doesn't blame them for the losses themselves (JPMorgan Chase, 2013).
59. Fannie Mae corrected its third-quarter 2003 results after finding computational errors made while applying a new accounting standard. Stockholders' equity went from $16.39 billion as first reported to $17.52 billion, a difference of about $1.1 billion (Fannie Mae filing with the SEC, 2003).
60. NASA lost the Mars Climate Orbiter in 1999 because one ground software file gave thruster data in pound-seconds when the navigation software expected newton-seconds, so each value was off by a factor of 4.45 (NASA Mishap Investigation Board, 1999).
61. Meaning to pay about $7.8 million of interest for Revlon, Citibank wired almost $900 million of its own money to Revlon's lenders in August 2020. Three employees signed off, each believing that setting one field to an internal account would keep the principal from going out. The system needed three fields set (US District Court, Southern District of New York, 2021).
62. A clerical error left mismatched bids in a TransAlta spreadsheet, and the company submitted them in a 2003 auction for New York transmission contracts. It expected the mistake to cost $24 million before tax (TransAlta filing with the SEC, 2003).
Bad data by industry#
Healthcare
63. A duplicate patient record adds an average of $1,950 to an inpatient stay and more than $800 to an emergency department visit, according to 1,392 users of patient-matching software at US hospitals and health systems (Black Book Market Research, 2018).
64. Respondents blamed inaccurate patient identification or information for 33% of denied claims, which they said cost the average hospital $1.5 million in 2017 (Black Book Market Research, 2018).
65. Only 22% of health information professionals had brought their organization's duplicate record rate to 1% or less, and 29% didn't know their rate (AHIMA, 2020).
66. 50% of US healthcare providers name missing or inaccurate data as a top reason claim denials are rising, up from 46% a year earlier, in a survey of 250 people who manage billing and claims (Experian Health, 2025).
67. Medicare fee-for-service had an improper payment rate of 7.66% in fiscal 2024, or $31.70 billion, and 6.55%, or $28.83 billion, in fiscal 2025 (CMS, 2025). CMS says most improper payments across its programs happen when reviewers can't tell whether a payment was proper because the documentation was insufficient (CMS, 2024).
Mortgage and lending
68. The critical defect rate in post-closing mortgage quality control rose to 1.71% in the first quarter of 2026, from 1.38% the quarter before (ACES Quality Management, 2026).
69. Legal, regulatory and compliance problems made up 26.02% of critical mortgage defects in that quarter, just ahead of income and employment problems at 20.07% (ACES Quality Management, 2026). For how lenders check income documents, see mortgage income verification.
Insurance
70. 78% of life insurers say data readiness is the biggest challenge in getting value from AI (LIMRA and Equisoft, 2025).
71. Commercial underwriters spend 17% of their week on manual admin, re-keying and moving between systems, and say about 8% would be reasonable, in a survey of 350 underwriting leaders and underwriters in the US and UK (hyperexponential and Coleman Parkes, 2026). More on reading submissions is on our insurance automation page.
Retail and supply chain
72. 65% of nearly 370,000 inventory records across 37 stores of one retailer were inaccurate when checked against what was on the shelves (DeHoratius and Raman, Management Science, 2008).
73. 69% of orders shipped from brands to retailers contained data errors when their barcode records were checked against RFID tag reads, in a one-year study of more than 1 million items across 5 retailers and 8 brand owners (Auburn University RFID Lab and GS1 US, 2018).
Accounts payable and finance
74. 18.4% of invoices hit an exception on average, and processing one costs AP teams $9.84 and takes 8.2 days (Ardent Partners, 2026).
75. High exception rates tie with slow approvals as the top accounts payable challenge in 2026, each cited by 48% of AP leaders (Ardent Partners, 2026).
76. The median organization sends 92.0% of its customer invoices out error-free the first time, so about 1 in 12 needs rework, across 2,342 organizations (APQC).
77. 18% of accountants make financial errors at least daily, a third make several a week, and 59% make several a month, in Gartner's survey of 497 people in controllership. The errors tracked with low capacity (Gartner, 2024). For invoice-level checks, see accounts payable automation and our accounts payable statistics.
Government payments
78. US federal agencies reported about $186 billion in improper payments for fiscal 2025, up about $24 billion from fiscal 2024, across 64 programs at 15 agencies. Agencies count a payment as improper when documentation is too thin to show it was proper (GAO, 2026).
Where bad data starts, and the 1-10-100 rule#
79. In TDWI's 2001 survey of data warehousing professionals, 76% named data entry by employees as a source of their data quality problems, the most common answer, ahead of changes to source systems at 53% (TDWI, 2002).
80. The 1-10-100 rule says it takes $1 to verify a record as it's entered, $10 to cleanse and de-dupe it later and $100 if nothing is done. SiriusDecisions described it as a rule known in data management circles, and the idea is usually credited to a 1992 book on quality costs by Labovitz, Chang and Rosansky. Treat it as a rule of thumb about when to catch errors, not a measured cost (SiriusDecisions).
81. People keying data once got 0.95% of values wrong in a controlled study where each keyed 1,260 values, against 0.03% with double entry (Barchard and Pace, 2011).
82. Copying data from medical records into research databases had a pooled error rate of 6.57% across 93 studies, against 0.29% for single data entry (Garza et al., 2025).
The calculator below prices only the errors someone finds and fixes, which is the cheap end of the 1-10-100 rule. Put in your own volume. Full error rates by entry method are in our data entry error statistics.
What manual keying errors cost you
Put in your own volume. The error-rate and wage defaults come from published research; change them to match your team.
- Errors keyed per month
- Hours spent fixing them per month
- Cost of fixing them per month
- Cost of fixing them per year
How it's worked out
- Errors are fields keyed times the share keyed wrong. Hours are errors times the minutes to fix one.
- Cost is hours times the hourly cost. Add benefits and overhead to the wage for a fuller figure.
- It counts only errors someone finds and fixes. Errors that reach payments, reports or decisions cost far more, which is the point of the 1-10-100 rule.
Source: Barchard & Pace, Preventing human error: the impact of data entry methods on data accuracy and statistical results, Computers in Human Behavior (2011); US Bureau of Labor Statistics: OEWS profile, Data Entry Keyers (43-9021), May 2025
Document AI moves the check to the point of capture. Docsumo reads each field of an invoice, bank statement or form with a confidence score and sends low-confidence fields to a person, with the source line highlighted. It reports 99% field-level accuracy on 250+ document types and 95%+ straight-through processing. On the Enterprise plan, cross-document validation checks one document against another, such as an invoice against its purchase order and receipt.
Our take. In lending and AP, the errors that cost the most are often two documents that disagree, like a pay stub and the deposits on a bank statement, or an invoice and the purchase order behind it. A perfect read of each page can still leave that mismatch in place. Judge any fix by whether the whole file agrees with itself.
How to put a number on your own cost of bad data#
Industry averages rarely survive a budget meeting, and your own numbers will. Start with the errors your team already finds. Count how many get fixed in a month, from a review queue, a ticket log or a sample, then multiply by the minutes each fix takes and a loaded hourly cost. That's the visible part, the $10 tier of the 1-10-100 rule.
Next, list what those errors did downstream in the same month. Payments reversed or sent twice, loan files sent back to underwriting, claims denied and resubmitted. These take longer to price, and they're the costs leadership notices.
Then estimate what nobody caught. Pull 100 recent records, check each against its source document, and apply the error rate you find to your monthly volume. Run the same audit a quarter later, after any change, so you measure the effect instead of assuming it.
How we picked and checked these figures#
We checked every figure on the publisher's own page or filing on September 27, 2026, and used the Internet Archive where a page has moved or blocks automated reading. Most figures are from 2024 to 2026. Older ones are here because people still quote them, and each carries its year and what it measured.
Many sources sell data products. Monte Carlo, Fivetran, Informatica, Qlik, Salesforce, Validity, Experian, ZeroBounce and Precisely sell data software or services. UiPath, Equisoft and hyperexponential sell automation or insurance software, Ardent Partners' AP research is sponsored by AP software vendors, and Black Book and ACES serve the industries they survey. We kept their figures because they're the newest data on these questions, and each figure names its source. Many are respondents' own estimates rather than measured losses, which matters most for the revenue figures.
You're welcome to use any figure here. Link to the original source, and to this page if it saved you the search.
If your bad data starts in documents#
Every survey above counts the damage after bad data is already in a system. If your team keys values from invoices, statements or claim forms, measure how many fields a person corrects today, then test automation against that number on your own files. Book a demo with a few real documents, or start a free trial.
Frequently asked questions#
How much does bad data cost a business?
Gartner's most quoted figure is $12.9 million a year for the average organization, from a 2020 survey of data quality vendors' reference customers. Thomas Redman estimates 15% to 25% of revenue for most companies. In Monte Carlo's 2023 survey, data professionals said data quality issues affected 31% of revenue on average.
Where does the $3.1 trillion cost of bad data come from?
The earliest source we found is a September 2011 article in SOA World Magazine by Hollis Tibbetts, who extrapolated it from a healthcare estimate and called it his own estimate. An IBM infographic repeated the number by 2013 without a method, and it became known as IBM's after Thomas Redman cited IBM in Harvard Business Review in 2016.
What is the 1-10-100 rule in data quality?
It's a rule of thumb that an error costs $1 to prevent at entry, $10 to find and fix later and $100 if nobody fixes it. It comes from a 1992 book on quality costs by Labovitz, Chang and Rosansky, and SiriusDecisions later applied it to customer records. It describes when to catch errors rather than a measured cost.
How does bad data affect AI projects?
Gartner predicts that through 2026 organizations will abandon 60% of AI projects that aren't supported by AI-ready data, and 63% of organizations lack the right data practices for AI or aren't sure they have them. In a 2024 Fivetran survey, organizations estimated they lose 6% of revenue to AI models built on inaccurate or low-quality data.
How do you calculate the cost of poor data quality?
Start with the errors your team already fixes. Multiply the number fixed each month by the time each fix takes and a loaded hourly cost, then add downstream costs such as reversed payments or denied claims. A sample audit of about 100 records against their source documents estimates the errors nobody caught. For error rates by entry method, see our data entry error statistics.
Sources
- Gartner, Magic Quadrant for Data Quality Solutions (Jul 27, 2020)
- SOA World Magazine (SYS-CON Media), Hollis Tibbetts, $3 Trillion Problem: Three Best Practices for Today's Dirty Data Pandemic (Internet Archive copy) (Sep 10, 2011)
- IBM Big Data & Analytics Hub, The Four V's of Big Data (infographic, Internet Archive copy) (Aug 8, 2013)
- Gartner, How to Stop Data Quality Undermining Your Business (Smarter with Gartner, Susan Moore) (Jan 18, 2018)
- Gartner, Findings From Primary Research Study: Organizations Perceive Significant Cost Impact From Data Quality Issues (G00170011) (Aug 14, 2009)
- MIT Sloan Management Review (Thomas C. Redman), Seizing Opportunity in Data Quality (Nov 27, 2017)
- Forrester, Data Snapshot: Millions Lost in 2023 Due To Poor Data Quality, Potential For Billions To Be Lost With AI Without Intervention (Jul 31, 2024)
- TDWI (The Data Warehousing Institute), Wayne W. Eckerson, Data Quality and the Bottom Line: Achieving Business Success through a Commitment to High Quality Data (TDWI Report Series) (Feb 2002)
- Business Horizons (Elsevier); Tadhg Nagle, Tom Redman, David Sammon, Assessing data quality: A managerial call to action (May 2020)
- CrowdFlower 2016 Data Science Report (usually via Forbes, 'data preparation accounts for about 80% of the work of data scientists') (Mar 2016)
- Harvard Business Review (Thomas C. Redman), Bad Data Costs the U.S. $3 Trillion Per Year (Sep 22, 2016)
- Gartner, Measuring the Business Value of Data Quality (Oct 10, 2011)
- MarketingSherpa (via HubSpot's Database Decay Simulation); the 30% variant is often credited to Gartner or called an 'industry benchmark' (unknown (in circulation by 2011; HubSpot page from 2013))
- Gartner (in most round-ups); sometimes ZoomInfo or Salesforce (2009 (study update; original 2002), described online 2015-02-13)
- Experian Data Quality / Experian QAS (2014) (Jan 22, 2014)
- Experian Data Quality, The data quality benchmark report (2015, Internet Archive copy) (Jan 2015)
- DiscoverOrg (now ZoomInfo), often with no date (Jul 19, 2011)
- KPMG / Forbes Insights, 2016 Global CEO Outlook (often undated) (Jun 30, 2016)
- Gartner, How to Create a Business Case for Data Quality Improvement (Jan 9, 2017)
- Gartner, The State of Data Quality: Current Practices and Evolving Trends (G00255625) and earlier/later Gartner notes (Dec 11, 2013)
- Monte Carlo, The Annual State of Data Quality Survey (May 2, 2023)
- Experian, Highlights from Our 2021 Global Data Management Research (Feb 25, 2021)
- Validity, The State of CRM Data Management in 2026 (Aug 25, 2026)
- Harvard Business Review (Tadhg Nagle, Thomas C. Redman, David Sammon), Only 3% of Companies’ Data Meets Basic Quality Standards (Sep 11, 2017)
- SiriusDecisions (acquired by Forrester in 2019), The Impact of Bad Data on Demand Creation (Nov 2008)
- Salesforce, State of Data and Analytics, 2nd Edition (2025)
- dbt Labs, The 2024 State of Analytics Engineering Report (2024)
- dbt Labs, The 2025 State of Analytics Engineering Report (Apr 2025)
- Drexel LeBow and Precisely, 2025 Outlook: Data Integrity Trends and Insights (Sep 2024)
- Qlik, Data Quality is Not Being Prioritized on AI Projects, a Trend that 96% of U.S. Data Professionals Say Could Lead to Widespread Crises (Mar 12, 2025)
- ZeroBounce, The Email List Decay Report for 2026 (Feb 26, 2026)
- Fivetran, Organizations Bullish on AI Adoption Despite Losing an Average of $406M Each Year Due to Underperforming AI Models (Mar 20, 2024)
- Anaconda, Anaconda Releases 2020 State of Data Science Survey Results (press release) / The State of Data Science 2020: Moving from Hype Toward Maturity (Jun 30, 2020)
- Anaconda, 2022 State of Data Science Report: Paving the Way for Innovation (Sep 2022)
- Monte Carlo, Data Quality Survey 2024 (2024 State of Reliable AI Survey) (Jun 4, 2024)
- Drexel LeBow and Precisely, 2026 State of Data Integrity and AI Readiness (Jan 2026)
- Salesforce, State of Sales, 6th Edition (Jul 25, 2024)
- Salesforce, Trust in Business Data Leaders Survey (May 2025)
- Gartner, Lack of AI-Ready Data Puts AI Projects at Risk (Feb 26, 2025)
- Gartner, Gartner Says Organizations With Successful AI Initiatives Invest Up to Four Times More in Data and Analytics Foundations (Apr 16, 2026)
- Deloitte, The State of Generative AI in the Enterprise: Now decides Next (Q3 report) (Sep 2024)
- IBM Institute for Business Value, The 2025 CDO Study: The AI multiplier effect (Nov 13, 2025)
- Informatica, New Global CDO Report Reveals Data Governance and AI Literacy as Key Accelerators in AI Adoption (CDO Insights 2026) (Jan 27, 2026)
- Institute of International Finance (IIF) and EY, IIF-EY Global Annual Survey Report on AI Use in Financial Services (2026) (Sep 17, 2026)
- UiPath, Stuck in Agentic AI Pilot Purgatory? UiPath Survey Points to Orchestration as Key to Scaling Enterprise Deployments (Sep 9, 2026)
- Salesforce, State of Sales, 7th Edition (Feb 3, 2026)
- Office of the Comptroller of the Currency (OCC), OCC Assesses $400 Million Civil Money Penalty Against Citibank (NR 2020-132) (Oct 7, 2020)
- Office of the Comptroller of the Currency (OCC); Federal Reserve Board, OCC Amends Enforcement Action Against Citibank, Assesses $75 Million Civil Money Penalty (NR 2024-76); Federal Reserve press release, July 10, 2024 (Jul 10, 2024)
- Basel Committee on Banking Supervision (BIS), Progress in adopting the Principles for effective risk data aggregation and risk reporting (Nov 28, 2023)
- European Central Bank (ECB Banking Supervision), Guide on effective risk data aggregation and risk reporting (May 2024)
- Consumer Financial Protection Bureau (CFPB), Consent Order, In the Matter of Washington Federal Bank, N.A. (2020-BCFP-0019) (Oct 27, 2020)
- Consumer Financial Protection Bureau (CFPB), CFPB Takes Action Against Repeat Offender Freedom Mortgage Corporation for Violating Law Enforcement Order and for Housing Data Errors (Jun 18, 2024)
- U.S. Securities and Exchange Commission (SEC), Goldman to Pay SEC $6 Million in Penalties for Providing Deficient Blue Sheet Data (Sep 22, 2023)
- Financial Industry Regulatory Authority (FINRA), Letter of Acceptance, Waiver, and Consent No. 2020068778601 (Citadel Securities LLC) (Oct 2024)
- Unity Software Inc. (Form 10-Q, SEC), Unity Software Inc. Form 10-Q for the quarter ended March 31, 2022 (May 10, 2022)
- Financial Conduct Authority (FCA); Prudential Regulation Authority (PRA), Bank of England, FCA fines CGML £27,766,200 for failures in its trading systems and controls; PRA fines Citigroup Global Markets Limited (CGML) £33,880,000 for failures in its trading systems and controls (May 22, 2024)
- Financial Services Commission (FSC), South Korea, 보도참고: 삼성증권의 배당사고 관련 검사결과에 대한 조치 (Reference Press Release: Measures on the Inspection Findings Related to Samsung Securities' Dividend Accident) (Jul 26, 2018)
- Equifax Inc., Equifax Statement on Recent Coding Issue (Aug 2, 2022)
- Public Health England / GOV.UK, PHE statement on delayed reporting of COVID-19 cases (Oct 4, 2020)
- JPMorgan Chase & Co. (Management Task Force), Report of JPMorgan Chase & Co. Management Task Force Regarding 2012 CIO Losses (Jan 16, 2013)
- Federal National Mortgage Association (Fannie Mae) / SEC EDGAR filings, Form 8-K/A (filed October 29, 2003) and original Form 8-K/Exhibit 99.1 earnings release (filed October 16, 2003) (Oct 29, 2003)
- NASA (Mars Climate Orbiter Mishap Investigation Board), Mars Climate Orbiter Mishap Investigation Board Phase I Report (Nov 10, 1999)
- U.S. District Court for the Southern District of New York (Judge Jesse M. Furman), In re Citibank August 11, 2020 Wire Transfers, Opinion and Order, No. 1:20-cv-06539-JMF (Feb 16, 2021)
- TransAlta Corporation, TransAlta Corporation press release / Form 6-K, “TransAlta announces one-time charge related to New York ISO transmission congestion contracts” (Jun 3, 2003)
- Black Book Market Research, Improving Provider Interoperability Congruently Increasing Patient Record Error Rates, Black Book Survey (Apr 10, 2018)
- AHIMA (American Health Information Management Association), A Realistic Approach to Achieving a 1% Duplicate Record Error Rate (Jul 2020)
- Experian Health, Experian Health's 3rd Annual State of Claims Survey Finds Denials Still on the Rise Amid Escalating Challenges (Sep 22, 2025)
- Centers for Medicare & Medicaid Services (CMS), Improper Payment Rates and Additional Data (CERT) (Nov 2025)
- Centers for Medicare & Medicaid Services (CMS), Fiscal Year 2024 Improper Payments Fact Sheet (Nov 15, 2024)
- ACES Quality Management, Q1 2026 ACES Mortgage QC Industry Trends Report (Aug 18, 2026)
- LIMRA and Equisoft, Assessing Data Readiness for AI in the Life Insurance Industry (Jan 15, 2025)
- hyperexponential, hx Underwriting Edge Report 2026 (Sep 10, 2026)
- INFORMS (Management Science), Inventory Record Inaccuracy: An Empirical Analysis (Apr 1, 2008)
- Auburn University RFID Lab and GS1 US, New Study from the Auburn University RFID Lab and GS1 US Confirms RFID Enables Nearly 100% Order Accuracy for Retail (Oct 10, 2018)
- Ardent Partners, State of ePayables (Part Nine): AP Benchmarks and Best-in-Class Performance (Jan 22, 2026)
- Ardent Partners, The State of AP 2026 Pt. 3: Challenges in 2026: Familiar Friction, Rising Stakes (Aug 11, 2026)
- APQC, Percentage of invoices processed error free the first time (Sep 2026)
- Gartner, Gartner Survey Shows That a Third of Accountants Make Several Financial Errors Per Week due to Capacity Constraints (Feb 21, 2024)
- US Government Accountability Office (GAO), Payment Integrity: Agencies' Estimated Improper Payments Increased to $186 Billion in Fiscal Year 2025 (Apr 27, 2026)
- Barchard and Pace, Preventing human error: The impact of data entry methods on data accuracy and statistical results. Computers in Human Behavior 27(5) (2011)
- Garza et al., Error rates of data processing methods in clinical research: a systematic review and meta-analysis. Int J Med Inform 195 (2025)
First published .