AI Data Readiness: Why Your Data Foundation Comes First

Learn why AI projects fail without clean, connected operational data and how to assess data readiness before choosing tools or platforms.

Connected operational data foundation prepared for enterprise AI

Every AI project starts with the same question, and it’s almost always the wrong one: “Which tool should we buy?”

Rolan Marco Garcia, CEO of Embiggen X, has sat through this conversation with manufacturing, construction, and industrial leaders across Switzerland and Asia more times than he can count. The pattern is always the same. A CEO wants to talk about AI agents, which large language model to use, which platform to buy. Nobody wants to talk about the unsexy part first: is your data even usable?

“If your data is garbage, garbage in, garbage out,” Garcia says. “If your data infrastructure is not good, if your data is unreliable, AI is not going to be very useful for you. It’s not going to get the efficiency gains or the topline gains that you want.”

That’s not a caveat. According to the research below, it’s the single best-documented reason AI projects fail.

The data behind the problem

Three independent studies, using three different methods, land on the same conclusion.

MIT’s NANDA initiative published The GenAI Divide: State of AI in Business 2025 in August 2025, based on 150 interviews with business leaders, a survey of 350 employees, and analysis of 300 publicly documented generative AI deployments. Its headline finding: despite an estimated $30–40 billion in enterprise generative AI investment, 95% of pilots failed to produce a measurable impact on revenue or P&L. The report attributes this primarily to what it calls a “learning gap” — an organizational failure to integrate AI into existing workflows and data — rather than to model quality.

RAND Corporation, in a 2024 research report based on interviews with 65 data scientists and engineers, found that more than 80% of AI projects fail — roughly twice the failure rate of non-AI IT projects. RAND’s researchers identified five recurring root causes: leadership-driven failures, data quality limitations, chasing technology bottom-up instead of starting from the business problem, underinvestment in deployment infrastructure, and applying AI to problems beyond its current capabilities. Data quality was one of the most consistently cited issues across the practitioner interviews.

Gartner, surveying 1,203 data management leaders in July 2024, found that 63% of organizations either lack the right data management practices for AI or are unsure whether they have them. Gartner’s resulting prediction, published in February 2025: through 2026, organizations will abandon 60% of AI projects due to a lack of AI-ready data. A separate Gartner prediction from mid-2024 put the number even higher for generative AI specifically — at least 30% of GenAI projects abandoned after proof-of-concept by the end of 2025, citing poor data quality, unclear business value, and escalating costs as the leading causes.

ResearchMethodologyKey finding
MIT NANDA, State of AI in Business 2025150 executive interviews, 350 employee surveys, 300 deployment case studies95% of enterprise GenAI pilots produce no measurable P&L impact
RAND Corporation, RRA2680-1 (2024)65 practitioner interviews80%+ of AI projects fail — 2x the rate of non-AI IT projects; data quality is a top-five recurring cause
Gartner (Feb 2025)Survey of 1,203 data management leaders63% lack AI-ready data practices; 60% of AI projects predicted to be abandoned through 2026 for this reason

Three different research teams, three different methods, one shared conclusion: the technology usually isn’t what breaks AI projects. The data underneath it is.

What is AI data readiness?

AI data readiness is how prepared your organization’s data actually is to be used by an AI system — structured consistently, centralized in a reliable source of truth, and accurate enough that an AI tool reading it produces correct outputs rather than confidently wrong ones.

It’s a different question from data quality in the narrow, technical sense (are individual records accurate, deduplicated, well-formatted). Data readiness is broader and more organizational: does a single reliable version of this data exist anywhere at all, is it structured the same way across departments, and does the AI system have a way to actually reach it? A company can have technically “clean” data in five different disconnected spreadsheets and still be nowhere near AI-ready. Gartner makes the same distinction in its own guidance to CIOs and chief data officers: traditional data management practices are “too slow, too structured, and too rigid” for AI use cases, which is exactly why 63% of the leaders it surveyed said they weren’t confident they had the right practices in place.

The tooling conversation is a distraction

It’s easy to see why leaders default to talking about tools. Tooling feels like progress — you can see a demo, sign a contract, put a launch date on a slide. Data infrastructure, by contrast, is invisible until it breaks something.

But AI doesn’t create capability out of nothing. It amplifies whatever is already running underneath it. A well-built forecasting model pointed at spreadsheets with inconsistent columns, siloed systems, and no shared source of truth won’t produce a fixed process — it will produce a faster, more confident version of the same bad process. This is consistent with what RAND found: AI project failure is rarely a single catastrophic error, and far more often a foundation problem that a good tool simply can’t compensate for. For AI to generate real P&L value, not just impressive demos, the data infrastructure underneath it has to be sound first.

Most companies overestimate their own data maturity

At industry conferences, Garcia has heard CIOs of large companies claim 70% of their data is “fixed” — meaning structured, reliable, ready to use. Maybe. But for mid-sized and smaller companies, the real number is often reversed: as much as 90% of operational data is unstructured, siloed across disconnected systems, or — worse — sitting only in the heads of a few long-tenured employees.

That number tracks with independent industry estimates. IDC and Gartner analyses have separately put unstructured data at roughly 80–90% of all enterprise data, growing an estimated 55–65% per year — nearly three times faster than structured data. In other words, Garcia’s field observation from Swiss and Asian industrial clients lines up with what analyst firms have been documenting at a broader scale.

That’s the uncomfortable version of data readiness: it’s not just a technical audit, it’s an organizational one. If the person who understands your production planning logic left tomorrow, would that knowledge still exist anywhere else?

The real cost of ignoring this

The cost of skipping the data question isn’t abstract. Gartner has estimated that poor data quality costs the average organization $12.9 million per year, based on a survey of 154 companies across 16 data-quality vendor implementations. Separately, MIT Sloan Management Review has cited research putting the cost of bad data at 15–25% of revenue for a typical company — driven by the hidden labor of people cross-checking numbers, correcting errors, and second-guessing reports they don’t fully trust.

Layer an AI initiative on top of that same unreliable foundation, and the cost compounds. You’re not just paying for the AI tool and the data problems separately — you’re paying for a tool that inherits, automates, and scales the underlying data problems, often before anyone notices.

Three questions to ask before you evaluate any AI tool

Before any enterprise AI transformation conversation, Garcia argues the first questions have to be about data, not tooling. This is the same gut-check we walk industrial clients through before any platform demo:

  1. Do you have a single, reliable source of truth? Not “a spreadsheet somewhere with most of it” — one place the whole company would point to if asked where the real numbers live.
  2. Are your systems of record actually connected to each other? ERP, HR platform, procurement, production planning — do they talk to one another, or does someone manually reconcile them every month?
  3. Is the data structured the same way across departments? If two teams both track the same thing but use different columns, formats, or definitions, no AI tool can quietly reconcile that for you.

If the honest answer to any of these is “no” or “sort of,” that’s not a reason to avoid AI — it’s the actual starting point of the project, before any tool gets chosen. It’s also, per Gartner’s five-step guidance to CIOs on building AI-ready data practices, the same starting point analysts are now recommending: align data to the specific use case, define governance requirements, and treat metadata as an active, continuously maintained practice rather than a one-time cleanup.

Data readiness vs. data quality — and why most AI ROI reports go wrong

Industry reports consistently show a gap between AI adoption and AI ROI — companies with high adoption rates that still aren’t seeing returns. The tools aren’t usually the problem. The foundation is. An AI system can only be as reliable as the data it’s reading, and if that data is inconsistent, unstructured, or scattered across disconnected systems, no amount of model sophistication will fix it downstream.

This is also where “data quality” and “data readiness” get confused. A data-quality tool can clean up formatting and catch duplicates inside one system. It won’t create a single source of truth across five disconnected ones, and it won’t tell you whether the AI project you’re about to fund is even solving a problem your data can support. Data readiness is the broader, earlier question — data quality is one input into it.

This is why data readiness deserves to be treated as its own project phase — not a footnote before the “real” AI work begins.

How Embiggen X approaches this

This is the reasoning behind Embiggen X’s three-phase method: diagnose the operation before prescribing technology. Before any AI tooling decision, the diagnostic phase maps what data actually exists, where it lives, how reliable it is, and whether departments are working from a shared source of truth. Data Core exists specifically to connect and structure that operational data into a single reliable layer — the foundation that makes every AI investment on top of it actually pay off, rather than just look impressive in a demo.

FAQ

What does “AI data readiness” actually mean?
It means your operational data is structured, reliable, and centralized enough for an AI system to read it accurately — consistent formats across departments, a single source of truth, and minimal reliance on tacit knowledge held by individual employees.

What percentage of AI projects actually fail?
Estimates vary by study and definition of “failure,” but they’re consistently high. MIT’s NANDA initiative found 95% of enterprise generative AI pilots produced no measurable P&L impact (2025). RAND Corporation found over 80% of AI projects fail outright — roughly double the failure rate of non-AI IT projects. Gartner predicts organizations will abandon 60% of AI projects through 2026 specifically due to a lack of AI-ready data.

What’s the difference between data readiness and data quality?
Data quality is a technical measure of whether individual data records are accurate and well-formatted. Data readiness is broader — it asks whether a reliable version of that data exists in one place at all, and whether it’s structured consistently enough across the company for an AI system to use it. You can have “clean” data in five disconnected places and still not be AI-ready.

How much of a typical mid-sized company’s data is actually usable for AI?
Often far less than leadership assumes. Industry estimates from IDC and Gartner put unstructured data at 80–90% of all enterprise data. Mid-sized and smaller companies frequently find that as much as 90% of their operational data is unstructured, siloed, or undocumented entirely.

Should we fix our data before or after choosing an AI tool?
Before. Assessing data maturity — reliability, structure, and a shared source of truth — should be the first step in any AI transformation, ahead of any tooling or platform decision. Gartner’s own guidance to CIOs recommends the same sequence: define what AI-ready data means for your use case first, then build the tooling and pipelines around it.

Sources

  1. MIT NANDA, The GenAI Divide: State of AI in Business 2025 (August 2025). mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf
  2. RAND Corporation, The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed: Avoiding the Anti-Patterns of AI, RRA2680-1 (2024). rand.org/pubs/research_reports/RRA2680-1.html
  3. Gartner, “Lack of AI-Ready Data Puts AI Projects at Risk” (February 26, 2025). gartner.com/en/newsroom/press-releases/2025-02-26-lack-of-ai-ready-data-puts-ai-projects-at-risk
  4. Gartner, “Gartner Predicts 30% of Generative AI Projects Will Be Abandoned After Proof of Concept By End of 2025” (July 29, 2024). gartner.com/en/newsroom/press-releases/2024-07-29-gartner-predicts-30-percent-of-generative-ai-projects-will-be-abandoned-after-proof-of-concept-by-end-of-2025
  5. Gartner, “Data Quality: Why It Matters and How to Achieve It” (cost-of-poor-data-quality estimate, ~$12.9M/year). gartner.com/en/data-analytics/topics/data-quality
  6. MIT Sloan Management Review, “Seizing Opportunity in Data Quality.” sloanreview.mit.edu/article/seizing-opportunity-in-data-quality
  7. IDC / Gartner estimates on unstructured data as a share of enterprise data (80–90%), as aggregated in industry roundups; original estimates are cited within IDC and Gartner unstructured-data research.

Not sure how ready your data actually is? Book a free AI assessment with Embiggen X and find out before you buy anything.

Share this:
Table of Contents
More Insights

Explore more insights in tech world

Insight
Factory worker using a role-aware AI agent to identify metal parts from a smartphone photo.

AI Agents in Manufacturing: Why Every Factory Worker Will Have One

Role-aware AI agents can help factory workers identify parts, resolve faults and capture operational knowledge—but only when data, permissions and workflows are ready.
Insight
Industrial leader ascending a staged transformation roadmap toward an illuminated destination

The Multi-Year Digital Transformation Roadmap Most Industrial Companies Get Backwards

Learn why industrial transformation programs lose budget when implementation precedes diagnosis—and how to build a staged, measurable roadmap.
General News
Embiggen X representatives receiving International Prime Awards Asia 2026

Embiggen X and Its CEO Recognised at the International Prime Awards Asia 2026

Embiggen X and CEO Rolan Marco Garcia received two International Prime Awards Asia 2026 for leadership in enterprise AI transformation.