AI Agent Need help?
Let's chat

We Value Your Privacy

We use cookies to enhance your browsing experience, serve personalized content, and analyze our traffic. By clicking "Accept All", you consent to our use of cookies. See our privacy policy. You can manage your preferences by clicking "customize".

Your AI Is Only as Good as Your Data (And Most Companies Aren't Ready)

Your AI Is Only as Good as Your Data (And Most Companies Aren't Ready)

Author Martin Wambui
2026-08-13
3 Views

 

Only about 1 in 10 organizations have a documented data strategy. In 2026, that gap is the single biggest reason AI initiatives stall before they ever reach production.

Here’s an uncomfortable fact. Only about 1 in 10 organizations actually have a documented data strategy. Not a data warehouse. Not a BI dashboard. A strategy, meaning a plan for how data gets collected, governed, and turned into decisions.

That number comes from data management veteran Peter Aiken, and it hasn’t moved much in twenty years. In 2002, 48% of companies had no data quality plan at all. By 2011, that number had only dropped to 30%. Two decades of "data is the new oil" think pieces, and most companies still haven’t written down how they’ll actually manage the stuff.

Now put that next to a newer number. In 2024, 99% of CEOs said they were planning to invest in GenAI. Nearly every company wants the upside of AI. Very few have done the unglamorous work that makes AI possible in the first place.

That gap is the whole story of data strategy right now.

The math nobody wants to do

According to EY’s research, 87% of data science projects never make it into production. Not because the models are bad. Because the data underneath them is siloed, ungoverned, or simply not there when it’s needed.

Gartner has put a sharper number on the AI version of this problem. Through 2026, the firm predicts organizations will abandon 60% of AI projects that aren’t supported by AI-ready data, and a 2025 Gartner survey found that 63% of organizations either don’t have, or aren’t sure they have, the right data management practices to support AI in the first place.

The 2026 data backs this up from a different angle. A Cloudera and Harvard Business Review Analytic Services study published in March 2026 found that only 7% of enterprises say their data is completely ready for AI, even though nearly every one of them is running AI initiatives already. Seventy-three percent admit their organization should be prioritizing AI data quality more than it currently does.

87%  of data science projects never reach production  (EY, 2024)

60%  of AI projects will be abandoned through 2026 without AI-ready data  (Gartner, 2025)

7%  of enterprises call their data completely AI-ready  (Cloudera / HBR, 2026)

So when people say "garbage in, garbage out," they’re not being cute. They’re describing the reason a majority of AI initiatives are on track to quietly die.

What a data strategy actually is (and isn’t)

A lot of executives hear "data strategy" and picture a slide with a cloud icon and some arrows. That’s not it.

DXC Technology’s research defines it more usefully. A data strategy is a common reference for how an organization acquires, stores, secures, manages, and operationalizes its data, paired with a clear target state and the success criteria to know if you’re getting there. It’s not a solution to any one technical problem, and it’s not just a leadership slide deck either. It has to be specific enough that a data engineer and a CFO can both use it to make decisions.

It’s also not static. As the business changes, the strategy has to change with it. Treating it as a "set and forget" document is one of the fastest ways to make it irrelevant within a year.

The four reasons companies actually build one

That research points to four recurring drivers behind why organizations finally commit to a data strategy:

  • Getting business and IT speaking the same language. Without a shared strategy, business teams and technical teams end up solving the same problem twice, differently.
  • Treating data as a shared asset instead of a departmental one. When every team defines “customer” or “revenue” its own way, nothing rolls up cleanly at the enterprise level.
  • Defining what “success” even means. Without agreed metrics, every project gets evaluated by whoever’s loudest in the room.
  • Paying down technology debt. Old systems that limp along because nobody wants to touch them quietly block every new initiative that depends on their data.

Notice that only one of these four is really a technology problem. The rest are organizational.

Three companies that took very different paths

The most interesting part of the data strategy research isn’t the frameworks. It’s watching real companies wrestle with the same tension from opposite directions.

Cisco grew through more than 140 acquisitions over 25 years, which meant 140 different ways of defining the same data. Their fix was to centralize: build a shared data foundation, standardize metrics across finance and field operations, and make data quality a measured, monitored discipline rather than an afterthought. As one of their data leaders put it, people were finally “rallying around the fact that data is critical.”

Intuit went the opposite direction. After building a single centralized enterprise data warehouse in the late 1990s, they found it created a massive project backlog and frustrated business units that used to manage their own data. So by 2008 they’d built a hybrid model instead: a shared foundation maintained centrally, with divisions allowed to build their own data marts on top of it using common tools and naming conventions. Their BI director described it as a balance between centralization and decentralization, enough rigor where it matters and enough flexibility where it doesn’t.

Harley-Davidson took a third route and prioritized speed. Rather than routing every data request through one slow, one-size-fits-all process, their information strategy leaned into self-service BI, analytic sandboxes, and a network of business-side “superusers” who could handle their own ad hoc analysis, freeing the core team to focus on harder problems.

Three different companies, three different starting points, three different answers. That’s the real lesson. There’s no universal template, only the right balance of control and flexibility for where your organization is today.

What "AI-ready" data actually requires

If your goal is specifically to get ready for AI, not just BI, the bar is higher. EY’s Data 4.0 research breaks AI readiness into seven pillars, but they boil down to five things that matter most:

  1. Metadata that’s actually current. Not a data dictionary from three years ago. Active, accurate context about what your data means and where it lives.
  2. Lineage and provenance. The ability to trace a piece of data back to its origin and every transformation along the way. Without this, you can’t debug a bad AI output, let alone trust a good one.
  3. Fit-for-purpose structuring. Different teams need different views of the same underlying data. AI-ready data is organized so the right slice is easy to find.
  4. Security and governance built in from the start, not added later. Industry surveys consistently show that data back-end readiness, not model sophistication, is the real bottleneck for most AI programs.
  5. A real compliance posture. Whether it’s a state privacy law, an industry regulation, or a client contract, data collected without a clear, lawful basis becomes a liability the moment you try to use it for AI.

Miss any of these and you get the pattern EY describes: AI initiatives that fail to deliver ROI, model inaccuracies, and, in the worst cases, regulatory exposure that turns a promising pilot into a legal problem.

The confidence gap hasn’t closed either

Here’s the part that should worry executives most. A 2026 survey of data and analytics leaders, run by Precisely with Drexel University’s LeBow College of Business, found that 87% of them believe they have the infrastructure to support AI. In the same survey, 42% named infrastructure as a top AI challenge. Those two numbers are describing the same organizations.

87%  say they have AI-ready infrastructure  (Precisely / Drexel LeBow, 2026)

42%  name infrastructure as a top AI challenge — same survey  (Precisely / Drexel LeBow, 2026)

That’s not a new phenomenon. It’s the same gap Wayne Eckerson documented back in 2011, when he found that a majority of BI directors thought their data quality was fine, right up until they were asked a harder question and admitted otherwise. Fifteen years and an entire generative AI boom later, the pattern hasn’t changed. Confidence in data quality still outpaces the reality of it, and that gap is exactly where AI projects go to die.

It’s not all bad news. The EDM Association’s 2026 Global Data Management Benchmark found that roughly 31% of organizations have reached advanced data strategy capability, real progress from the "1 in 10" figure Eckerson cited in 2011. The floor has risen. It just hasn’t risen fast enough to keep pace with how quickly companies are trying to bolt AI onto what they’ve got.

The part nobody puts on a slide

Every one of these studies, whether written in 2011 or 2024, circles back to the same uncomfortable truth. The hard part of data strategy was never the technology.

It’s convincing a leader that data quality is worth funding before something breaks. It’s getting a data steward, a business person rather than an IT person, to actually own a dataset instead of just attending meetings about it. It’s having the discipline to enforce a governance policy even when enforcing it slows someone down this quarter.

"If no one has ever gotten fired for misusing corporate data or violating established data policies, you don’t have governance. You have a document."

Where to actually start

If you’re staring down a blank page, the research is fairly consistent on sequencing:

  • Document your current state honestly before designing the future one.
  • Get executive sponsorship before you get a tool.
  • Pick one cross-functional, data-intensive project as your proving ground, not a company-wide rollout.
  • Define your data governance model and your reference architecture before you pick platforms, not after.
  • Build in a change management plan from day one. Resistance isn’t a sign you’re doing it wrong. It’s normal, and ignoring it is what actually kills these programs.

None of that requires a seven-figure platform purchase. It requires a decision that data is worth treating like the asset every company claims it is, and then actually following through.

The organizations building GenAI pilots on top of siloed, unstandardized, poorly governed data are, statistically, going to be part of Gartner’s 60% abandonment number. The ones who spent the boring months before that fixing their data foundation are the 7% Cloudera and Harvard Business Review found already calling their data AI-ready. That gap between the two groups is the entire data strategy conversation, and it’s only getting wider.

 

Sources

  1. EY and EDM Council, Data 4.0: Making Your Data AI-Ready (September 2024)
  2. DXC Technology, Defining a Data Strategy: An Essential Component of Your Transformation Journey (2021)
  3. Wayne Eckerson, Creating an Enterprise Data Strategy: Managing Data as a Corporate Asset, TechTarget (June 2011)
  4. Gartner, How to Evaluate AI Data Readiness, by Mark Beyer, Ehtisham Zaidi, and Roxane Edjlali (January 2025); Gartner press release, “Lack of AI-Ready Data Puts AI Projects at Risk” (February 2025)
  5. Cloudera and Harvard Business Review Analytic Services, The Data Readiness Index 2026: Understanding the Foundations for Successful AI (March 2026)
  6. Precisely and the Center for Applied AI and Business Analytics, Drexel University LeBow College of Business, 2026 State of Data Integrity and AI Readiness (January 2026)
  7. EDM Association, 2026 Global Data Management Benchmark Report (May 2026)