Back to Blog

    Data Readiness for AI: What Actually Has to Be True

    Most companies are told they need a data platform before they can use AI. For most use cases, four narrower conditions matter far more.

    Erin Moore

    Erin Moore

    Fractional Chief AI Officer

    |September 15, 20264 min read
    Data Readiness for AI: What Actually Has to Be True
    Share:
    Share:

    Data readiness for AI means four things are true about the specific data your use case touches: it exists, someone owns it, it is accurate enough for the decision, and you can get at it without a project. Notice that "we have a data warehouse" is not on the list — plenty of companies with excellent platforms have unusable data, and plenty without one automate successfully.

    The four conditions

    It exists, in a retrievable form. Not "we have that somewhere". If the information lives in someone's inbox or in the memory of a long-serving employee, the first task is capturing it, and that is a process change rather than an AI project.

    Someone owns it. Ownership means a named person who decides what a field means and whether a value is correct. Without an owner, every disagreement about the output becomes an unresolvable argument about the input.

    It is accurate enough for this decision. Precision requirements vary enormously. Routing a support ticket tolerates far more error than calculating a refund. Ask what happens when the data is wrong, and size the accuracy requirement to that consequence rather than to a general standard.

    You can access it without a project. If extracting the data requires a quarter of engineering work, that cost belongs in the business case. Many use cases that look attractive stop looking attractive once integration is priced honestly.

    Why the warehouse question is usually a distraction

    Building a data platform before knowing which decisions it will serve is how companies spend a year preparing and produce nothing. The platform is justified by the volume of downstream uses, and you do not know that volume until you have run a few. The short version: start with the use case and let the platform earn itself.

    The exception is when three or four candidate use cases all need the same joined data. Then the platform work is shared, and building it first is efficient rather than premature.

    Testing readiness cheaply

    Take one candidate use case and try to assemble a week of its data by hand. If a competent person can do it in a morning, the data is accessible. If it takes three days and two interruptions of an engineer, you have found the real cost, and you found it for the price of a morning rather than a quarter.

    Then look at the sample honestly. Count the rows you would have to discard. If a third of the records are unusable, the data quality problem is the project, and an AI audit is a better next step than an automation pilot.

    Where this sits in the bigger picture

    Data is one of four dimensions in the AI readiness assessment — the others being process, people and governance. It is worth knowing that data is the dimension companies most often assume is their constraint, and comparatively rarely is. More often the process is undocumented or disputed, which no amount of data engineering fixes.

    If you would rather work through the sequence yourself, the 90-Day AI Playbook covers the readiness step and what to do with each answer.

    Where this sits in a risk framework

    The inventory work described here is what NIST's AI Risk Management Framework calls the "map" function — establishing what systems and data you actually have before deciding what to do with them. Doing it once serves both the readiness question and the governance one.

    Frequently asked questions

    What does data readiness for AI actually mean? That the data your specific use case needs exists in retrievable form, has a named owner, is accurate enough for the consequence of being wrong, and can be accessed without a separate engineering project.

    Do we need a data warehouse before using AI? Usually not. Build the platform when several use cases need the same joined data, not before — building it first means spending a year preparing for uses you have not validated.

    How accurate does our data need to be? As accurate as the consequence demands. Ticket routing tolerates far more error than billing. Size the requirement to what happens when a value is wrong, rather than to an abstract standard.

    How do we test data readiness quickly? Assemble one week of the data by hand. If that takes a morning, you are ready; if it takes days and repeated engineering help, you have just priced the integration honestly and cheaply.

    Further reading

    Erin Moore

    Written by

    Erin Moore

    Fractional Chief AI Officer

    Army Veteran turned Fractional Chief AI Officer. Founder of AutomateNexus. I help growing businesses implement enterprise-grade AI solutions that deliver ROI in 90 days or less. Author of "The AI Automation Field Manual."

    Ready to Automate Your Business?

    Let's discuss how AI automation can deliver measurable ROI for your organization in 90 days or sooner.