Data and Information

    Creating the data basis for AI

    By Redaktion techport.ai, IT-Beratung · Last updated on

    AI initiatives in the mid-market rarely fail on technology. They fail because an assistant draws on a body of material consisting of five versions of the same price list, a manual from 2019 and the notes of a colleague who left. The result is a plausibly worded wrong answer, and after two of those nobody uses the tool any more.

    Preparing the data basis is the least spectacular and most important part of any AI initiative. It determines the quality of the results more than the choice of model.

    How you notice it

    • A tested assistant produces answers based on outdated documents.
    • Nobody can say which source is the valid one for a given topic.
    • It is unclear whether an assistant shows information the asking person should not see.
    • Results cannot be traced to a source and therefore cannot be checked.

    Why this happens

    Company knowledge has accumulated in many places over the years and nobody ever decided which version applies. For people that is manageable, because they know the context and recognise outdated documents. A system has no such context. It treats every source equally, weighting by similarity rather than by validity, and therefore gives wrong answers precisely when the body of material is contradictory.

    How we go about it

    1. Use case first, sources second. We determine the concrete case, for example answers in customer service or support for quotations, and derive from it which sources are actually needed. Not the entire body of material belongs in a knowledge base.
    2. Clean and label the sources. We remove outdated versions, determine the valid document per topic and add information on status, ownership and validity. Those attributes later form the basis for weighting and verifiability.
    3. Mirror the permissions. We ensure that the access rights of the source system also apply in the knowledge base. An assistant that exposes HR or costing data to everyone is a data protection incident, not progress.
    4. Secure currency and verifiability. We set up refresh so that changes in the sources arrive, and make sure answers cite their sources so users can check what a statement rests on.

    What you gain

    • Answers that are correct and whose origin is traceable.
    • Trust from users that survives the first weeks.
    • A knowledge base that delivers value independently of AI.

    From our projects

    We recommend starting with a clearly bounded area rather than with the whole of company knowledge. An assistant for a single subject area with maintained sources produces usable results within a short time, whereas trying to include everything at once regularly ends in a body of material nobody can curate. The second reliable finding concerns permissions: almost every network drive contains areas readable by more people than intended. In daily work that goes unnoticed because nobody searches there. An assistant searches everywhere it has access to and exposes those gaps within days.

    Häufige Fragen

    Will our data be used to train the model?

    That depends on the contract and belongs settled before use. In business offerings from reputable providers, use of your input for training is generally excluded, while with free accounts it frequently is not. Have that point confirmed in writing and check it again whenever the contract changes.

    Do we have to tidy up all documents before we start?

    No, that would prevent any initiative. What gets tidied is the area the first use case needs. That scope is usually manageable, and the experience from the first area makes the following ones faster. What matters is that the boundary of the knowledge base is set technically and not merely intended.

    Let us talk about Creating the data basis for AI

    In a thirty minute first call we work out where your biggest lever sits and whether we are the right people for it.

    Further reading

    Back to the field Data and Information