Cleansing master data and keeping it clean
By Redaktion techport.ai, IT-Beratung · Last updated on
Master data is the quiet cost driver in the mid-market. A customer address that exists three times produces three invoices to the same company, distorts the revenue analysis and sends the reminder to the wrong place. An article without a weight makes shipping cost calculation impossible. Each individual case is small, and in total they create considerable effort that nobody recognises as a data problem because it disguises itself as rework.
The hard part is not the cleansing. The hard part is holding the level after cleansing.
How you notice it
- The same company exists several times, in slightly different spellings.
- Mandatory fields are filled with placeholders because they were in the way at creation.
- Analyses have to be corrected by hand before they can be shown.
- A year after a cleansing exercise the quality is back where it was.
Why this happens
Data quality arises at the point of entry, and there it conflicts with speed. Someone creating a customer under time pressure does not spend three minutes looking for a possible duplicate. If mandatory fields are additionally required that nobody knows at the moment of creation, placeholders appear. Without feedback to the person entering the data, every cleansing remains a one-off exercise whose effect fades within months.
How we go about it
- Measure quality rather than estimate it. We check completeness, duplicates, format compliance and currency per data object and express the result in a few metrics. Only those numbers turn a feeling into something that can be decided about.
- Cleanse where it counts. We cleanse where the effect is greatest, usually on active records and on fields that block processes. Inactive legacy data gets flagged and archived rather than laboriously corrected.
- Regulate creation and maintenance. We simplify creation, remove mandatory fields that cannot be known at the time of entry, introduce a duplicate check and define who may create and who reviews.
- Watch quality permanently. We set up a regular report of the quality metrics that goes to the responsible areas. Feedback to the place where data arises is the only measure that works over time.
What you gain
- Noticeably less rework in sales, purchasing, shipping and finance.
- Analyses usable without correction.
- A data basis on which automation and AI can work at all.
From our projects
Cleansing projects without subsequent rules are wasted time. We therefore measure quality before cleansing, immediately after and one year later, and make that development visible. In companies where feedback to the entering areas was set up, the level achieved stays largely stable. Where it is missing, quality returns to close to the starting value after about a year. The second recurring finding concerns placeholders in mandatory fields: they are almost always a sign that a field is required at the wrong moment, and they can be removed by changing the workflow rather than by admonition.
Häufige Fragen
Should we merge duplicates automatically?
Automatically only in unambiguous cases with strict criteria and traceable logging. Everything else belongs on a review list for people, because a wrongly merged company with different delivery addresses causes more damage than two records. More important than merging afterwards is the check at creation.
How do we persuade departments to enter data more carefully?
Not with appeals but by making the benefit visible and the effort smaller. If creation becomes easier, if the person entering notices that a well maintained field saves them work later, and if quality per area is reported visibly, behaviour changes. Without those three elements no instruction lasts longer than a quarter.
Let us talk about Cleansing master data and keeping it clean
In a thirty minute first call we work out where your biggest lever sits and whether we are the right people for it.
Further reading
Back to the field Data and Information