There are generally two types of conversation when someone contacts me these days. The first is an initial pre-emptive enquiry such as how we go about our business, typical milestones, likely costs and team size. The second is from someone facing an impending disaster and very keen to know how quickly we can start.
Why the haste? Usually because the immutable truth of a data migration is about to come to pass: the data quality isn’t right and things aren’t going to work.
The reality of a new system implementation or the bringing together of datasets due a merger, involves so many complex moving parts that there’s usually been an assumption that the data will be okay, around variations of “things work now, so why wouldn’t they work post-migration?”, “we’ve bigger fish to fry right now” and “let’s migrate what we have and we’ll fix it afterwards”.
I talked in the last edition of Housing Technology about the interoperability of systems (or lack thereof) so I won’t re-cover that ground, but why do we instinctively want to assume that the data we have is fit for purpose? After all, if it was easy, everyone would be doing it.
Messy and complicated
For one thing, data quality is messy. It’s long, repetitive, tedious and complicated. The quietest room in the world is the one in which a CEO has just asked, “so who’s actually responsible for our data?”.
Improving the quality of data implies first understanding what’s important among all the data that is held. Not all of it is used so we have to separate the necessary data from the nice-to-have data.
We then have to understand how to test the quality. The Data Management Association (dama.org) defines six data-quality dimensions: accuracy, completeness, consistency, timeliness, validity and uniqueness. To test the quality, we must now define and build the rules for each field, considering each dimension.
In theory, it’s quite easy to test data quality with these six dimensions. Does this value make sense in this field; for example, do I have a valid postcode in a ‘Postcode’ field. However, to truly test the quality of a field we need its context.
I may have a postcode but does it match the ‘City or Town’ data provided. More than that, does it match the whole address? Going even further, does it belong to an asset that I have maintenance records or a tenancy agreement for?
To understand that context for every field in every table that we need to do business isn’t the work of a moment and it’s not usually within the scope of one person’s knowledge either. That means we need to involve people across the organisation and that takes time.
It’s this complexity that experience helps to overcome.
Quality is different for migrations
When we look at data quality through a migration lens, there are three categories we need to think about.
First of all, the ‘system mandatory’ data – in other words, the data fields the new system needs to be populated to work.
This data is usually specified by your systems integrator. They’ll look at their data model and the processes they’re implementing and provide a list of the mandatory data items. There’s not much point in moving forward without this data because the system won’t work as expected.
You’re likely to have most of these fields; after all, we’re working with housing systems and housing data so the two will naturally align closely, and if we’re missing data for a field in this category entirely, we can default it with no harm done.
The second camp isn’t so easy to identify but just as important. We call these fields ‘business mandatory’; data that the new system doesn’t necessarily need but your business processes definitely do need.
We can only discover these by working closely with subject-matter experts (SMEs). We may need to take in spreadsheets or other ‘off-system’ data and we may come across processes that are run regularly but aren’t actually understood by anyone currently in the business. It’s common to hear something along the lines of, “I don’t know what it does, but I know we need it.”
This category of data item is probably the hardest to fully quality assess in a migration context. We’ll often only find out that fields in this category are missing when user testing begins which, of course, invites delay.
The third and final category is, well, everything else. This group is usually where we would spend the least time, but strongly encourage our clients to have a very close look to see why the data is being held if it’s not mandatory for the system or the business.
Holding data of any type has a cost, both financially and ecologically, so we need to test why we’re paying to keep data we don’t think we use. Sometimes there are good reasons but often there aren’t.
Maximum value for minimum effort
The elephant in the room here is, of course, the cleansing part – the important part and the immutable truth part.
Having defined the data needed and how it will be tested, the cleansing can begin. I recommend working on the principle of ‘maximum value for minimum effort’; we want to do the least amount of cleansing possible to enjoy a successful outcome. Our suggestions include:
- Maximise the use of rules-based cleansing in the source system, because we can automate it.
- Maximise touchpoints with our data creators and consumers, for example, implementing quality checks into customer interactions because it’s cheap and effective.
- Maximise efficiency by not aiming for 100 per cent clean, but clean enough to leave only a manageable number of exceptions, because you can deal with that.
Despite adding to the already-complex picture, addressing data quality as part of a transformation programme is an ideal moment to do so. Like a good spring clean, it allows the new system to start lean, green and ready to deliver.
From disaster to success
Ultimately, experience tells you where to look, what to check and how best to fix it in order to move forward fast, and that’s what can turn an impending disaster into transformational success.
David Bamford is the commercial director of Migra Data.

