Spectrum Technology Platform

 View Only

Evolution of Data Quality – Getting Smart in the ‘Age of the Customer’

By Himanshu Verma posted 12-10-2018 13:27

  

Evolution of Data Quality – Getting Smart in the ‘Age of the Customer’


Recently I was thinking back to the date when I joined Pitney Bowes. It’s been more than 9 years now and I think that’s a long time. Facebook at that time was 5 years old, WhatsApp was just launched and Instagram wasn’t even born. Today, Facebook is not only alive and kicking, but also owns WhatsApp and Instagram. That’s a great example of evolution!

But hey, why I am talking about Facebook? Two reasons – a) It is our customer and b) we are helping them evolve… :) 

Pitney Bowes Data Quality evolution goes back to 30 years. In 1993, we launched stand-alone list products which supported data normalization, merge/purge and casing features. Fast forward to 2002, we launched Enterprise Data Quality tool ‘DataSight’ which was built upon services oriented architecture, followed by ‘Sagent’ in 2003 which offered strong Data Integration capabilities. In 2006, we launched ‘CDQP’ (Customer Data Quality Platform), which was an all-in-one data profiling, data cleansing and data enrichment suite. This is where we highlighted our focused market, i.e., Customer Data Management. CDQP laid the foundations for today’s Spectrum Technology Platform.

As per Forrester, year 2010 marked the beginning of ‘Age of the Customer’, which they define as an age where organizations need to be obsessed about their customers and their needs. If you see our evolution, it coincides with Forrester’s market analysis. Today, Spectrum offers a comprehensive set of data management and analytics capabilities running on a single, unified platform for managing Customer related information. Graph-based MDM, Big Data Quality, Applied Machine Learning capabilities are all examples of that. Clearly, we have come a long way.

The timeline view of our Data Management evolution journey below gives you a glimpse of major milestones we’ve achieved over past 3 decades:

Evolution of Data Quality in the age of the customer

As a regular user of Spectrum I’ve always found it capable of solving customer related challenges, but it has always been suited more to IT users as compared to Business users. Even for IT users, there is a learning curve involved that requires time and effort to grow their knowledge from a beginner level to an expert. For instance, I’ve always found it challenging to apply the right matching algorithms to find duplicates when my data contains a mix of Names, Phone numbers, Emails and Address fields. There are numerous algorithms offered by Spectrum to choose from. You only get better with time, not to forget the expertise required to use matching stages within Enterprise Designer – MatchKey Generator, Intraflow, Interflow, Transactional Match, Candidate Finder, and more. If you are a newbie to Spectrum, you really need a jumpstart to get going.

The good news is, things are changing thick and fast. We are no longer looking to build features thinking of only IT-users. Rather, we are focused more on building persona-driven features, web-based applications with guided tours, and automating repetitive steps by leveraging Machine Learning capabilities. 

Going back to my example of finding duplicates with Spectrum; how about if I get a web-based interface, where I can select a data source, let the application do all the intelligent work for me to identify the right pairs of duplicates, and I am required only to tag those duplicates as matches and non-matches as per my use case? That sounds brilliant! I am all for it :)

If that excites you too, then you are in for a surprise: Smart Data Quality works exactly like that. It comes with a brand new interface and lets you find duplicates in your data source in 5 easy steps –

  1. Select a data source
  2. Choose columns from your data required for finding duplicates
  3. Select an optimal threshold for creating logical groups of similar records
  4. View and tag possible duplicates generated by the underpinned machine learning model
  5. Tag all the proposed duplicates correctly as matches and non-matches, then leverage the machine generated MatchKeys and MatchRules in your existing dataflows

That’s really cool! But aren’t you interested to know how trustworthy this process is? Can it match the accuracy and quality of the manual process? Does it yield the same benefit as manually selecting each algorithm carefully and taking multiple iterations to cover all the variations and outliers? If you’ve spent enough time with Spectrum, I think all these questions are fair to be asked.

I’ve used Smart Data Quality and I see a lot of intelligence built around it –

  • The moment you select columns in your data source, the semantic classification process kicks in to let the model know which algorithms to use for creating match rules. This means the model will not try to apply a name-dependent algorithm like Soundex on a column that contains phone numbers or emails.
  • While grouping the similar records together, the model makes sure to cover all the variations and present it to you in the form of possible duplicate pairs. Based on internal logical calculations, the model picks two records from same groups, close groups, and far away groups so that the relevant match, non-match, and unsure records are shown to you on the user interface for tagging purposes.
  • Once you tag the proposed duplicate pairs as a match or non-match, the model runs various matching algorithms on those pairs on the basis of detected semantic type.
  • The more pairs you tag as per your business use case, the more the machine learning model learns to differentiate between a match and non-match pair.

For me this is revolutionary, and I see Smart Data Quality as a prime example of evolution. It is a step in right direction and just the beginning. If you are interested in learning more about Smart Data Quality, please check out the user guide at this link.

I encourage you to start using Smart Data Quality to build your demos for clients and prospects. If you’ve feedback or suggestions to improve it further, please submit your ideas in the Idea portal here or you can reach out to me directly at Himanshu.verma@pb.com

2 comments
79 views

Permalink

Comments

02-05-2019 08:22

Hi Himanshu, this sounds really great!
I have been quite impressed with the demo and would love to see this used with "real" customer data. I think 15 years ago we would have this a "wizard", but the magic is gone, so now we call it "smart" (and it is actually much better than a wizard).
Congratulations to this innovation!

12-13-2018 15:28

Exciting new innovation Himanshu!  It is noteworthy that while ML has a lot of promise in general, there is still reluctance to let machine driven black box approaches drive business decisions without any transparency around the eventual logic that led to a decision.  I like that our approach addresses this uniquely, making Machine-Learning versus Rules-Driven entity resolution a false choice!  Looking forward to hearing from our clients.