In the CRM: 20,000 contacts. Of those, usable: Maybe 300.
The rest: questionable, incomplete, error-ridden, and not maintained. Completely useless for sales and marketing campaigns. Yet customer data – or so they say – is a company’s most valuable asset! So, what is to be done?
A European software company’s CRM held over 20,000 contacts – and just roughly 300 of them were actually usable. Years of neglected data hygiene meant no personalized campaigns, no regional targeting, and no realistic way to fix it by hand (three student assistants and a few months of manual cleanup would’ve cost tens of thousands of euros).
Instead, working with the company’s data science team, we built an algorithmic scraper that cross-referenced the incomplete CRM records against other data sources already available inside the company, filling gaps and normalizing the data automatically.
The result: a fully populated, GDPR-compliant CRM with a data error rate under 5%, built and deployed in about two weeks – and a sales lead and marketing team who finally had something to work with.
Client
B2B Software Company, Enterprise Hospitality Sector
Year
2023
Scope
Gap Analysis
Problemsolving
Data Science
Testing
Process Development
Educating
Read full story
Tech Used
Hubspot CRM
HQ revenue suite
The Problem: A "Single Source of Truth" Nobody Could Actually Use
Everyone talks about how a complete data base is mandatory. But when it comes to building it, the motivation often thins.
The client is a leading software vendor operating across Europe, serving a genuinely mixed customer base – small businesses, mid-market, and enterprise accounts – with offerings that vary heavily by region. For sales and marketing, that means campaigns need to be segmented by region, customer type, and classification, on top of the usual expectation of personalization.
The CRM couldn’t support any of that. Records were missing last names, addresses, cities, postal codes, or countries seemingly at random.
Sometimes the company name was entered as a first name, or vice versa. Sometimes a record was missing almost everything.
Years of neglected data maintenance had quietly caught up with the business — and it was now blocking expansion plans. You can’t target a new market if you can’t even filter your own CRM by country.
Pantone® 532 C
C71 M65 Y64 K72
#232323
Pantone® 1795 C
C19 M90 Y79 K9
#a50834
Pantone® 128 C
C3 M14 Y76 K0
#f9d55c
Pantone® 663 C
C02 M01 Y01 K00
#f7f7f7
The Solution: Let the Data Do the Work
The first idea on the table was straightforward: have a few student assistants clean the data by hand, record by record. It sounded like a bad idea from the start, but it seemed worth scoping out anyway. Four hours after kickoff, the estimate was in: several months of work and tens of thousands of euros — and three very burned-out students at the end of it. Not viable.
The real solution came out of a conversation with the company’s data science team: could an algorithmic approach handle this instead? Together, we built what the data scientists called an “advanced algorithm” — paired with a custom scraper — that cross-referenced the incomplete CRM records against other data sources the company already had access to, filling in gaps and normalizing formatting automatically. (And yes, all sources were confirmed legal to use — that question came up more than once.)
The Result: A CRM People Can Actually Work With
Development took about two weeks, implementation a few days, and a full run of the process took just minutes. The data science working student got an unusually interesting project for their portfolio – and there was even time left over for a proper GDPR review.
The final error rate landed around 5% – roughly 1k records that still needed manual review or deletion. Compared to the manual alternative, that was a strong outcome. The data foundation wasn’t just restored; in a lot of ways, it existed properly for the first time.
Marketing and sales could ‘suddenly’ build and launch region-specific campaigns, and filter the contact base by criteria that simply hadn’t been usable before.
Just as importantly, the project raised awareness inside the company about why data quality matters in the first place – and sparked enough ambition that a proper CRM data governance process could then be built, communicated, and rolled out across departments. As far as anyone can tell, it’s still being followed today.




