The short answer
Deduplication has three parts and most teams only do the first. Choose a match key that fits how your data actually arrives, decide which value wins when two records disagree, then validate the forms and integrations creating the records. Merging without the second and third steps buys you about eighteen months.
Why the same cleanup keeps getting commissioned
A deduplication project runs. Thousands of records merge. The database looks correct for a quarter. Then the count starts climbing again, and two years later someone proposes a deduplication project.
Nothing went wrong with the merge. The merge was the easy part, and it was the only part anybody did. Duplicates are produced continuously by forms, imports, integrations and sales reps typing a company name slightly differently. Cleaning the output without changing the input is washing a floor while the tap runs.
Part one: the match key, and what each one costs you
Every deduplication tool asks you to choose what counts as the same person. That choice decides everything else, and each option is wrong in a different direction.
- Exact email address is safe and misses most of them. The same person appears as a work address, a personal address, an alias and a typo. You will merge almost nothing and feel like nothing worked.
- Email domain is aggressive and will merge two unrelated people at the same large company into one record. On a database full of enterprise contacts this destroys data.
- Normalised company name catches Acme Ltd, Acme Limited and ACME, which is genuinely useful, and will also merge two separate businesses that happen to share a name.
- Fuzzy name matching sounds appealing and should be treated with suspicion. It merges fathers and sons, and it merges common names across unrelated companies.
In practice the usable answer is a combination with a confidence threshold: exact email as the strong signal, normalised company plus surname as a weaker one that creates a review queue rather than an automatic merge. The review queue is not a failure of the system. It is the system admitting that some calls need a person.
Part two: survivorship, which almost nobody decides on purpose
When two records merge and their values disagree, one has to win. Most teams inherit whatever the platform does by default, which is usually newest wins.
Newest wins is frequently the wrong rule. The newest value often came from a web form where somebody typed their name in lower case and their company as "acme". The older value came from a signed contract, where the legal entity name was correct.
Source based survivorship holds up better. Values from the billing system beat values from a form, which beat values from an enrichment tool. Job title from a manual sales update beats job title from an automated append. It takes an afternoon to agree with sales and finance, and it stops the quiet degradation of every record that passes through a merge.
Write the rule down. It becomes the thing you check when someone says the data got worse after the cleanup.
Part three: the forms, which is where they come back from
This is the part that makes the difference between a cleanup and a fix.
Most duplicates enter through a handful of predictable doors. A form that only asks for an email and a name, creating a new record because nothing matched. An import where someone mapped company to the wrong column. An integration with no match rule of its own, inserting rather than updating. A sales rep creating a contact because search did not find the one that already existed.
Each has an unglamorous fix. Normalise email case and strip plus addressing before matching. Require a company field on forms that feed sales, and validate its format. Give every integration an explicit match rule and test what it does with a near miss. Make search in the CRM actually usable, because reps create duplicates when search fails them, not because they want to.
None of that is interesting work. It is the difference between fixing this once and fixing it every two years.
What good looks like afterwards
A documented match key with a confidence threshold. A written survivorship order. Validation on every inbound door. A review queue someone owns, checked weekly, that takes ten minutes because the volume is small once the inputs are fixed.
And a number you can quote: duplicates created per month, trending down and then flat. If nobody is measuring that, the cleanup has no way of proving it worked, and you will be having this conversation again.
Where this comes from
We do this work
This article is drawn from how we scope and run crm integration. If you recognised your own setup in any of it, that page covers what an engagement looks like, what is included, and what we will not take on.
Also here
More from insights
- Why your HubSpot and Salesforce records do not match Two systems that can both write the same field will eventually disagree. Here is why, and what to do instead.
- Will server-side tagging recover the data you lost to consent? Short answer: no. Longer answer: it fixes several real problems, and it is worth knowing which.
- GA4 revenue does not match your shop platform The two numbers will never match exactly. The goal is a small gap you can explain in one sentence.
- Lead scoring that sales will actually use If sales ignores the score, the score is wrong. That is the whole feedback loop, and most models never get it.
- Your new website produced fewer leads than the old one It is rarely the design. It is almost always something that was carried across incompletely.
- Zapier, Make or n8n for a marketing team The feature lists are converging. The real difference is operational, and it shows up in the second year.