Home  /  Insights

CRM integration

Duplicate contacts in HubSpot, and how to stop them coming back

Every cleanup works. Most of them get repeated two years later, because nobody fixed the thing producing the duplicates.

· 6 minute read

The short answer

Deduplication has three parts and most teams only do the first. Choose a match key that fits how your data actually arrives, decide which value wins when two records disagree, then validate the forms and integrations creating the records. Merging without the second and third steps buys you about eighteen months.

Why the same cleanup keeps getting commissioned

A deduplication project runs. Thousands of records merge. The database looks correct for a quarter. Then the count starts climbing again, and two years later someone proposes a deduplication project.

Nothing went wrong with the merge. The merge was the easy part, and it was the only part anybody did. Duplicates are produced continuously by forms, imports, integrations and sales reps typing a company name slightly differently. Cleaning the output without changing the input is washing a floor while the tap runs.


Part one: the match key, and what each one costs you

Every deduplication tool asks you to choose what counts as the same person. That choice decides everything else, and each option is wrong in a different direction.

  • Exact email address is safe and misses most of them. The same person appears as a work address, a personal address, an alias and a typo. You will merge almost nothing and feel like nothing worked.
  • Email domain is aggressive and will merge two unrelated people at the same large company into one record. On a database full of enterprise contacts this destroys data.
  • Normalised company name catches Acme Ltd, Acme Limited and ACME, which is genuinely useful, and will also merge two separate businesses that happen to share a name.
  • Fuzzy name matching sounds appealing and should be treated with suspicion. It merges fathers and sons, and it merges common names across unrelated companies.

In practice the usable answer is a combination with a confidence threshold: exact email as the strong signal, normalised company plus surname as a weaker one that creates a review queue rather than an automatic merge. The review queue is not a failure of the system. It is the system admitting that some calls need a person.

Three deduplication match keys compared: exact email misses variants, email domain wrongly merges colleagues, normalised company name catches spelling variants but merges unrelated firms
The match key decides what you catch and what you wrongly merge. There is no setting that does both.

Part two: survivorship, which almost nobody decides on purpose

When two records merge and their values disagree, one has to win. Most teams inherit whatever the platform does by default, which is usually newest wins.

Newest wins is frequently the wrong rule. The newest value often came from a web form where somebody typed their name in lower case and their company as "acme". The older value came from a signed contract, where the legal entity name was correct.

Source based survivorship holds up better. Values from the billing system beat values from a form, which beat values from an enrichment tool. Job title from a manual sales update beats job title from an automated append. It takes an afternoon to agree with sales and finance, and it stops the quiet degradation of every record that passes through a merge.

Write the rule down. It becomes the thing you check when someone says the data got worse after the cleanup.


Part three: the forms, which is where they come back from

This is the part that makes the difference between a cleanup and a fix.

Most duplicates enter through a handful of predictable doors. A form that only asks for an email and a name, creating a new record because nothing matched. An import where someone mapped company to the wrong column. An integration with no match rule of its own, inserting rather than updating. A sales rep creating a contact because search did not find the one that already existed.

Each has an unglamorous fix. Normalise email case and strip plus addressing before matching. Require a company field on forms that feed sales, and validate its format. Give every integration an explicit match rule and test what it does with a near miss. Make search in the CRM actually usable, because reps create duplicates when search fails them, not because they want to.

None of that is interesting work. It is the difference between fixing this once and fixing it every two years.


What good looks like afterwards

A documented match key with a confidence threshold. A written survivorship order. Validation on every inbound door. A review queue someone owns, checked weekly, that takes ten minutes because the volume is small once the inputs are fixed.

And a number you can quote: duplicates created per month, trending down and then flat. If nobody is measuring that, the cleanup has no way of proving it worked, and you will be having this conversation again.

Where this comes from

We do this work

This article is drawn from how we scope and run crm integration. If you recognised your own setup in any of it, that page covers what an engagement looks like, what is included, and what we will not take on.

Tell us where the chain breaks

Also here

More from insights