Skip to main content

Identity resolution and merge: end-to-end walkthrough

The individual guides — identity resolution, contact merge, the survivorship policy concept, the Customer-360 workspace, anonymous identity stitching, and account relations — each go deep on one surface. This walkthrough strings them into the one operator loop that keeps a golden profile golden: ingest with identifiers, scan for duplicates, preview the merge, apply it with a survivorship plan, and confirm the survivor reads right in Customer-360. Run the steps in this order. Identity work downstream of segments is the expensive order; identity work before segments is the cheap one.

1. Why identity precedes ingestion

One person should produce one golden profile. When ingestion races ahead of identity rules, the same person lands as two contact rows — a web signup keyed on email, an inbound SMS keyed on phone — and cleaning that up later is merge debt: every segment built on the dirty base counted them twice, every campaign that sent to the twin has delivery history you now have to reconcile, and every opt-out recorded on one row was reachable through the other. The cheaper order is deterministic rules first (step 2), then the review queue clears what the rules can-not settle, then segments and scores build on the single profile.

2. Ingest with the dedupe surfaces in mind

Identity resolution can only match over the identifiers your contacts actually carry, so decide which dedupe surface resolves new identifiers BEFORE the first import or SDK call: Deterministic rules are the always-on layer. PUT /api/v1/cdp/identity-rules authors one rule per identifier type (external_id, email, phone, anonymous_id) with a priority; ingestion resolves to an existing contact on the strongest enabled rule the moment an identifier arrives. Prefer your own system’s stable key in external_id when one exists — it beats phone and email formats:
Feeding identifiers at contact creation is the second half of the same decision — pass phone, email, and external_id on POST /api/v1/contacts, on imports, and on webhook-driven upserts:
A contact created with the three identifiers above is the anchor record the rest of this walkthrough folds duplicates into. Its profileId (cnt_01f3…) never changes across merges — pick the record your external systems already reference as the survivor, not the oldest row.

3. Map the ten dedupe and merge surfaces

The Audience hub splits identity across ten consoles; knowing which one answers which duplicate shape is the first operator skill. The dashboard routes sit under /audience (see the Audience hub orientation for the full tile map):

4. Decision tree — which surface picks up a duplicate

A new duplicate row needs one of three resolutions, and the tree is stable across the ten surfaces:
When a candidate could go either way, default to the review queue. Rejecting a bad pair costs a re-scan; accepting a bad merge folds one person into somebody else’s record and corrupts every downstream reader at once.

5. Configure identity resolution — rules, review, and bands

The identity resolution guide is the deep dive; the operator walk is short:
  1. Enable the deterministic rule per identifier. The example in step 2 enabled external_id at priority 1. Repeat for email and phone at lower priorities so the strongest signal wins first.
  2. Simulate before enabling. POST /api/v1/cdp/identity-resolution/simulate returns the what-if preview — which pairs the proposed rule would merge, the per-field survivorship preview, and an over-merge warning when a rule would collapse clearly-different people. Run this against the full policy before it goes live.
  3. Let the scan clear the residue. GET /api/v1/cdp/identity/merge-candidates with include_auto_merge=false returns only the review band — the ambiguous remainder the rules could not settle. Bound the scan with min_confidence (0–100) when a noisy import fills the queue.
Each candidate carries a confidence score and a band:
  • High confidence (auto_merge) — prepared to merge automatically, for example normalized phone and email corroborating each other.
  • Review band (review) — a steward confirms each pair before it merges; a phone-only link whose members carry different emails lands here by design.

6. Read the side-by-side diff in the Merge queue

Open Audience → Merge duplicates (/audience/merge). The scan finds duplicate groups with either strategy:
Each group renders as a side-by-side diff: the two records’ scalar fields (name, email, phone, WhatsApp/Viber ids, company, country, timezone, language, external id) line up per row, and the custom-field surface shows the same per-key split. The candidate below — anchor record cnt_01f3… vs. an import twin cnt_02a8… — shows a confidence of 91 and matched fields phone and email, so the diff reads like the anchor’s row with a second column: The decision per conflicting field is the survivorship rule the pair honors; equal values need none, null-vs-value fills under prefer_non_null, and the genuinely conflicting fields (display_name here) need a pin or a tenant-policy rule.

7. Work the survivorship policy

The tenant survivorship policy — authored on the Identity resolution page or GET/PUT/DELETE /api/v1/cdp/survivorship-policy — fills in every field you do not pin per merge. Three strategies cover the worked-example cases: Last-touch wins (most_recently_updated). Take the value from the record with the newer updated_at. Use it when the freshest write is the most trustworthy — a web-preference form the person just filled beats a stale CRM row. First-touch wins (most_recently_created inverted, or prefer_target where the target is the older record). Use it when the original registration is canonical and later imports polluted the row. Provenance wins (prefer_source_system). Prefer whichever record’s source column matches a named system (crm, salesforce, import) — the strategy to reach for when “our CRM is the source of truth for name and company.”
The governed fields are a closed set — the scalar contact columns the merge actually honors: phone, email, whatsapp_id, viber_id, first_name, last_name, display_name, company, country_code, timezone, language, external_id. Channel memberships, tags, conversation history, and CDP events fold in as unions with no pin available; the per-field decision applies only to scalars. Timestamps conflicts settle per field and per rule — run the rules simulator against the full policy before enabling it, so a last-touch rule whose timestamps disagree catches the over-merge warning rather than sweeping two different people.

8. Write the merge

POST the pair with the survivorship plan from step 6. A per-request fieldStrategies pin beats the tenant policy, and any unpinned field falls to the policy then to mergeStrategy:
The response returns the survivor and a merge_id:
(surviving_source is the equivalent label when you author the merge payload against the concept’s naming — the field names on the wire are primaryId / secondaryIds; inside the survivorship policy the same pair reads as target/source.) The 30-minute undo_window is the rollback path: POST /api/v1/contacts/unmerge/:mergeId reverts the fold inside the window; GET /api/v1/contacts/merge-history keeps the audit row (survivor, folded-in ids, acting user, strategy) forever. After the window the fold is durable.

9. Confirm the survivor in Customer-360

The merge resolves identity; Customer-360 confirms every downstream reader picked the survivor up. Open Audience → Contacts → <id>/360 or call the snapshot endpoint:
Expect the folded-in record’s conversations, calls, and events to now read on the survivor — the endpoint fans every source in parallel and returns the full envelope, so a missing section is an empty array rather than a failed read. When segments or personalization reference the folded-in ids, point them at the survivor id before you trust the base. See the Customer-360 workspace guide for the full envelope map.

10. Edge cases that gate every merge

Soft-deleted and merged-away rows. A folded-in record stops being reachable as a separate profile, but a re-imported identical contact re-enters the duplicates queue — run the scan after every bulk import so a re-imported twin lands in the review band, not back in segments. Suppressed contacts. A fold does not clear suppression. When either side of the pair carried an opt-out, verify the survivor against GET /api/v1/cdp/consent/:contactId and POST /api/v1/cdp/consent/check before any campaign sends — the authoritative consent answer, not either record’s raw flags. PII scoping. The identity, merge, and Customer-360 surfaces render only for owner, admin, and developer seats; a viewer sees none of the tiles, and the destination page 403s them regardless. The Audience hub orientation maps the permission scopes per tile. Consent-aware merging. Consent is tenant-owned and survives the fold — the survivor inherits every recorded grant or revocation from both sides, and a revocation on the folded-in record continues to suppress the survivor on that channel. For the auditor-facing export, see export consent and suppression. B2B account relations. When the contact is a member of golden accounts (/audience/accounts), the fold updates the contact↔account membership — the account’s golden record does not change, but the Members list now names the survivor. Re-check the account’s detail pane when the merged contact appears under multiple relations. See account relations.

See also