Skip to main content

The CDP data flow: sources → identity → governance → activation → measurement

The CDP is a pipeline: each stage’s output is the next stage’s input. Events enter from sources, identity folds them onto one profile per person, governance decides what may leave and what must be deleted, activation publishes audiences to the destinations you spend against, and measurement reads the same event stream the sources wrote. This page is the map. Each stage section names the dashboard surface, the public endpoints, and the guide that goes deep. If you want one continuous pass instead of the map, run the first activation walkthrough against a test source and destination first.

Stage 1: Sources

Three roads converge on the same ingestion primitives:
  • Signed live ingest. Your SDK or server posts to POST /cdp/v1/<ingest_id>/track with the three signature headers. Ingest signing covers the wire format; source and destinations covers minting and rotation.
  • Batch file ingest. POST /api/v1/cdp/file-ingest carries up to 200 CSV, JSON, or JSONL rows per request and returns per-row errors and counters. Batch-load guide.
  • Scheduled object-storage sources. Drop files into your own S3, GCS, or Azure Blob prefix and the platform polls them on a cadence. Browse providers and validate a config from Integrations → CDP → File sources (GET /api/v1/cdp/file-sources, POST /api/v1/cdp/file-sources/:source_id/validate-config); manage the connected source with GET/PATCH /api/v1/cdp/object-storage-source-config/:provider. Credentials are write-only and never returned by a read.
Every row, live or backfilled, passes the same identity rules, the same deterministic message-id dedupe, and the same erasure gate. Confirm arrival in the event debugger before trusting anything downstream.

Stage 2: Identity

Identity decides which identifiers belong to one person. Everything downstream reads the contact row it produced.
  • Identity resolution console (Audience → Identity resolution): deterministic preview, auto-merge above your threshold, and the probabilistic review queue (GET /api/v1/cdp/identity/merge-candidates, ranked with confidence, band, and per-signal breakdown). Approvals run through the standard merge with its 30-minute undo. Identity resolution concept.
  • Simulation gate: POST /api/v1/cdp/identity-resolution/simulate returns clusters formed, largest cluster, delta versus active rules, and an over-merge guard naming the offending rule. The dashboard refuses to save a match-shape change until a simulation has run in your session. Simulation guide.
  • Identity graph: GET /api/v1/cdp/profiles/by-user-id/:user_id/identity-graph shows which identifiers and merged contacts were stitched into a profile, with provenance per fold. Identity graph and device graph.
  • Linked Audiences (Audience → Linked audiences): model warehouse entities and relationships in the data graph once, then build audiences that traverse them. Linked Audiences guide.

Stage 3: Governance

Three surfaces, all tenant-owned: you set the policy, the platform enforces it and keeps the trail.
  • Erasure propagation: per-destination policy (GET/PUT /api/v1/cdp/erasure/destination-modes, delete or suppress) and per-erasure fan-out (POST /api/v1/cdp/erasure/:erasureId/propagate), with a per-destination trail reconstructed from immutable audit rows. Use it after a DSAR completes inside Orbit. Erasure propagation guide.
  • Consent revocation propagation: POST /api/v1/cdp/consent/:contactId/:channel/propagate-revocation fans a channel opt-out out to your connected downstream tools, with the same trail surface. It governs copies in tools you connected; sends on Orbit channels stay gated by the per-send consent ledger. Consent inspector.
  • Clean rooms: POST /api/v1/cdp/clean-room/match measures audience overlap from hashed identifiers only, returns counts and rates, and suppresses results below the k-anonymity threshold. Clean room model.

Stage 4: Activation

An activation publishes a segment to a destination on a schedule you control, consent-checked per member before anything leaves.
  • Paid media: Meta, Google, TikTok, LinkedIn, Snapchat, Pinterest, Reddit, The Trade Desk, and Criteo over OAuth. Activation pipelines.
  • Webhook sink: signed JSON POST to your own HTTPS endpoint, HMAC-verified on receipt. Same guide.
  • Reverse ETL: object-storage destinations (S3, GCS, Azure Blob) serialize segments, profiles, or events to CSV/JSONL on schedule into your own bucket; ERP sync writes segments into Oracle NetSuite, SAP S/4HANA, Workday, or QuickBooks Online as native object upserts. Reverse ETL and warehouse exports, ERP object sync, source and destinations.
  • Server-side conversion APIs: Meta Conversions API and TikTok Events API destinations forward hashed conversions for attribution. Meta and TikTok conversion forwarding.
  • Predictive schedules: store a per-model cadence and top-N so churn, conversion, LTV, and fatigue audiences re-materialize as segments automatically. Predictive models, activation schedule concept.

Stage 5: Measurement

  • Funnels and retention: POST /api/v1/cdp/analytics/funnel and POST /api/v1/cdp/analytics/cohort-retention run on demand over the collected event stream, counting resolved subjects per step. Funnels and retention guide.
  • Predictive models: train, evaluate, and score churn, conversion intent, lifetime value, and fatigue over your own history; scores become segments stage 4 activates. Predictive models guide.
  • Operational proof: the sync run log records per-destination run history; the event debugger’s test-event tooling fires a synthetic event down a destination before real traffic flows; the Insights hub is the tile grid for everything else.
When a number looks wrong, walk the stages backward: measurement misreads usually trace to an identity over-merge or a source gap, not to the report itself.