Skip to Content
IdentityProbabilistic Matching

Probabilistic Matching

Probabilistic matching extends identity resolution beyond exact-match rules by using fuzzy and statistical signals to link records. Each probabilistic match produces a confidence score between 0.0 and 1.0 that flows through to audiences and syncs, letting you control the precision/recall tradeoff at query time without re-running the identity pipeline.

Graph Types

An identity graph’s type is set in step 1 of the creation wizard and decides whether confidence scoring exists at all.

TypeDescriptionConfidence features
DeterministicExact-match rules only. The default. Every link carries confidence 1.0Hidden — no confidence controls, columns or thresholds
ProbabilisticExact and fuzzy rules in the same graph, chosen per merge ruleEnabled — confidence columns, and thresholds at graph, audience and sync level
Step 1 of the wizard with Probabilistic selected, revealing the default confidence threshold slider

Deterministic graphs behave exactly as they always have. A probabilistic graph adds confidence throughout the pipeline: on links, on profile aggregates, on golden record rows, and as a filter in audience queries.

Graph type cannot be changed after creation. To move a deterministic graph to probabilistic, create a new identity graph and run it. Adding a probabilistic rule to a deterministic graph is rejected outright rather than stored as a rule that silently produces nothing.

A probabilistic graph always rebuilds in full. Incremental resolution is not available for it, whatever the run mode says, and every run records probabilistic as the reason. Budget for that before choosing the type.

Configuring a Probabilistic Rule

On a probabilistic graph, every merge rule gains a Match control. Set it to Probabilistic and a configuration panel appears under the rule:

ControlRangeDefaultMeaning
Scoring FunctionFuzzy Name · IP Cluster · HouseholdFuzzy NameHow candidate record pairs are compared
Base Confidence0.10 – 1.00, in steps of 0.050.70The starting confidence assigned to links this rule produces
Min Similarity (Fuzzy Name)0.50 – 1.000.85Minimum similarity for two names to count as a match
Time Window (hours) (IP Cluster)1 – 16824How far apart two events on one IP may be
Name Similarity Threshold (Household)0.50 – 1.000.80How similar surnames at one address must be

A scoring function reads fixed identifier families — not the family its own merge rule is on. A probabilistic rule on email that selects Fuzzy Name still scores the name family. If the families it needs are not mapped, the rule produces no links at all and the run succeeds having done nothing. The rule editor warns before you save, naming the families to map.

Scoring Functions

Fuzzy Name

Matches records with similar names.

  • Reads: the name family — a first_name, last_name or full_name column mapped on participating models.
  • How it works: records sharing a phonetic code are selected as candidates, which keeps the comparison from being all-pairs. Confidence is then the similarity score the warehouse can compute — Snowflake uses Jaro-Winkler, Databricks and ClickHouse use a normalised edit distance. BigQuery and Redshift have no usable similarity function, so every phonetically-matched pair scores the flat base confidence instead, and Min Similarity becomes an all-or-nothing gate on that constant rather than a graded threshold.

Example: “Alice Smith” and “Alice M. Smith” share a phonetic code. On Snowflake the similarity scores 0.92, which clears the 0.85 threshold, so a link is created carrying confidence 0.92.

IP Cluster

Links records seen on the same IP address close together in time.

  • Reads: the ip_address family, plus a column mapped as timestamp or event_timestamp on the same model.
  • How it works: exact IP match, with the time window as a second condition. Every matched pair carries the rule’s base confidence.
  • Without a timestamp: the window is dropped and the links carry half the base confidence, because “same IP, any time” is a far weaker claim than “same IP within a day”.

Example: two records on IP 203.0.113.7 eighteen hours apart fall inside a 24-hour window, and are linked at the rule’s base confidence.

Household

Links records at the same address with similar names.

  • Reads: the address family (postal_code or zip_code) and the name family, and both must be mapped on the same model — the two columns are read from one row. A graph with names on one model and postal codes on another produces nothing, and the rule editor warns about exactly this.
  • How it works: exact postal-code match plus a phonetic name match. Where the warehouse can compute a similarity score, the Name Similarity Threshold is applied as a second condition; on BigQuery and Redshift it cannot be enforced and the phonetic match alone decides. Every link carries the rule’s base confidence.

Example: two records in postal code 10001 named “John Smith” and “J. Smith” match on postal code and phonetic code, and are linked at the rule’s base confidence.

Household links a household, not a person. Treat it as a way to reach a family unit for direct mail or connected TV, not as evidence that two records are the same individual — and give it a low base confidence so a threshold can separate it from person-level matches.

Confidence Scores

Link typeConfidence
Deterministic — email, phone, customer id, any exact ruleAlways 1.0
Fuzzy NameThe computed similarity score, where the warehouse can compute one; otherwise the rule’s base confidence
IP ClusterThe rule’s base confidence, halved when no timestamp column is available
HouseholdThe rule’s base confidence

Propagation

Confidence propagates through the connected-components algorithm. When two records are linked through one or more intermediate links, the confidence of the path between them is the lowest confidence on that path. One weak link therefore lowers the confidence of everything reachable through it.

The output tables in your warehouse carry this:

TableColumnMeaning
_IDENTITY_GRAPHmin_edge_confidenceThe lowest confidence on the path from that record to the rest of its profile
_IDENTITY_PROFILESmin_confidenceThe lowest min_edge_confidence of any record in the profile
_IDENTITY_PROFILESavg_confidenceThe average across the profile’s records
_GOLDEN_RECORD_min_confidenceThe lowest confidence across the records that contributed to the golden record row

Confidence Thresholds

A threshold excludes low-confidence profiles from audience queries and syncs. Thresholds are evaluated at query time, so changing one takes effect on the next query with no re-run of resolution.

Three levels

The effective threshold is resolved highest-precedence first:

  1. Sync override — set on an individual audience sync, for destination-specific tuning
  2. Audience threshold — set on the audience
  3. Graph default — Default Confidence Threshold on the identity graph itself, which the wizard sets to 0.50 for a new probabilistic graph

A threshold of 0.0 at every level means no filtering: all profiles are included whatever their confidence. Thresholds apply only to probabilistic graphs — a deterministic graph has no confidence to filter on.

Example

LevelThresholdEffect
Graph default0.5Base filtering — anything below 0.5 is excluded wherever nothing else is set
Email campaign audience0.8Stricter: only high-confidence profiles qualify for this audience
Google Ads sync(none)Falls back to the audience’s 0.8
Braze sync0.9Strictest override — personalised messaging demands accuracy

Where thresholds apply

  • Audience size estimates, previews and membership — profiles below the threshold are excluded from all three, so estimates move when you change it.
  • Syncs — the same filter is applied when the sync compiles, using the sync’s own override if it has one.
  • Golden records — an audience reading through the golden record filters on _min_confidence instead, so both routes to the same profile agree.

Threshold changes need no pipeline re-run. They do change audience sizes, so review an audience’s estimate after adjusting one.

Best Practices

  1. Start deterministic, add probabilistic later — get the deterministic graph right and validate its profiles first. Once you understand your data’s identity patterns, create a probabilistic graph to expand reach.
  2. Map the families the scoring functions read — a probabilistic rule with no name, ip_address or address columns behind it produces nothing, and the run still succeeds.
  3. Keep limits on the weak families — fuzzy signals chain records together through transitive closure more readily than exact ones. Probabilistic families take no per-profile limit by default, so set one explicitly if a profile is accumulating variants, and keep the shared-value limit tight on IP addresses and postal codes. See Limit Rules.
  4. Review the confidence distribution before setting a threshold — after a run, look at min_confidence in _IDENTITY_PROFILES to see the shape of your data, and set the graph default from that rather than from an arbitrary number.
  5. Tune per destination — ad platforms tolerate a looser threshold in exchange for reach; email and SMS platforms deserve a stricter one, because a wrong match damages sender reputation.
  6. Watch avg_confidence over time — a declining average across runs points at degrading source data or a rule producing weak links at scale.
  7. Pair with Audience Boost — probabilistic resolution expands your first-party graph by linking records you already hold. Audience Boost enriches with third-party identifiers at sync time. They close different gaps.

Next Steps

  • Merge Rules — Configure deterministic and probabilistic rules and set rule priorities
  • Identifier Families — Map the name, address and IP columns the scoring functions need
  • Golden Records — Unified profiles with confidence metadata and survivorship rules
  • Running Resolution — Execute the pipeline and review the output tables
Last updated on