Skip to Content
InsightsSync Health

Sync Health

The Sync Health dashboard provides a composite health score for every sync in your workspace, combining success rate, error trends, latency, and data freshness into a single at-a-glance view. Use it to identify failing syncs, diagnose degradation, and monitor throughput and SLA compliance.

Health Score

Each sync receives a health score from 0 to 100, computed from four weighted components:

ComponentWeightWhat it measures
Success rate40%The share of runs in the period that completed
Failure streak25%How many runs have failed consecutively. Five in a row zeroes this component; the penalty rises linearly to that point.
Duration stability20%How consistent run durations are — the standard deviation of duration against its mean. A sync that takes the same time every run scores full marks; one whose duration swings scores less, whether it is getting slower or faster.
Freshness15%Whether the sync ran successfully at all within the period

The streak component is separate from success rate on purpose. Nineteen successes and one failure and nineteen successes with the failures at the end give the same success rate and mean very different things; the streak is what tells them apart.

Score Ranges

ScoreStatusMeaning
80–100HealthySync is operating normally. No action needed.
60–79WarningSync is showing signs of degradation. Investigate soon.
40–59DegradedSync has significant issues. Immediate investigation recommended.
0–39CriticalSync is failing or severely impaired. Requires urgent attention.

Reading the Dashboard

The Overview tab carries the health score alongside the numbers it was computed from: total runs in the period, how many succeeded and failed, the success rate, and the average duration. Sync Analytics narrows the same cards to one sync.

Two things on the card are worth reading before the score itself:

  • The daily breakdown — one entry per day with its own totals and average duration. A score of 62 tells you something is wrong; the daily breakdown tells you when it started.
  • Top errors — the recurring error messages across the period, each with a count, when it was last seen, and which destination type produced it. Ten runs failing with one message is a different problem from ten runs failing with ten.

Throughput

Throughput is reported as daily row changes — rows added, updated and deleted per day — for the period you select.

Read as a series it answers the questions a single run cannot:

  • Is the volume normal? A day with an order of magnitude more rows than its neighbours usually means the model changed, not the customers.
  • Did the data stop moving? Days of zeroes on a sync that is still reporting successful runs means the model has stopped returning changes.
  • Is it growing? Steady growth in daily volume is what warrants a look at warehouse compute before the syncs start overrunning their schedules.

Narrow it to one sync with sync_id, or read it across the workspace.

Duration

Duration is reported per day, as three numbers:

MetricWhat it is
AverageMean run duration that day
P50The median run
P95The 95th percentile — the slow tail

The gap between P50 and P95 is the useful part. A sync whose median is a minute and whose P95 is twenty is not a slow sync; it is a sync that occasionally does something very different, and the run history is where you find out what.

Duration also feeds the health score through its stability component, so a sync whose durations swing scores lower even when every run succeeds.

Schedule Adherence

The SLA view answers one question per sync: is it running as often as its schedule says it should?

Zeotap reads the sync’s own schedule, works out the interval it implies, and compares that to when the sync last ran:

FieldMeaning
ScheduleThe sync’s configured schedule
Last run atWhen it last ran, or nothing if it never has
On scheduleWhether the time since that run is within one interval
Overdue minutesHow far past one interval it is. Zero when on schedule; for a sync that has never run, the whole interval.

A sync that is overdue is not the same as a sync that is failing — nothing ran, so nothing failed. It is the case the run-based metrics on this page cannot see, which is why it has a view of its own.

Diagnosing a Failing Sync

When a sync’s health score drops, work down from the score to the runs behind it.

Step 1: Work out which component moved

The score is one number over four components, so start by asking which one fell:

  • Runs are failing — success rate is down. Go straight to the errors.
  • The recent runs are failing — the failure streak is what is costing you the most, and it means the problem is current rather than historical.
  • Durations are swinging — stability is down. The P50/P95 gap on the Duration card says how far.
  • Nothing ran — freshness is zero and so is everything else. Check schedule adherence: an overdue sync has not failed, it has not started.

Step 2: Read the top errors

The health card lists the recurring error messages with a count and the destination type each came from. One message across every failure is a single cause; a scatter of different messages usually is not.

Step 3: Inspect the runs

Open the sync’s Runs tab for the individual runs:

  • Error message — click it to read the whole thing
  • Batches — tells a wholesale failure apart from a few rejected batches
  • Row counts — compare across runs to spot an unexpected spike or collapse in volume

Common Causes

SymptomLikely causeResolution
Runs failing with auth errorsOAuth token expired or revokedReconnect the destination
Runs failing with rate limit errorsDestination API throttlingReduce sync frequency, or space the schedules of syncs sharing a destination
Runs timing outWarehouse query performanceCheck warehouse resource usage and the model’s query plan
Runs completing with batches failingSome records rejected at the destinationCheck the field mapping for type mismatches and required fields
Overdue, but recent runs succeededThe schedule is not firing as often as expectedCheck the sync’s schedule and the workspace’s schedule timezone

API Endpoints

Every insights endpoint is workspace-scoped and takes an optional period — a number of days written as 7d, 30d, 90d and so on. It defaults to 30d, and any positive number of days is accepted.

Sync Health

GET /api/v1/workspaces/{id}/insights/sync-health?period=30d

Returns the aggregate over the period: total_runs, success_count, failed_count, success_rate, avg_duration_sec, the composite health_score, a daily_breakdown (one entry per day with its own totals and average duration), and top_errors — the recurring error messages with a count, when each was last seen, and the destination type it came from.

Narrow it with sync_id, or with sync_type to one kind of run (model, audience, journey, store_feed).

Sync Throughput

GET /api/v1/workspaces/{id}/insights/sync-throughput?period=30d

Daily row-change time series. Takes sync_id.

Sync Duration

GET /api/v1/workspaces/{id}/insights/sync-duration?period=30d

Daily duration distribution. Takes sync_id.

Schedule Adherence

GET /api/v1/workspaces/{id}/insights/sync-sla?period=30d

One entry per sync: its schedule, last run, whether it is on schedule, and how many minutes overdue it is.

Last updated on