Sync Health
The Sync Health dashboard provides a composite health score for every sync in your workspace, combining success rate, error trends, latency, and data freshness into a single at-a-glance view. Use it to identify failing syncs, diagnose degradation, and monitor throughput and SLA compliance.
Health Score
Each sync receives a health score from 0 to 100, computed from four weighted components:
| Component | Weight | What it measures |
|---|---|---|
| Success rate | 40% | The share of runs in the period that completed |
| Failure streak | 25% | How many runs have failed consecutively. Five in a row zeroes this component; the penalty rises linearly to that point. |
| Duration stability | 20% | How consistent run durations are — the standard deviation of duration against its mean. A sync that takes the same time every run scores full marks; one whose duration swings scores less, whether it is getting slower or faster. |
| Freshness | 15% | Whether the sync ran successfully at all within the period |
The streak component is separate from success rate on purpose. Nineteen successes and one failure and nineteen successes with the failures at the end give the same success rate and mean very different things; the streak is what tells them apart.
Score Ranges
| Score | Status | Meaning |
|---|---|---|
| 80–100 | Healthy | Sync is operating normally. No action needed. |
| 60–79 | Warning | Sync is showing signs of degradation. Investigate soon. |
| 40–59 | Degraded | Sync has significant issues. Immediate investigation recommended. |
| 0–39 | Critical | Sync is failing or severely impaired. Requires urgent attention. |
Reading the Dashboard
The Overview tab carries the health score alongside the numbers it was computed from: total runs in the period, how many succeeded and failed, the success rate, and the average duration. Sync Analytics narrows the same cards to one sync.
Two things on the card are worth reading before the score itself:
- The daily breakdown — one entry per day with its own totals and average duration. A score of 62 tells you something is wrong; the daily breakdown tells you when it started.
- Top errors — the recurring error messages across the period, each with a count, when it was last seen, and which destination type produced it. Ten runs failing with one message is a different problem from ten runs failing with ten.
Throughput
Throughput is reported as daily row changes — rows added, updated and deleted per day — for the period you select.
Read as a series it answers the questions a single run cannot:
- Is the volume normal? A day with an order of magnitude more rows than its neighbours usually means the model changed, not the customers.
- Did the data stop moving? Days of zeroes on a sync that is still reporting successful runs means the model has stopped returning changes.
- Is it growing? Steady growth in daily volume is what warrants a look at warehouse compute before the syncs start overrunning their schedules.
Narrow it to one sync with sync_id, or read it across the workspace.
Duration
Duration is reported per day, as three numbers:
| Metric | What it is |
|---|---|
| Average | Mean run duration that day |
| P50 | The median run |
| P95 | The 95th percentile — the slow tail |
The gap between P50 and P95 is the useful part. A sync whose median is a minute and whose P95 is twenty is not a slow sync; it is a sync that occasionally does something very different, and the run history is where you find out what.
Duration also feeds the health score through its stability component, so a sync whose durations swing scores lower even when every run succeeds.
Schedule Adherence
The SLA view answers one question per sync: is it running as often as its schedule says it should?
Zeotap reads the sync’s own schedule, works out the interval it implies, and compares that to when the sync last ran:
| Field | Meaning |
|---|---|
| Schedule | The sync’s configured schedule |
| Last run at | When it last ran, or nothing if it never has |
| On schedule | Whether the time since that run is within one interval |
| Overdue minutes | How far past one interval it is. Zero when on schedule; for a sync that has never run, the whole interval. |
A sync that is overdue is not the same as a sync that is failing — nothing ran, so nothing failed. It is the case the run-based metrics on this page cannot see, which is why it has a view of its own.
Diagnosing a Failing Sync
When a sync’s health score drops, work down from the score to the runs behind it.
Step 1: Work out which component moved
The score is one number over four components, so start by asking which one fell:
- Runs are failing — success rate is down. Go straight to the errors.
- The recent runs are failing — the failure streak is what is costing you the most, and it means the problem is current rather than historical.
- Durations are swinging — stability is down. The P50/P95 gap on the Duration card says how far.
- Nothing ran — freshness is zero and so is everything else. Check schedule adherence: an overdue sync has not failed, it has not started.
Step 2: Read the top errors
The health card lists the recurring error messages with a count and the destination type each came from. One message across every failure is a single cause; a scatter of different messages usually is not.
Step 3: Inspect the runs
Open the sync’s Runs tab for the individual runs:
- Error message — click it to read the whole thing
- Batches — tells a wholesale failure apart from a few rejected batches
- Row counts — compare across runs to spot an unexpected spike or collapse in volume
Common Causes
| Symptom | Likely cause | Resolution |
|---|---|---|
| Runs failing with auth errors | OAuth token expired or revoked | Reconnect the destination |
| Runs failing with rate limit errors | Destination API throttling | Reduce sync frequency, or space the schedules of syncs sharing a destination |
| Runs timing out | Warehouse query performance | Check warehouse resource usage and the model’s query plan |
| Runs completing with batches failing | Some records rejected at the destination | Check the field mapping for type mismatches and required fields |
| Overdue, but recent runs succeeded | The schedule is not firing as often as expected | Check the sync’s schedule and the workspace’s schedule timezone |
API Endpoints
Every insights endpoint is workspace-scoped and takes an optional period — a number of days written as 7d, 30d, 90d and so on. It defaults to 30d, and any positive number of days is accepted.
Sync Health
GET /api/v1/workspaces/{id}/insights/sync-health?period=30dReturns the aggregate over the period: total_runs, success_count, failed_count, success_rate, avg_duration_sec, the composite health_score, a daily_breakdown (one entry per day with its own totals and average duration), and top_errors — the recurring error messages with a count, when each was last seen, and the destination type it came from.
Narrow it with sync_id, or with sync_type to one kind of run (model, audience, journey, store_feed).
Sync Throughput
GET /api/v1/workspaces/{id}/insights/sync-throughput?period=30dDaily row-change time series. Takes sync_id.
Sync Duration
GET /api/v1/workspaces/{id}/insights/sync-duration?period=30dDaily duration distribution. Takes sync_id.
Schedule Adherence
GET /api/v1/workspaces/{id}/insights/sync-sla?period=30dOne entry per sync: its schedule, last run, whether it is on schedule, and how many minutes overdue it is.
Related Resources
- Reverse ETL — Configure and manage warehouse-to-destination syncs
- Syncs — Audience-to-destination sync configuration
- Activation Coverage — See which audiences are reaching which destinations
- API Reference — Full API documentation