Loaders
Loaders pull data from SaaS applications and third-party services into your data warehouse on a recurring schedule. Instead of building and maintaining custom ETL pipelines, you configure a loader in Zeotap, and it handles extraction, schema mapping, and incremental synchronization automatically.
What Is a Loader?
A Loader is a managed ingestion pipeline that connects to a SaaS application’s API, extracts data from one or more objects (e.g., Salesforce Contacts, Stripe Subscriptions), and writes it into your data warehouse. Once the data is in your warehouse, it becomes available to Warehouses, Models, and the rest of the Zeotap platform.
Loaders complement Warehouses. While a Warehouse gives Zeotap read access to data already in your warehouse, a Loader brings external data into your warehouse in the first place.
Supported Connectors
Zeotap supports 15+ loader connectors across CRM, marketing automation, advertising, payments, support, and productivity categories.
| Connector | Category | Authentication | Key Objects |
|---|---|---|---|
| Salesforce | CRM | OAuth 2.0 | Contacts, Leads, Accounts, Opportunities |
| Salesforce Service Cloud | CRM | OAuth 2.0 | Cases, Case History, Emails, Messaging Sessions, Knowledge |
| HubSpot | CRM | OAuth 2.0 / API Key | Contacts, Companies, Deals |
| Stripe | Payments | API Key | Customers, Charges, Subscriptions |
| Zendesk | Support | API Token | Tickets, Users, Organizations |
| Intercom | Support | OAuth 2.0 | Contacts, Companies, Conversations |
| Marketo | Marketing | OAuth 2.0 | Leads, Lists, Programs, Activities |
| Google Ads | Advertising | OAuth 2.0 | Campaigns, Ad Groups, Performance |
| Facebook Ads | Advertising | OAuth 2.0 | Campaigns, Ad Sets, Insights |
| LinkedIn Ads | Advertising | OAuth 2.0 | Campaigns, Creatives, Analytics |
| Shopify | E-commerce | OAuth 2.0 | Orders, Products, Customers |
| GitHub | Developer | OAuth 2.0 / PAT | Repos, Issues, Pull Requests |
| Jira | Project Management | OAuth 2.0 / API Token | Issues, Projects, Sprints |
| Google Sheets | Productivity | OAuth 2.0 | Spreadsheet Tabs |
| Slack | Collaboration | OAuth 2.0 | Messages, Channels, Users |
How Loaders Work
1. Connect
Authenticate with the source application using OAuth 2.0, API keys, or access tokens. Zeotap securely stores credentials and handles token refresh automatically.
2. Discover
Once connected, Zeotap discovers the available objects and streams from the application’s API. You select which objects to sync — there’s no need to extract everything.
3. Map
Zeotap automatically maps the source application’s schema to warehouse-compatible table definitions. Each selected object becomes a table in your target schema. You can customize column names, types, and which fields to include or exclude.
4. Schedule
Configure a sync schedule that determines how frequently data is pulled. Options range from every 15 minutes to daily. Zeotap uses incremental sync by default, pulling only records that have changed since the last run.
5. Load
On each scheduled run, Zeotap extracts changed records from the source API, transforms them into the target schema, and writes them to your warehouse. Failed runs are automatically retried, and you can monitor progress from the Loaders dashboard.
Sync Modes
Loaders support three sync modes, chosen per stream (see Sync mode per stream):
| Mode | Description | Use Case |
|---|---|---|
| Full Refresh | Replaces the entire table with a fresh extract on each run | Small reference tables, lookup data, or when the source API doesn’t support change tracking |
| Incremental (append) | Pulls only new and updated records since the last successful sync, and appends them | Append-only data — event streams, logs, transactions that are never revised |
| Incremental (merge on key) | Pulls only new and updated records, then upserts them on a primary key | Records that get updated — contacts, accounts, orders — where you want one row per record holding its current values |
Both incremental modes use a cursor field (usually a timestamp like updated_at or modified_date) to track progress. Zeotap persists the cursor value between runs, so each execution picks up exactly where the last one left off.
Sync mode per stream
The sync mode is chosen per stream, because one source usually mixes stream shapes. A support desk’s tickets are edited for weeks and want merge on their id; the ticket field-history stream is never edited after it is written and wants a cheap append; a team-membership list whose rows are hard-deleted at the source has no change timestamp at all and wants a full refresh. One mode for all three would either duplicate history, never drop removed members, or re-read everything every run.
A loader has two settings that together decide each stream’s mode:
| Setting | What it is |
|---|---|
Default sync mode (sync_mode) | The loader-wide mode, used by every stream that has no mode of its own. |
Stream sync mode (stream_sync_modes) | An explicit mode for one stream, which wins over the default. |
A stream’s effective mode is its own mode if it has one, otherwise the loader’s default. That one rule is applied everywhere a mode matters — the run, merge keys, change detection, and Data Prep’s view of the landing table.
In the loader form, each selected stream has its own Sync mode select (Full refresh / Append / Merge). Changing the Default sync mode offers Apply to all streams, which sets every selected stream to the new default; without it, only streams with no mode of their own follow the change. When the streams do not all share one mode, the loader’s summary shows Mixed.
Connector defaults
Connectors can recommend a mode for each stream they know — for example, the Salesforce Service Cloud loader recommends Merge for cases and contacts, Append for CaseHistory and Full refresh for CaseTeamMember and CaseMilestone. The stream list pre-selects the recommendation, and the connector guides list it per stream.
Recommendations are applied once, when the loader is created, and saved on the loader as that stream’s mode. A loader never picks up a changed recommendation from a later Zeotap release by itself — what happens to your tables changes only when you change a mode. Loaders created before per-stream modes existed keep running every stream in the loader’s single mode until you edit them. A connector with no recommendation for a stream leaves it on the loader’s default.
Append on a stream without a cursor duplicates the table every run. A stream that cannot be read incrementally — the stream list shows it without a cursor — is read in full on every run whatever its mode. Under Append, each run adds another complete copy of it. The form warns when you choose Append for such a stream; use Full refresh (the table always matches the source) or Merge (one row per record, but rows deleted at the source are never removed).
Changing a stream’s mode takes effect on the next run and does not rewrite rows already in the table: switching a stream to Merge does not collapse duplicates an earlier Append left behind (run a full refresh of it once if you need that), and switching to Full refresh replaces the table on the next run.
Append or merge?
They read identically and differ only in what lands in the table.
Append writes a row per read. A contact updated on Monday, Wednesday and Friday arrives as three rows, and nothing in the table marks which is current — every model, audience and computed attribute reading that table sees all three. That is the right behaviour when each read really is a separate fact, and the wrong one when it is the same record again.
Merge upserts on a primary key, so that contact occupies one row holding its newest values. Within a single batch, the record with the highest cursor value wins.
Merge mode needs to know what identifies a row. Zeotap works it out in this order:
- The primary key you configure for the stream, on the loader. Multiple columns are supported for compositely-keyed data (for example
order_id+sku). - The key the connector reports. Database loaders read it from the source catalog; a composite key is used in full.
- An
idcolumn, if the stream has one.
A stream where none of these applies fails its run with a message naming the columns available, rather than quietly appending.
Merge mode upserts — it does not delete. A record removed at the source stays in the table until you reload it with a full refresh.
Two things to know before switching an existing loader over. It does not clean up the rows already in the table: what happens to records that are already duplicated there depends on the warehouse, so if you need one row per record from day one, run a full refresh once after switching. And the primary key you configure is trusted rather than checked — if it is not unique in the source, records sharing a value are collapsed to one and the others are dropped.
Target Warehouse Configuration
Loader data is written to a schema in your connected data warehouse. You configure the target location when creating a loader:
| Setting | Description | Example |
|---|---|---|
| Target Warehouse | The warehouse connection to write into | Production Snowflake |
| Schema | The target schema for loader tables | SALESFORCE_RAW, HUBSPOT_DATA |
| Table Prefix | Optional prefix applied to all table names | sf_, hs_ |
Zeotap creates tables automatically in the target schema. If a table already exists, what the loader does with it follows that stream’s sync mode: full refresh replaces its contents, append adds rows, and merge upserts them on the stream’s primary key. Merge mode also creates a short-lived staging table alongside the target (suffixed __ss_merge_...) for the duration of each run.
Reserved columns
Beyond the columns your source declares, every pull loader — every connector on this page that Zeotap calls on a schedule — stamps two of its own onto the rows it lands:
| Column | Type | What it holds |
|---|---|---|
_ss_loaded_at | TIMESTAMP | The moment Zeotap received the row — the platform’s own clock, not the source’s |
_ss_run_id | STRING | The loader run that landed the row, so a row can be traced back to its run and logs |
Both names are reserved. If one of your own streams declares a column called _ss_loaded_at or _ss_run_id, your column wins and that table is left unstamped — nothing of yours is overwritten, but the guarantees below do not apply to it.
They are ordinary, visible columns: they appear in the schema browser, in a model’s column picker, and in a Data Prep recipe like any other. Two things make _ss_loaded_at particularly useful:
- It only moves forward. One value is written per batch, and batches become visible in order, so
MAX(_ss_loaded_at)is exactly the boundary of what has landed. A source-sideupdated_athas no such property — a row edited three days ago can land today. This is why an incremental prepared table over a loader landing defaults its watermark to it. - It is a timestamp for sources that have none. Set it as the model’s timestamp column (Schema → the model → Entity Config) and the model can rank its own rows — golden-record Most Recent survivorship, and latest/earliest selection when a sync maps a column through a relationship.
Rows landed before this column existed, and rows in a table whose widening the warehouse refused, carry NULL. A NULL ranks last everywhere it is read, so it is never mistaken for a recent row.
Push loaders land a different column. The generic HTTP push loader and Braze Currents receive data rather than fetching it, so they do not run the pull pipeline that stamps the two columns above. Their tables carry received_at — a native TIMESTAMP holding the moment the event reached Zeotap, which serves the same purpose wherever this page recommends _ss_loaded_at.
Scheduling
Loaders run on configurable schedules:
| Interval | Description |
|---|---|
| Every 15 minutes | Near real-time for critical data |
| Hourly | Good balance of freshness and API usage |
| Every 6 hours | Suitable for most operational data |
| Daily | Best for reference data or high-volume extracts |
| Custom cron | Full cron expression for advanced scheduling |
All schedules are evaluated in UTC. You can pause and resume loaders at any time without losing cursor state.
Monitoring
Each loader run produces a detailed execution log:
- Status — Success, Failed, or Running
- Records extracted — Number of records pulled from the source API
- Records loaded — Number of records written to the warehouse
- Duration — Wall-clock time of the run
- Errors — Any API errors, rate limit hits, or schema conflicts
You can view run history from the Loaders dashboard or query the API for programmatic monitoring.
API Reference
Loaders are managed through the Zeotap REST API. Every path below is workspace-scoped — {id} is your workspace ID. See Base URL for your instance’s API base URL and Authentication for the required Authorization and X-Workspace-ID headers.
# List all loaders
GET /api/v1/workspaces/{id}/loaders
# Get a single loader
GET /api/v1/workspaces/{id}/loaders/{loaderId}
# Create a loader
POST /api/v1/workspaces/{id}/loaders
# Update a loader
PUT /api/v1/workspaces/{id}/loaders/{loaderId}
# Delete a loader
DELETE /api/v1/workspaces/{id}/loaders/{loaderId}
# Trigger a manual run
POST /api/v1/workspaces/{id}/loaders/{loaderId}/trigger
# Get run history
GET /api/v1/workspaces/{id}/loaders/{loaderId}/runs
# Get a single run
GET /api/v1/workspaces/{id}/loaders/{loaderId}/runs/{runId}
# Cancel a running load
POST /api/v1/workspaces/{id}/loaders/{loaderId}/runs/{runId}/cancel
# Test a loader's connection
POST /api/v1/workspaces/{id}/loaders/{loaderId}/test
# Discover the streams a loader can pull
POST /api/v1/workspaces/{id}/loaders/{loaderId}/discoverSee the API Reference for full request/response schemas.
Best Practices
- Use a dedicated schema — Write loader data into a separate schema (e.g.,
SALESFORCE_RAW) to keep it isolated from your curated models and analytics tables. - Start with incremental sync — Incremental mode is faster, cheaper, and puts less load on both the source API and your warehouse.
- Monitor API quotas — Some source applications have API rate limits. If you’re loading many objects at high frequency, check that your API plan supports the volume.
- Schedule off-peak — For large full-refresh loads, schedule runs during off-peak hours to minimize impact on your warehouse.
- Use table prefixes — If multiple loaders write to the same schema, use prefixes to avoid naming collisions and make tables easy to identify.
Next Steps
- Create your first loader
- Choose a connector guide: Salesforce | HubSpot | Stripe | Zendesk