Skip to Content
LoadersOverview

Loaders

Loaders pull data from SaaS applications and third-party services into your data warehouse on a recurring schedule. Instead of building and maintaining custom ETL pipelines, you configure a loader in Zeotap, and it handles extraction, schema mapping, and incremental synchronization automatically.

What Is a Loader?

A Loader is a managed ingestion pipeline that connects to a SaaS application’s API, extracts data from one or more objects (e.g., Salesforce Contacts, Stripe Subscriptions), and writes it into your data warehouse. Once the data is in your warehouse, it becomes available to Warehouses, Models, and the rest of the Zeotap platform.

Loaders complement Warehouses. While a Warehouse gives Zeotap read access to data already in your warehouse, a Loader brings external data into your warehouse in the first place.

Supported Connectors

Zeotap supports 15+ loader connectors across CRM, marketing automation, advertising, payments, support, and productivity categories.

ConnectorCategoryAuthenticationKey Objects
SalesforceCRMOAuth 2.0Contacts, Leads, Accounts, Opportunities
Salesforce Service CloudCRMOAuth 2.0Cases, Case History, Emails, Messaging Sessions, Knowledge
HubSpotCRMOAuth 2.0 / API KeyContacts, Companies, Deals
StripePaymentsAPI KeyCustomers, Charges, Subscriptions
ZendeskSupportAPI TokenTickets, Users, Organizations
IntercomSupportOAuth 2.0Contacts, Companies, Conversations
MarketoMarketingOAuth 2.0Leads, Lists, Programs, Activities
Google AdsAdvertisingOAuth 2.0Campaigns, Ad Groups, Performance
Facebook AdsAdvertisingOAuth 2.0Campaigns, Ad Sets, Insights
LinkedIn AdsAdvertisingOAuth 2.0Campaigns, Creatives, Analytics
ShopifyE-commerceOAuth 2.0Orders, Products, Customers
GitHubDeveloperOAuth 2.0 / PATRepos, Issues, Pull Requests
JiraProject ManagementOAuth 2.0 / API TokenIssues, Projects, Sprints
Google SheetsProductivityOAuth 2.0Spreadsheet Tabs
SlackCollaborationOAuth 2.0Messages, Channels, Users

How Loaders Work

1. Connect

Authenticate with the source application using OAuth 2.0, API keys, or access tokens. Zeotap securely stores credentials and handles token refresh automatically.

2. Discover

Once connected, Zeotap discovers the available objects and streams from the application’s API. You select which objects to sync — there’s no need to extract everything.

3. Map

Zeotap automatically maps the source application’s schema to warehouse-compatible table definitions. Each selected object becomes a table in your target schema. You can customize column names, types, and which fields to include or exclude.

4. Schedule

Configure a sync schedule that determines how frequently data is pulled. Options range from every 15 minutes to daily. Zeotap uses incremental sync by default, pulling only records that have changed since the last run.

5. Load

On each scheduled run, Zeotap extracts changed records from the source API, transforms them into the target schema, and writes them to your warehouse. Failed runs are automatically retried, and you can monitor progress from the Loaders dashboard.

Sync Modes

Loaders support three sync modes, chosen per stream (see Sync mode per stream):

ModeDescriptionUse Case
Full RefreshReplaces the entire table with a fresh extract on each runSmall reference tables, lookup data, or when the source API doesn’t support change tracking
Incremental (append)Pulls only new and updated records since the last successful sync, and appends themAppend-only data — event streams, logs, transactions that are never revised
Incremental (merge on key)Pulls only new and updated records, then upserts them on a primary keyRecords that get updated — contacts, accounts, orders — where you want one row per record holding its current values

Both incremental modes use a cursor field (usually a timestamp like updated_at or modified_date) to track progress. Zeotap persists the cursor value between runs, so each execution picks up exactly where the last one left off.

Sync mode per stream

The sync mode is chosen per stream, because one source usually mixes stream shapes. A support desk’s tickets are edited for weeks and want merge on their id; the ticket field-history stream is never edited after it is written and wants a cheap append; a team-membership list whose rows are hard-deleted at the source has no change timestamp at all and wants a full refresh. One mode for all three would either duplicate history, never drop removed members, or re-read everything every run.

A loader has two settings that together decide each stream’s mode:

SettingWhat it is
Default sync mode (sync_mode)The loader-wide mode, used by every stream that has no mode of its own.
Stream sync mode (stream_sync_modes)An explicit mode for one stream, which wins over the default.

A stream’s effective mode is its own mode if it has one, otherwise the loader’s default. That one rule is applied everywhere a mode matters — the run, merge keys, change detection, and Data Prep’s view of the landing table.

In the loader form, each selected stream has its own Sync mode select (Full refresh / Append / Merge). Changing the Default sync mode offers Apply to all streams, which sets every selected stream to the new default; without it, only streams with no mode of their own follow the change. When the streams do not all share one mode, the loader’s summary shows Mixed.

Connector defaults

Connectors can recommend a mode for each stream they know — for example, the Salesforce Service Cloud loader recommends Merge for cases and contacts, Append for CaseHistory and Full refresh for CaseTeamMember and CaseMilestone. The stream list pre-selects the recommendation, and the connector guides list it per stream.

Recommendations are applied once, when the loader is created, and saved on the loader as that stream’s mode. A loader never picks up a changed recommendation from a later Zeotap release by itself — what happens to your tables changes only when you change a mode. Loaders created before per-stream modes existed keep running every stream in the loader’s single mode until you edit them. A connector with no recommendation for a stream leaves it on the loader’s default.

Append on a stream without a cursor duplicates the table every run. A stream that cannot be read incrementally — the stream list shows it without a cursor — is read in full on every run whatever its mode. Under Append, each run adds another complete copy of it. The form warns when you choose Append for such a stream; use Full refresh (the table always matches the source) or Merge (one row per record, but rows deleted at the source are never removed).

Changing a stream’s mode takes effect on the next run and does not rewrite rows already in the table: switching a stream to Merge does not collapse duplicates an earlier Append left behind (run a full refresh of it once if you need that), and switching to Full refresh replaces the table on the next run.

Append or merge?

They read identically and differ only in what lands in the table.

Append writes a row per read. A contact updated on Monday, Wednesday and Friday arrives as three rows, and nothing in the table marks which is current — every model, audience and computed attribute reading that table sees all three. That is the right behaviour when each read really is a separate fact, and the wrong one when it is the same record again.

Merge upserts on a primary key, so that contact occupies one row holding its newest values. Within a single batch, the record with the highest cursor value wins.

Merge mode needs to know what identifies a row. Zeotap works it out in this order:

  1. The primary key you configure for the stream, on the loader. Multiple columns are supported for compositely-keyed data (for example order_id + sku).
  2. The key the connector reports. Database loaders read it from the source catalog; a composite key is used in full.
  3. An id column, if the stream has one.

A stream where none of these applies fails its run with a message naming the columns available, rather than quietly appending.

Merge mode upserts — it does not delete. A record removed at the source stays in the table until you reload it with a full refresh.

Two things to know before switching an existing loader over. It does not clean up the rows already in the table: what happens to records that are already duplicated there depends on the warehouse, so if you need one row per record from day one, run a full refresh once after switching. And the primary key you configure is trusted rather than checked — if it is not unique in the source, records sharing a value are collapsed to one and the others are dropped.

Target Warehouse Configuration

Loader data is written to a schema in your connected data warehouse. You configure the target location when creating a loader:

SettingDescriptionExample
Target WarehouseThe warehouse connection to write intoProduction Snowflake
SchemaThe target schema for loader tablesSALESFORCE_RAW, HUBSPOT_DATA
Table PrefixOptional prefix applied to all table namessf_, hs_

Zeotap creates tables automatically in the target schema. If a table already exists, what the loader does with it follows that stream’s sync mode: full refresh replaces its contents, append adds rows, and merge upserts them on the stream’s primary key. Merge mode also creates a short-lived staging table alongside the target (suffixed __ss_merge_...) for the duration of each run.

Reserved columns

Beyond the columns your source declares, every pull loader — every connector on this page that Zeotap calls on a schedule — stamps two of its own onto the rows it lands:

ColumnTypeWhat it holds
_ss_loaded_atTIMESTAMPThe moment Zeotap received the row — the platform’s own clock, not the source’s
_ss_run_idSTRINGThe loader run that landed the row, so a row can be traced back to its run and logs

Both names are reserved. If one of your own streams declares a column called _ss_loaded_at or _ss_run_id, your column wins and that table is left unstamped — nothing of yours is overwritten, but the guarantees below do not apply to it.

They are ordinary, visible columns: they appear in the schema browser, in a model’s column picker, and in a Data Prep recipe like any other. Two things make _ss_loaded_at particularly useful:

  • It only moves forward. One value is written per batch, and batches become visible in order, so MAX(_ss_loaded_at) is exactly the boundary of what has landed. A source-side updated_at has no such property — a row edited three days ago can land today. This is why an incremental prepared table over a loader landing defaults its watermark to it.
  • It is a timestamp for sources that have none. Set it as the model’s timestamp column (Schema → the model → Entity Config) and the model can rank its own rows — golden-record Most Recent survivorship, and latest/earliest selection when a sync maps a column through a relationship.

Rows landed before this column existed, and rows in a table whose widening the warehouse refused, carry NULL. A NULL ranks last everywhere it is read, so it is never mistaken for a recent row.

Push loaders land a different column. The generic HTTP push loader and Braze Currents receive data rather than fetching it, so they do not run the pull pipeline that stamps the two columns above. Their tables carry received_at — a native TIMESTAMP holding the moment the event reached Zeotap, which serves the same purpose wherever this page recommends _ss_loaded_at.

Scheduling

Loaders run on configurable schedules:

IntervalDescription
Every 15 minutesNear real-time for critical data
HourlyGood balance of freshness and API usage
Every 6 hoursSuitable for most operational data
DailyBest for reference data or high-volume extracts
Custom cronFull cron expression for advanced scheduling

All schedules are evaluated in UTC. You can pause and resume loaders at any time without losing cursor state.

Monitoring

Each loader run produces a detailed execution log:

  • Status — Success, Failed, or Running
  • Records extracted — Number of records pulled from the source API
  • Records loaded — Number of records written to the warehouse
  • Duration — Wall-clock time of the run
  • Errors — Any API errors, rate limit hits, or schema conflicts

You can view run history from the Loaders dashboard or query the API for programmatic monitoring.

API Reference

Loaders are managed through the Zeotap REST API. Every path below is workspace-scoped — {id} is your workspace ID. See Base URL for your instance’s API base URL and Authentication for the required Authorization and X-Workspace-ID headers.

# List all loaders GET /api/v1/workspaces/{id}/loaders # Get a single loader GET /api/v1/workspaces/{id}/loaders/{loaderId} # Create a loader POST /api/v1/workspaces/{id}/loaders # Update a loader PUT /api/v1/workspaces/{id}/loaders/{loaderId} # Delete a loader DELETE /api/v1/workspaces/{id}/loaders/{loaderId} # Trigger a manual run POST /api/v1/workspaces/{id}/loaders/{loaderId}/trigger # Get run history GET /api/v1/workspaces/{id}/loaders/{loaderId}/runs # Get a single run GET /api/v1/workspaces/{id}/loaders/{loaderId}/runs/{runId} # Cancel a running load POST /api/v1/workspaces/{id}/loaders/{loaderId}/runs/{runId}/cancel # Test a loader's connection POST /api/v1/workspaces/{id}/loaders/{loaderId}/test # Discover the streams a loader can pull POST /api/v1/workspaces/{id}/loaders/{loaderId}/discover

See the API Reference for full request/response schemas.

Best Practices

  • Use a dedicated schema — Write loader data into a separate schema (e.g., SALESFORCE_RAW) to keep it isolated from your curated models and analytics tables.
  • Start with incremental sync — Incremental mode is faster, cheaper, and puts less load on both the source API and your warehouse.
  • Monitor API quotas — Some source applications have API rate limits. If you’re loading many objects at high frequency, check that your API plan supports the volume.
  • Schedule off-peak — For large full-refresh loads, schedule runs during off-peak hours to minimize impact on your warehouse.
  • Use table prefixes — If multiple loaders write to the same schema, use prefixes to avoid naming collisions and make tables easy to identify.

Next Steps

Last updated on