Salesforce Loader
The Salesforce loader pulls Sales Cloud records — contacts, leads, accounts, opportunities, campaigns, tasks and events — from your Salesforce org into your data warehouse using the Salesforce REST API. It syncs incrementally on each record’s LastModifiedDate and lands a fixed set of commonly used standard fields for each object.
For service data — cases, case history, emails, messaging sessions, entitlements, knowledge — use the Salesforce Service Cloud loader. Both loaders can point at the same org.
Prerequisites
- A Salesforce org with API access (Enterprise, Unlimited, Performance or Developer Edition; Professional Edition only with the API add-on). Production and Developer Edition orgs only — sandbox orgs are not supported yet.
- A Salesforce user with API access (the API Enabled permission must be active) and read access to the objects you want to load. A field the user cannot see is loaded as
NULL. - A connected Warehouse (target warehouse) with write permissions on the target schema
Authentication
The Salesforce loader uses OAuth 2.0 for authentication, through Zeotap’s own Salesforce connected app.
OAuth 2.0 Setup
- In Zeotap, click Add Loader and select Salesforce
- Click Connect with Salesforce
- You’ll be redirected to the Salesforce login page
- Log in with your Salesforce credentials and click Allow to grant Zeotap access
- You’ll be redirected back to Zeotap with the connection established
Zeotap requests the following OAuth scopes:
| Scope | Purpose |
|---|---|
api | Access Salesforce REST API |
refresh_token | Maintain a long-lived connection without re-authenticating |
offline_access | Refresh tokens when the access token expires |
Zeotap automatically refreshes the access token before it expires, and refreshes and retries once if Salesforce invalidates the session in the middle of a run. If the refresh token is revoked (e.g., user password reset, admin action), you’ll need to re-authenticate.
Sandbox Connections
Sandbox orgs are not supported yet. Sign-in always goes through login.salesforce.com, which does not accept sandbox credentials. The loader has a Sandbox Environment toggle, but it currently has no effect — leave it off.
Available Objects
The Salesforce loader syncs the following standard objects. Each lands in a table named after the object, with Id as the primary key and LastModifiedDate as the incremental cursor.
| Object | API Name | Description | Selected by Default | Default sync mode |
|---|---|---|---|---|
| Contact | Contact | Individual people associated with accounts | Yes | Merge |
| Lead | Lead | Prospective customers not yet associated with an account | Yes | Merge |
| Account | Account | Companies or organizations | Yes | Merge |
| Opportunity | Opportunity | Sales deals with stages, amounts, and close dates | Yes | Merge |
| Campaign | Campaign | Marketing campaigns and their metadata | No | Merge |
| Task | Task | Activities like calls, emails, and to-dos | No | Merge |
| Event | Event | Calendar events and meetings | No | Merge |
Columns
Each object lands a fixed set of commonly used standard fields — for example FirstName, LastName, Email, AccountId, OwnerId, the mailing address components, CreatedDate, LastModifiedDate and IsDeleted on Contact; StageName, Amount, CloseDate, IsWon on Opportunity. The column set is the same in every org:
- A field your org does not expose — a feature is off, the configured API version predates it, or field-level security hides it from the connecting user — is left out of the query and loaded as
NULL, rather than failing the run. - Custom fields and custom objects are not loaded. For Account and Contact, the Salesforce Service Cloud loader also loads every readable custom field in a
CustomFieldsJSON column.
Configuration
| Field | Type | Required | Description |
|---|---|---|---|
| Instance URL | Text | Yes | Your org’s My Domain URL, e.g. https://yourorg.my.salesforce.com. Must start with https:// and end in .salesforce.com. |
| API Version | Text | No | Salesforce REST API version. Default: v59.0. Must not be newer than your org supports. |
| Sandbox Environment | Toggle | No | Has no effect yet — sandbox orgs are not supported. |
| Include deleted records | Toggle | No | Default: off. When on, the Contact, Lead and Account streams also return soft-deleted records (those in the Recycle Bin) as rows with IsDeleted = true. See Deleted records. |
Target schema, stream selection, sync mode (per stream — see Sync Modes) and schedule are set on the loader as for every connector — see Creating a Loader.
Sync Modes
| Mode | Supported |
|---|---|
| Full Refresh | Yes |
| Incremental (append) | Yes |
| Incremental (merge on key) | Yes — the default for every stream |
The sync mode is set per stream. A new loader gives every stream Incremental (merge on key), which upserts on Id so each record occupies one row holding its latest values — the right shape for all seven objects, which are edited after they are created. The stream list shows the mode pre-selected per stream; change any of them before or after you save, for example to Incremental (append) to keep one row per change as a history, or to Full Refresh to rebuild a table from scratch every run (which costs a full read of the object each time). Streams you do not set follow the loader’s default sync mode — see Loaders → Sync mode per stream.
Loaders created before per-stream sync modes existed keep running every stream in the loader’s single mode until you edit them.
How It Works
- Incremental cursor. Each incremental run reads records whose
LastModifiedDateis after the last run’s saved cursor and no later than two minutes before the run started, ordered byLastModifiedDateand thenId. Holding back the last two minutes means a record committed in the same second the loader reads is picked up by the next run instead of being skipped. - No skipped ties. Many records can share one
LastModifiedDate(a bulk update touches thousands in the same second). The cursor saved during a run is the last timestamp it has fully delivered, so a run that stops partway through such a group resumes at the start of the group rather than after it. A resumed run may re-read a few records; under merge mode they collapse to one row. LastModifiedDate, notSystemModstamp.LastModifiedDatechanges when a user or integration edits a record. Changes Salesforce makes on its own — for example a roll-up summary recalculation — update onlySystemModstamp, so they are picked up the next time the record is edited, or by a Full Refresh.- Typed values. Datetime fields land as real UTC timestamps on every warehouse, numbers as numbers and checkboxes as booleans.
- REST query API. The loader reads with SOQL through the REST query endpoint, following Salesforce’s result pages and writing each page to the warehouse as it arrives. It does not use the Bulk API.
Deleted Records
With Include deleted records on, the Contact, Lead and Account streams are read with Salesforce’s queryAll endpoint, which also returns records in the Recycle Bin; a deleted record arrives as a row with IsDeleted = true (and, in merge mode, updates the existing row in place). Filter on IsDeleted = false in your models to see only live records. Opportunity, Campaign, Task and Event have no IsDeleted column in their tables and always use the standard query, which never returns deleted records.
Salesforce keeps deleted records in the Recycle Bin for 15 days. Run the loader at least that often, or deletions purged from the bin in between are never seen. With the option off (the default, and the behaviour of every loader created before the option existed), deleted records are not returned and rows already loaded for them stay in the table unchanged.
Scheduling Notes
- Rate limits: Salesforce enforces API call limits based on your edition and licence count. Enterprise and Unlimited editions provide 100,000 API calls per rolling 24 hours plus an allowance per user licence. Each run of each stream costs one describe call plus roughly one call per page of results, so the first full-history run of a large object is the expensive one.
- Schedule: Hourly or daily schedules suit most orgs. Keep the interval under 15 days if you rely on Include deleted records.
Schema Mapping
Salesforce field types are mapped to warehouse-compatible types:
| Salesforce Type | Warehouse Type | Notes |
|---|---|---|
id, reference | STRING / VARCHAR | 18-character Salesforce ID |
string, textarea, picklist, email, phone, url | STRING / VARCHAR | |
boolean | BOOLEAN | |
int | BIGINT (INT64 on BigQuery) | |
double, currency, percent | DOUBLE (FLOAT64 on BigQuery) | |
date | DATE | |
datetime | TIMESTAMP | Normalized to UTC |
Troubleshooting
| Issue | Solution |
|---|---|
”INVALID_SESSION_ID” or HTTP 401 that persists | Zeotap refreshes the token and retries once automatically. If runs keep failing, re-authenticate by clicking Reconnect on the loader detail page |
| ”REQUEST_LIMIT_EXCEEDED” | You’ve hit Salesforce’s daily API limit. Reduce sync frequency or number of objects, or upgrade your Salesforce edition |
A column is NULL in every row | Your org does not expose that field to the connecting user (feature off, field-level security, or an API version older than the field). Grant read access or raise the API version |
HTTP 404 on the connectivity test | The configured API version is newer than your org supports. Use v59.0 |
| Custom fields or custom objects missing | Not supported by this loader — it loads a fixed set of standard fields on seven standard objects |
| Sandbox connection fails | Sandbox orgs are not supported yet; the Sandbox toggle has no effect |
| Deleted records never appear | Turn on Include deleted records (Contact, Lead and Account only) and run at least every 15 days. Records permanently removed from the Recycle Bin before a run cannot be captured |
Datetime columns were NULL in tables loaded before October 2026 | Earlier versions of the loader wrote Salesforce datetimes in a format some warehouses could not parse on bulk loads. New runs land them correctly; run a Full Refresh once to repair rows already loaded |
| Slow initial sync | Large orgs with millions of records may take several hours for the initial backfill. Subsequent incremental syncs will be much faster |
Next Steps
- Create a model to transform your raw Salesforce data
- Build an audience using Salesforce contacts and leads
- Sync audiences back to Salesforce as a destination
- Load cases, emails and other service objects with the Salesforce Service Cloud loader