Skip to Content
LoadersSalesforce

Salesforce Loader

The Salesforce loader pulls Sales Cloud records — contacts, leads, accounts, opportunities, campaigns, tasks and events — from your Salesforce org into your data warehouse using the Salesforce REST API. It syncs incrementally on each record’s LastModifiedDate and lands a fixed set of commonly used standard fields for each object.

For service data — cases, case history, emails, messaging sessions, entitlements, knowledge — use the Salesforce Service Cloud loader. Both loaders can point at the same org.

Prerequisites

  • A Salesforce org with API access (Enterprise, Unlimited, Performance or Developer Edition; Professional Edition only with the API add-on). Production and Developer Edition orgs only — sandbox orgs are not supported yet.
  • A Salesforce user with API access (the API Enabled permission must be active) and read access to the objects you want to load. A field the user cannot see is loaded as NULL.
  • A connected Warehouse (target warehouse) with write permissions on the target schema

Authentication

The Salesforce loader uses OAuth 2.0 for authentication, through Zeotap’s own Salesforce connected app.

OAuth 2.0 Setup

  1. In Zeotap, click Add Loader and select Salesforce
  2. Click Connect with Salesforce
  3. You’ll be redirected to the Salesforce login page
  4. Log in with your Salesforce credentials and click Allow to grant Zeotap access
  5. You’ll be redirected back to Zeotap with the connection established

Zeotap requests the following OAuth scopes:

ScopePurpose
apiAccess Salesforce REST API
refresh_tokenMaintain a long-lived connection without re-authenticating
offline_accessRefresh tokens when the access token expires

Zeotap automatically refreshes the access token before it expires, and refreshes and retries once if Salesforce invalidates the session in the middle of a run. If the refresh token is revoked (e.g., user password reset, admin action), you’ll need to re-authenticate.

Sandbox Connections

Sandbox orgs are not supported yet. Sign-in always goes through login.salesforce.com, which does not accept sandbox credentials. The loader has a Sandbox Environment toggle, but it currently has no effect — leave it off.

Available Objects

The Salesforce loader syncs the following standard objects. Each lands in a table named after the object, with Id as the primary key and LastModifiedDate as the incremental cursor.

ObjectAPI NameDescriptionSelected by DefaultDefault sync mode
ContactContactIndividual people associated with accountsYesMerge
LeadLeadProspective customers not yet associated with an accountYesMerge
AccountAccountCompanies or organizationsYesMerge
OpportunityOpportunitySales deals with stages, amounts, and close datesYesMerge
CampaignCampaignMarketing campaigns and their metadataNoMerge
TaskTaskActivities like calls, emails, and to-dosNoMerge
EventEventCalendar events and meetingsNoMerge

Columns

Each object lands a fixed set of commonly used standard fields — for example FirstName, LastName, Email, AccountId, OwnerId, the mailing address components, CreatedDate, LastModifiedDate and IsDeleted on Contact; StageName, Amount, CloseDate, IsWon on Opportunity. The column set is the same in every org:

  • A field your org does not expose — a feature is off, the configured API version predates it, or field-level security hides it from the connecting user — is left out of the query and loaded as NULL, rather than failing the run.
  • Custom fields and custom objects are not loaded. For Account and Contact, the Salesforce Service Cloud loader also loads every readable custom field in a CustomFields JSON column.

Configuration

FieldTypeRequiredDescription
Instance URLTextYesYour org’s My Domain URL, e.g. https://yourorg.my.salesforce.com. Must start with https:// and end in .salesforce.com.
API VersionTextNoSalesforce REST API version. Default: v59.0. Must not be newer than your org supports.
Sandbox EnvironmentToggleNoHas no effect yet — sandbox orgs are not supported.
Include deleted recordsToggleNoDefault: off. When on, the Contact, Lead and Account streams also return soft-deleted records (those in the Recycle Bin) as rows with IsDeleted = true. See Deleted records.

Target schema, stream selection, sync mode (per stream — see Sync Modes) and schedule are set on the loader as for every connector — see Creating a Loader.

Sync Modes

ModeSupported
Full RefreshYes
Incremental (append)Yes
Incremental (merge on key)Yes — the default for every stream

The sync mode is set per stream. A new loader gives every stream Incremental (merge on key), which upserts on Id so each record occupies one row holding its latest values — the right shape for all seven objects, which are edited after they are created. The stream list shows the mode pre-selected per stream; change any of them before or after you save, for example to Incremental (append) to keep one row per change as a history, or to Full Refresh to rebuild a table from scratch every run (which costs a full read of the object each time). Streams you do not set follow the loader’s default sync mode — see Loaders → Sync mode per stream.

Loaders created before per-stream sync modes existed keep running every stream in the loader’s single mode until you edit them.

How It Works

  • Incremental cursor. Each incremental run reads records whose LastModifiedDate is after the last run’s saved cursor and no later than two minutes before the run started, ordered by LastModifiedDate and then Id. Holding back the last two minutes means a record committed in the same second the loader reads is picked up by the next run instead of being skipped.
  • No skipped ties. Many records can share one LastModifiedDate (a bulk update touches thousands in the same second). The cursor saved during a run is the last timestamp it has fully delivered, so a run that stops partway through such a group resumes at the start of the group rather than after it. A resumed run may re-read a few records; under merge mode they collapse to one row.
  • LastModifiedDate, not SystemModstamp. LastModifiedDate changes when a user or integration edits a record. Changes Salesforce makes on its own — for example a roll-up summary recalculation — update only SystemModstamp, so they are picked up the next time the record is edited, or by a Full Refresh.
  • Typed values. Datetime fields land as real UTC timestamps on every warehouse, numbers as numbers and checkboxes as booleans.
  • REST query API. The loader reads with SOQL through the REST query endpoint, following Salesforce’s result pages and writing each page to the warehouse as it arrives. It does not use the Bulk API.

Deleted Records

With Include deleted records on, the Contact, Lead and Account streams are read with Salesforce’s queryAll endpoint, which also returns records in the Recycle Bin; a deleted record arrives as a row with IsDeleted = true (and, in merge mode, updates the existing row in place). Filter on IsDeleted = false in your models to see only live records. Opportunity, Campaign, Task and Event have no IsDeleted column in their tables and always use the standard query, which never returns deleted records.

Salesforce keeps deleted records in the Recycle Bin for 15 days. Run the loader at least that often, or deletions purged from the bin in between are never seen. With the option off (the default, and the behaviour of every loader created before the option existed), deleted records are not returned and rows already loaded for them stay in the table unchanged.

Scheduling Notes

  • Rate limits: Salesforce enforces API call limits based on your edition and licence count. Enterprise and Unlimited editions provide 100,000 API calls per rolling 24 hours plus an allowance per user licence. Each run of each stream costs one describe call plus roughly one call per page of results, so the first full-history run of a large object is the expensive one.
  • Schedule: Hourly or daily schedules suit most orgs. Keep the interval under 15 days if you rely on Include deleted records.

Schema Mapping

Salesforce field types are mapped to warehouse-compatible types:

Salesforce TypeWarehouse TypeNotes
id, referenceSTRING / VARCHAR18-character Salesforce ID
string, textarea, picklist, email, phone, urlSTRING / VARCHAR
booleanBOOLEAN
intBIGINT (INT64 on BigQuery)
double, currency, percentDOUBLE (FLOAT64 on BigQuery)
dateDATE
datetimeTIMESTAMPNormalized to UTC

Troubleshooting

IssueSolution
”INVALID_SESSION_ID” or HTTP 401 that persistsZeotap refreshes the token and retries once automatically. If runs keep failing, re-authenticate by clicking Reconnect on the loader detail page
”REQUEST_LIMIT_EXCEEDED”You’ve hit Salesforce’s daily API limit. Reduce sync frequency or number of objects, or upgrade your Salesforce edition
A column is NULL in every rowYour org does not expose that field to the connecting user (feature off, field-level security, or an API version older than the field). Grant read access or raise the API version
HTTP 404 on the connectivity testThe configured API version is newer than your org supports. Use v59.0
Custom fields or custom objects missingNot supported by this loader — it loads a fixed set of standard fields on seven standard objects
Sandbox connection failsSandbox orgs are not supported yet; the Sandbox toggle has no effect
Deleted records never appearTurn on Include deleted records (Contact, Lead and Account only) and run at least every 15 days. Records permanently removed from the Recycle Bin before a run cannot be captured
Datetime columns were NULL in tables loaded before October 2026Earlier versions of the loader wrote Salesforce datetimes in a format some warehouses could not parse on bulk loads. New runs land them correctly; run a Full Refresh once to repair rows already loaded
Slow initial syncLarge orgs with millions of records may take several hours for the initial backfill. Subsequent incremental syncs will be much faster

Next Steps

Last updated on