Databricks
This guide covers how to configure Databricks as a warehouse in Zeotap, including workspace setup, SQL warehouse configuration, authentication, and required permissions.
Prerequisites
- A Databricks workspace (AWS, Azure, or GCP)
- A SQL warehouse or all-purpose cluster running in the workspace
- A personal access token (or service principal) with read access to your data and write access to Zeotap’s dedicated working schemas (see Required Permissions)
- Network access from Zeotap to your Databricks workspace
Connection Configuration
Required Fields
| Field | Description | Example |
|---|---|---|
| Host | The Databricks workspace URL (without https://) | mycompany.cloud.databricks.com |
| HTTP Path | The HTTP path for the SQL warehouse or cluster | /sql/1.0/warehouses/abc123def456 |
| Token | A Databricks personal access token | dapi1234567890abcdef |
| Catalog | The Unity Catalog name (or hive_metastore for legacy) | main |
| Schema | The default schema to query | default |
Host
The workspace URL is visible in your browser’s address bar when logged in to Databricks. Use just the hostname without https://:
# AWS
mycompany.cloud.databricks.com
# Azure
adb-1234567890123456.7.azuredatabricks.net
# GCP
mycompany.gcp.databricks.comHTTP Path
The HTTP path identifies which compute resource to use for running queries. You can find it in the Databricks UI:
For SQL Warehouses (recommended):
- Go to SQL Warehouses in the left sidebar
- Click on your warehouse
- Go to the Connection Details tab
- Copy the HTTP Path value
The format is typically:
/sql/1.0/warehouses/<warehouse-id>For All-Purpose Clusters:
- Go to Compute in the left sidebar
- Click on your cluster
- Go to Advanced Options > JDBC/ODBC
- Copy the HTTP Path value
The format is typically:
/sql/protocolv1/o/<org-id>/<cluster-id>SQL Warehouses are recommended over all-purpose clusters for Zeotap because they start faster, scale automatically, and provide better cost isolation.
Catalog and Schema
If your Databricks workspace uses Unity Catalog, specify the catalog name (e.g., main, production). If you’re using the legacy Hive Metastore, set the catalog to hive_metastore.
-- Unity Catalog: catalog.schema.table
SELECT * FROM main.customer_data.users
-- Hive Metastore: schema.table (catalog is hive_metastore)
SELECT * FROM customer_data.usersDatabricks is case-insensitive for schema and table names, so Customer_Data and customer_data reference the same schema.
Authentication
Personal Access Token
Zeotap uses Databricks personal access tokens (PATs) for authentication.
To create a PAT:
- In the Databricks workspace, click your username in the top-right corner
- Select Settings
- Go to Developer > Access Tokens
- Click Generate New Token
- Set a description (e.g., “Zeotap CDP”) and expiration
- Copy the token — it is only shown once
Token format: dapi1234567890abcdef1234567890abcdefToken expiration: Set a reasonable expiration (e.g., 90 days) and establish a rotation process. When the token expires, update the source configuration in Zeotap and re-test the connection.
Service Principal (Alternative)
For production environments, consider using a Databricks service principal instead of a personal access token:
- Create a service principal in Account Console > User Management
- Generate an OAuth secret or PAT for the service principal
- Grant the service principal access to the required catalog, schema, and tables
Required Permissions
Zeotap needs read access to your source data, plus the ability to create and write the operational schemas it manages inside the connection’s catalog.
Unity Catalog Permissions
-- Grant catalog access
GRANT USE CATALOG ON CATALOG main TO `zeotap-service-principal`;
-- Grant schema access
GRANT USE SCHEMA ON SCHEMA main.customer_data TO `zeotap-service-principal`;
-- Grant table read access
GRANT SELECT ON SCHEMA main.customer_data TO `zeotap-service-principal`;Platform Schemas (Write Access)
Zeotap keeps all of its working state in dedicated schemas inside the connected catalog — your own tables are never modified. Read access to your data is sufficient everywhere else; write access is only needed on these schemas:
| Schema | Purpose | Needed for |
|---|---|---|
cdp_planner | Temporary computation tables | All functionality |
cdp_audit | Sync run audit logs | All functionality |
cdp_journey | Orchestration execution state and send logs | Orchestrations |
cdp_identity | Identity resolution output tables (default schema; configurable per identity graph) | Identity resolution |
cdp_raw | Landing tables for data pulled in by loaders (default schema; configurable per loader) | Loaders |
audit_logs | Delivery and observability event logs | Event delivery / observability logging |
cdp_prep | Prepared tables built by Data Prep, plus their transient rebuild siblings | Only if Data Prep is enabled for the workspace |
cdp_metadata | The audiences table written by the metadata export | Only if the metadata export is enabled for the workspace |
The simplest setup grants the principal permission to create schemas in the catalog — Zeotap creates each schema on first use, and as its creator it owns it, so no further grants are needed:
GRANT CREATE SCHEMA ON CATALOG main TO `zeotap-service-principal`;Alternatively, pre-create the schemas and grant the connecting principal full access to them:
CREATE SCHEMA IF NOT EXISTS main.cdp_planner;
CREATE SCHEMA IF NOT EXISTS main.cdp_audit;
CREATE SCHEMA IF NOT EXISTS main.cdp_journey;
CREATE SCHEMA IF NOT EXISTS main.cdp_identity;
CREATE SCHEMA IF NOT EXISTS main.cdp_raw;
CREATE SCHEMA IF NOT EXISTS main.audit_logs;
GRANT ALL PRIVILEGES ON SCHEMA main.cdp_planner TO `zeotap-service-principal`;
GRANT USE SCHEMA ON SCHEMA main.cdp_planner TO `zeotap-service-principal`;
GRANT ALL PRIVILEGES ON SCHEMA main.cdp_audit TO `zeotap-service-principal`;
GRANT USE SCHEMA ON SCHEMA main.cdp_audit TO `zeotap-service-principal`;
GRANT ALL PRIVILEGES ON SCHEMA main.cdp_journey TO `zeotap-service-principal`;
GRANT USE SCHEMA ON SCHEMA main.cdp_journey TO `zeotap-service-principal`;
GRANT ALL PRIVILEGES ON SCHEMA main.cdp_identity TO `zeotap-service-principal`;
GRANT USE SCHEMA ON SCHEMA main.cdp_identity TO `zeotap-service-principal`;
GRANT ALL PRIVILEGES ON SCHEMA main.cdp_raw TO `zeotap-service-principal`;
GRANT USE SCHEMA ON SCHEMA main.cdp_raw TO `zeotap-service-principal`;
GRANT ALL PRIVILEGES ON SCHEMA main.audit_logs TO `zeotap-service-principal`;
GRANT USE SCHEMA ON SCHEMA main.audit_logs TO `zeotap-service-principal`;cdp_prep — only if Data Prep is enabled for the workspace
Data Prep is a per-workspace entitlement, off by default. A workspace that has it builds prepared tables into one more schema, cdp_prep, and its connection test then carries a seventh write probe (write_prep). Until Data Prep is enabled nothing creates or reads this schema, so there is nothing to grant.
CREATE SCHEMA IF NOT EXISTS main.cdp_prep;
GRANT ALL PRIVILEGES ON SCHEMA main.cdp_prep TO `zeotap-service-principal`;
GRANT USE SCHEMA ON SCHEMA main.cdp_prep TO `zeotap-service-principal`;ALL PRIVILEGES is the same grant the six schemas above take, and every part of it is load-bearing for a prepared table — which, unlike the other schemas’ tables, is REBUILT and, when it is incremental, updated in place: CREATE TABLE for the table and its transient <slug>__next / <slug>__delta siblings, MODIFY for the appends, merges and delete_when deletes, SELECT for the delta computation and for the Models that read the result, and the ability to drop the siblings once the swap is done.
Revoking the entitlement does not drop the schema or its tables, so this grant can be withdrawn at your convenience rather than urgently.
cdp_metadata — only if the metadata export is enabled for the workspace
The metadata export is a per-workspace entitlement, off by default, that keeps the workspace’s audience catalogue as the table cdp_metadata.audiences. When it is enabled for a workspace, the connection test of the source chosen to hold the table carries one more write probe (write_metadata); the workspace’s other sources are tested exactly as before. Until then nothing creates or reads this schema, so there is nothing to grant.
CREATE SCHEMA IF NOT EXISTS main.cdp_metadata;
GRANT ALL PRIVILEGES ON SCHEMA main.cdp_metadata TO `zeotap-service-principal`;
GRANT USE SCHEMA ON SCHEMA main.cdp_metadata TO `zeotap-service-principal`;Each export creates audiences__next, inserts into it, replaces audiences with it and drops it, so it needs the same ALL PRIVILEGES as the schemas above. Give readers USE SCHEMA and SELECT on the schema rather than on the table, since the table is replaced by every export.
Revoking the entitlement does not drop the schema or its table.
The connection test verifies write access to every one of these schemas — mirroring how the runtime creates them — and reports the exact statements to run if a step fails. A principal that can only read source data passes authentication but fails the write steps, which would otherwise surface later as a permission error on its first orchestration, identity, loader, or observability run.
Optional: MODIFY on event tables (faster orchestration entry)
This one grant is optional, is on your own tables rather than on Zeotap’s schemas, and nothing fails without it.
A reactive orchestration narrows each evaluation to the people with a new event since the last one, which it decides from the event’s timestamp. That misses a late arrival: a row written today but carrying last week’s timestamp, which a backfill, an offline mobile app flushing its queue, or a partner file that arrives a day late all produce routinely. Such a person enters the orchestration on its next full reconciliation pass instead of within minutes.
Granting MODIFY on the event tables your orchestrations read lets Zeotap turn
on Delta Change Data Feed for them, which answers the other question — which
rows appeared since the last evaluation — so a late arrival enters on the next
wake:
-- Only the event tables your orchestrations read; nothing else needs MODIFY.
GRANT MODIFY ON TABLE main.customer_data.events TO `zeotap-service-principal`;What the grant is actually used for, in full: Zeotap runs
ALTER TABLE … SET TBLPROPERTIES (delta.enableChangeDataFeed = true) once per
table, and never writes rows to it. Enabling Change Data Feed is a permanent
change to your table: from then on Delta also records change data for every
write, which occupies storage you own. Both are worth knowing before granting
it.
Without the grant, orchestrations behave exactly as they did before — the feed is recorded as unavailable with Databricks’ own message, and the reconciliation pass still admits late arrivals. The feature is also off by default per workspace; ask your Zeotap administrator to enable it once the grant is in place.
Hive Metastore Permissions
-- Grant database access
GRANT USAGE ON DATABASE customer_data TO `cdp_user`;
-- Grant table read access
GRANT SELECT ON DATABASE customer_data TO `cdp_user`;On the legacy Hive Metastore the platform schemas above are databases; grant CREATE at the metastore level (or pre-create the six databases and grant the user full access to each).
SQL Warehouse Access
The user or service principal must also have Can Use permission on the SQL warehouse:
- Go to SQL Warehouses
- Click on the warehouse
- Go to the Permissions tab
- Add the user/service principal with Can Use permission
Bulk Loading
When Zeotap loads or forwards data into Databricks, it first stages the records as files in a Google Cloud Storage (GCS) bucket. How those records then reach your tables depends on which cloud your Databricks workspace runs on, and Zeotap picks the right path automatically from the workspace hostname.
| Workspace | Default path | What you configure |
|---|---|---|
Azure (*.azuredatabricks.net) | Zeotap reads the staged files and writes the rows over your SQL warehouse | Nothing |
AWS (*.cloud.databricks.com) | Same | Nothing |
GCP (*.gcp.databricks.com) | Your warehouse reads the staging bucket directly with read_files() — faster | A Unity Catalog storage credential (optional; without it, the Azure/AWS path is used) |
Only a Databricks-on-GCP workspace can be given access to a GCS bucket:
Unity Catalog issues a GCP service-account storage credential there, while on
Azure it governs abfss:// locations and on AWS s3:// ones. On those clouds
there is nothing to set up — Zeotap relays the rows itself, and this works
with any catalog, including hive_metastore.
Optional: direct GCS access on a GCP workspace
On a Databricks-on-GCP workspace you can let the warehouse read the staging
bucket itself, which is faster than relaying the rows. It requires Unity
Catalog — the legacy hive_metastore catalog cannot govern GCS access this
way — so set the connection’s Catalog to a Unity Catalog catalog.
-
A Unity Catalog storage credential for GCS. A Databricks admin creates this in Unity Catalog. When created, Databricks generates a managed Google service-account email that looks like
db-uc-credential-xxxxx@<region>.iam.gserviceaccount.com. Paste that email into the GCS Service Account field of the Zeotap Databricks connection. This is the service account Databricks impersonates to read GCS — it is not a key you download. -
An external location for the staging bucket, bound to the storage credential above. Zeotap auto-creates this external location for each staging bucket at load time if the connecting principal (the token or service principal used in the connection) has the CREATE EXTERNAL LOCATION privilege on the metastore. If it does not, a Databricks admin must pre-create the external location manually under Catalog → External Data → External Locations, pointing it at the staging bucket URL (
gs://<bucket>) and using the storage credential above. -
Bucket read access for the credential’s service account. The storage credential’s service account needs read access on the staging bucket. Zeotap grants this automatically.
Leaving the GCS Service Account field empty is always safe: loading falls back to the relayed path rather than failing.
Overriding the choice
The Bulk Load Mode connection field pins the path instead of inferring it:
| Value | Behaviour |
|---|---|
auto (default) | Chooses from the workspace hostname and whether a GCS service account is set |
external_location | Always have the warehouse read gs:// directly |
insert | Always relay the rows through Zeotap |
Set it only to work around a specific problem — for example pinning insert on
a GCP workspace whose external location has been removed, or
external_location on a workspace reached through a custom domain that
Zeotap cannot classify.
Data Types
| Databricks Type | Zeotap Handling |
|---|---|
STRING | Mapped as text |
INT, BIGINT, DOUBLE, DECIMAL | Mapped as number |
BOOLEAN | Mapped as boolean |
TIMESTAMP, DATE | Mapped as date/datetime |
STRUCT | Supported in queries; flattened for sync |
ARRAY | Supported in queries; flattened for sync |
MAP | Supported in queries; serialized as JSON for sync |
Delta Lake Features
Databricks tables use Delta Lake format, which provides:
- Time travel — Query historical versions of tables in your model SQL
- Schema evolution — Tables can change schema over time; Zeotap detects column changes
- ACID transactions — Consistent reads even during concurrent writes
-- Time travel: query data as of a specific timestamp
SELECT * FROM customer_data.users TIMESTAMP AS OF '2025-01-01'
-- Time travel: query a specific version
SELECT * FROM customer_data.users VERSION AS OF 42Example Configuration
curl -X POST https://agentic.zeotap.com/api/v1/sources \
-H "Authorization: Bearer $API_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "Production Databricks",
"type": "databricks",
"config": {
"host": "mycompany.cloud.databricks.com",
"http_path": "/sql/1.0/warehouses/abc123def456",
"token": "dapi1234567890abcdef",
"catalog": "main",
"schema": "customer_data"
}
}'Network Configuration
If your Databricks workspace uses private networking (Private Link, VNet injection), ensure that Zeotap can reach the workspace’s public or private endpoint. Options include:
- Public endpoint with IP allowlisting — Add Zeotap IPs to the workspace’s IP access list
- Private connectivity — Contact Zeotap support for private link options
All Zeotap connections to your workspace originate from these static egress IPs:
| Egress IP | Region |
|---|---|
34.76.7.172 | Europe (europe-west1) |
34.22.225.249 | Europe (europe-west1) |
To configure IP access lists in Databricks:
- Go to Admin Console > Workspace Settings
- Enable IP Access Lists
- Add Zeotap’s egress IPs (
34.76.7.172,34.22.225.249) to the allowlist
These addresses are stable — Zeotap does not rotate them. If the list ever changes, this page is updated first.
Troubleshooting
| Issue | Solution |
|---|---|
| ”Invalid access token” | Verify the token is correct and has not expired; generate a new one if needed |
| ”SQL warehouse is not running” | Start the SQL warehouse or enable auto-start in its configuration |
| ”Catalog ‘X’ not found” | Verify the catalog name; use hive_metastore if not using Unity Catalog |
| ”Connection timed out” | Check network access — ensure Zeotap IPs are allowed and the workspace is reachable |
| ”Insufficient privileges” | Verify USE CATALOG, USE SCHEMA, and SELECT permissions are granted |
| ”Schema not found” or a permission error on the first orchestration, identity resolution, or loader run | The feature’s platform schema does not exist and the principal cannot create it — pre-create the schema or grant CREATE SCHEMA on the catalog. See Platform Schemas |
| ”HTTP Path is invalid” | Verify the HTTP path from the warehouse/cluster Connection Details tab |
| ”Journey feed unavailable” / “enabling Change Data Feed needs MODIFY” in a journey run log | Expected, and harmless: the principal cannot enable Delta Change Data Feed on that event table, so the journey keeps its ordinary entry scan and late-arriving events enter on the next reconciliation pass. Grant MODIFY on the table to enable it — see Optional: MODIFY on event tables. After granting it, allow up to 24 hours before the feed is retried: the refusal is recorded with a one-day backoff so a missing grant does not cost a metadata statement on every single wake. Re-saving the journey does not reset that timer — the verdict lives on the journey row and journey edits deliberately do not touch it — and neither does toggling the workspace setting off and on. Waiting is the supported answer |
”Error getting access token from metadata server at http://169.254.169.254/...” during a load | No external location governs the staging bucket, so Databricks fell back to the cluster’s default service account (which has no access). This only affects connections using direct GCS access: either fix the Unity Catalog storage credential and external location, or clear the GCS Service Account field (or set Bulk Load Mode to insert) to have Zeotap relay the rows instead — see Bulk Loading |
Next Steps
- Create a model using your Databricks warehouse
- Use the SQL editor to write Delta Lake queries
- Set up a Reverse ETL sync to activate your data