Skip to Content
WarehousesDatabricks

Databricks

This guide covers how to configure Databricks as a warehouse in Zeotap, including workspace setup, SQL warehouse configuration, authentication, and required permissions.

Prerequisites

  • A Databricks workspace (AWS, Azure, or GCP)
  • A SQL warehouse or all-purpose cluster running in the workspace
  • A personal access token (or service principal) with read access to your data and write access to Zeotap’s dedicated working schemas (see Required Permissions)
  • Network access from Zeotap to your Databricks workspace

Connection Configuration

Required Fields

FieldDescriptionExample
HostThe Databricks workspace URL (without https://)mycompany.cloud.databricks.com
HTTP PathThe HTTP path for the SQL warehouse or cluster/sql/1.0/warehouses/abc123def456
TokenA Databricks personal access tokendapi1234567890abcdef
CatalogThe Unity Catalog name (or hive_metastore for legacy)main
SchemaThe default schema to querydefault

Host

The workspace URL is visible in your browser’s address bar when logged in to Databricks. Use just the hostname without https://:

# AWS mycompany.cloud.databricks.com # Azure adb-1234567890123456.7.azuredatabricks.net # GCP mycompany.gcp.databricks.com

HTTP Path

The HTTP path identifies which compute resource to use for running queries. You can find it in the Databricks UI:

For SQL Warehouses (recommended):

  1. Go to SQL Warehouses in the left sidebar
  2. Click on your warehouse
  3. Go to the Connection Details tab
  4. Copy the HTTP Path value

The format is typically:

/sql/1.0/warehouses/<warehouse-id>

For All-Purpose Clusters:

  1. Go to Compute in the left sidebar
  2. Click on your cluster
  3. Go to Advanced Options > JDBC/ODBC
  4. Copy the HTTP Path value

The format is typically:

/sql/protocolv1/o/<org-id>/<cluster-id>

SQL Warehouses are recommended over all-purpose clusters for Zeotap because they start faster, scale automatically, and provide better cost isolation.

Catalog and Schema

If your Databricks workspace uses Unity Catalog, specify the catalog name (e.g., main, production). If you’re using the legacy Hive Metastore, set the catalog to hive_metastore.

-- Unity Catalog: catalog.schema.table SELECT * FROM main.customer_data.users -- Hive Metastore: schema.table (catalog is hive_metastore) SELECT * FROM customer_data.users

Databricks is case-insensitive for schema and table names, so Customer_Data and customer_data reference the same schema.

Authentication

Personal Access Token

Zeotap uses Databricks personal access tokens (PATs) for authentication.

To create a PAT:

  1. In the Databricks workspace, click your username in the top-right corner
  2. Select Settings
  3. Go to Developer > Access Tokens
  4. Click Generate New Token
  5. Set a description (e.g., “Zeotap CDP”) and expiration
  6. Copy the token — it is only shown once
Token format: dapi1234567890abcdef1234567890abcdef

Token expiration: Set a reasonable expiration (e.g., 90 days) and establish a rotation process. When the token expires, update the source configuration in Zeotap and re-test the connection.

Service Principal (Alternative)

For production environments, consider using a Databricks service principal instead of a personal access token:

  1. Create a service principal in Account Console > User Management
  2. Generate an OAuth secret or PAT for the service principal
  3. Grant the service principal access to the required catalog, schema, and tables

Required Permissions

Zeotap needs read access to your source data, plus the ability to create and write the operational schemas it manages inside the connection’s catalog.

Unity Catalog Permissions

-- Grant catalog access GRANT USE CATALOG ON CATALOG main TO `zeotap-service-principal`; -- Grant schema access GRANT USE SCHEMA ON SCHEMA main.customer_data TO `zeotap-service-principal`; -- Grant table read access GRANT SELECT ON SCHEMA main.customer_data TO `zeotap-service-principal`;

Platform Schemas (Write Access)

Zeotap keeps all of its working state in dedicated schemas inside the connected catalog — your own tables are never modified. Read access to your data is sufficient everywhere else; write access is only needed on these schemas:

SchemaPurposeNeeded for
cdp_plannerTemporary computation tablesAll functionality
cdp_auditSync run audit logsAll functionality
cdp_journeyOrchestration execution state and send logsOrchestrations
cdp_identityIdentity resolution output tables (default schema; configurable per identity graph)Identity resolution
cdp_rawLanding tables for data pulled in by loaders (default schema; configurable per loader)Loaders
audit_logsDelivery and observability event logsEvent delivery / observability logging
cdp_prepPrepared tables built by Data Prep, plus their transient rebuild siblingsOnly if Data Prep is enabled for the workspace
cdp_metadataThe audiences table written by the metadata exportOnly if the metadata export is enabled for the workspace

The simplest setup grants the principal permission to create schemas in the catalog — Zeotap creates each schema on first use, and as its creator it owns it, so no further grants are needed:

GRANT CREATE SCHEMA ON CATALOG main TO `zeotap-service-principal`;

Alternatively, pre-create the schemas and grant the connecting principal full access to them:

CREATE SCHEMA IF NOT EXISTS main.cdp_planner; CREATE SCHEMA IF NOT EXISTS main.cdp_audit; CREATE SCHEMA IF NOT EXISTS main.cdp_journey; CREATE SCHEMA IF NOT EXISTS main.cdp_identity; CREATE SCHEMA IF NOT EXISTS main.cdp_raw; CREATE SCHEMA IF NOT EXISTS main.audit_logs; GRANT ALL PRIVILEGES ON SCHEMA main.cdp_planner TO `zeotap-service-principal`; GRANT USE SCHEMA ON SCHEMA main.cdp_planner TO `zeotap-service-principal`; GRANT ALL PRIVILEGES ON SCHEMA main.cdp_audit TO `zeotap-service-principal`; GRANT USE SCHEMA ON SCHEMA main.cdp_audit TO `zeotap-service-principal`; GRANT ALL PRIVILEGES ON SCHEMA main.cdp_journey TO `zeotap-service-principal`; GRANT USE SCHEMA ON SCHEMA main.cdp_journey TO `zeotap-service-principal`; GRANT ALL PRIVILEGES ON SCHEMA main.cdp_identity TO `zeotap-service-principal`; GRANT USE SCHEMA ON SCHEMA main.cdp_identity TO `zeotap-service-principal`; GRANT ALL PRIVILEGES ON SCHEMA main.cdp_raw TO `zeotap-service-principal`; GRANT USE SCHEMA ON SCHEMA main.cdp_raw TO `zeotap-service-principal`; GRANT ALL PRIVILEGES ON SCHEMA main.audit_logs TO `zeotap-service-principal`; GRANT USE SCHEMA ON SCHEMA main.audit_logs TO `zeotap-service-principal`;

cdp_prep — only if Data Prep is enabled for the workspace

Data Prep is a per-workspace entitlement, off by default. A workspace that has it builds prepared tables into one more schema, cdp_prep, and its connection test then carries a seventh write probe (write_prep). Until Data Prep is enabled nothing creates or reads this schema, so there is nothing to grant.

CREATE SCHEMA IF NOT EXISTS main.cdp_prep; GRANT ALL PRIVILEGES ON SCHEMA main.cdp_prep TO `zeotap-service-principal`; GRANT USE SCHEMA ON SCHEMA main.cdp_prep TO `zeotap-service-principal`;

ALL PRIVILEGES is the same grant the six schemas above take, and every part of it is load-bearing for a prepared table — which, unlike the other schemas’ tables, is REBUILT and, when it is incremental, updated in place: CREATE TABLE for the table and its transient <slug>__next / <slug>__delta siblings, MODIFY for the appends, merges and delete_when deletes, SELECT for the delta computation and for the Models that read the result, and the ability to drop the siblings once the swap is done.

Revoking the entitlement does not drop the schema or its tables, so this grant can be withdrawn at your convenience rather than urgently.

cdp_metadata — only if the metadata export is enabled for the workspace

The metadata export is a per-workspace entitlement, off by default, that keeps the workspace’s audience catalogue as the table cdp_metadata.audiences. When it is enabled for a workspace, the connection test of the source chosen to hold the table carries one more write probe (write_metadata); the workspace’s other sources are tested exactly as before. Until then nothing creates or reads this schema, so there is nothing to grant.

CREATE SCHEMA IF NOT EXISTS main.cdp_metadata; GRANT ALL PRIVILEGES ON SCHEMA main.cdp_metadata TO `zeotap-service-principal`; GRANT USE SCHEMA ON SCHEMA main.cdp_metadata TO `zeotap-service-principal`;

Each export creates audiences__next, inserts into it, replaces audiences with it and drops it, so it needs the same ALL PRIVILEGES as the schemas above. Give readers USE SCHEMA and SELECT on the schema rather than on the table, since the table is replaced by every export.

Revoking the entitlement does not drop the schema or its table.

The connection test verifies write access to every one of these schemas — mirroring how the runtime creates them — and reports the exact statements to run if a step fails. A principal that can only read source data passes authentication but fails the write steps, which would otherwise surface later as a permission error on its first orchestration, identity, loader, or observability run.

Optional: MODIFY on event tables (faster orchestration entry)

This one grant is optional, is on your own tables rather than on Zeotap’s schemas, and nothing fails without it.

A reactive orchestration narrows each evaluation to the people with a new event since the last one, which it decides from the event’s timestamp. That misses a late arrival: a row written today but carrying last week’s timestamp, which a backfill, an offline mobile app flushing its queue, or a partner file that arrives a day late all produce routinely. Such a person enters the orchestration on its next full reconciliation pass instead of within minutes.

Granting MODIFY on the event tables your orchestrations read lets Zeotap turn on Delta Change Data Feed for them, which answers the other question — which rows appeared since the last evaluation — so a late arrival enters on the next wake:

-- Only the event tables your orchestrations read; nothing else needs MODIFY. GRANT MODIFY ON TABLE main.customer_data.events TO `zeotap-service-principal`;

What the grant is actually used for, in full: Zeotap runs ALTER TABLE … SET TBLPROPERTIES (delta.enableChangeDataFeed = true) once per table, and never writes rows to it. Enabling Change Data Feed is a permanent change to your table: from then on Delta also records change data for every write, which occupies storage you own. Both are worth knowing before granting it.

Without the grant, orchestrations behave exactly as they did before — the feed is recorded as unavailable with Databricks’ own message, and the reconciliation pass still admits late arrivals. The feature is also off by default per workspace; ask your Zeotap administrator to enable it once the grant is in place.

Hive Metastore Permissions

-- Grant database access GRANT USAGE ON DATABASE customer_data TO `cdp_user`; -- Grant table read access GRANT SELECT ON DATABASE customer_data TO `cdp_user`;

On the legacy Hive Metastore the platform schemas above are databases; grant CREATE at the metastore level (or pre-create the six databases and grant the user full access to each).

SQL Warehouse Access

The user or service principal must also have Can Use permission on the SQL warehouse:

  1. Go to SQL Warehouses
  2. Click on the warehouse
  3. Go to the Permissions tab
  4. Add the user/service principal with Can Use permission

Bulk Loading

When Zeotap loads or forwards data into Databricks, it first stages the records as files in a Google Cloud Storage (GCS) bucket. How those records then reach your tables depends on which cloud your Databricks workspace runs on, and Zeotap picks the right path automatically from the workspace hostname.

WorkspaceDefault pathWhat you configure
Azure (*.azuredatabricks.net)Zeotap reads the staged files and writes the rows over your SQL warehouseNothing
AWS (*.cloud.databricks.com)SameNothing
GCP (*.gcp.databricks.com)Your warehouse reads the staging bucket directly with read_files() — fasterA Unity Catalog storage credential (optional; without it, the Azure/AWS path is used)

Only a Databricks-on-GCP workspace can be given access to a GCS bucket: Unity Catalog issues a GCP service-account storage credential there, while on Azure it governs abfss:// locations and on AWS s3:// ones. On those clouds there is nothing to set up — Zeotap relays the rows itself, and this works with any catalog, including hive_metastore.

Optional: direct GCS access on a GCP workspace

On a Databricks-on-GCP workspace you can let the warehouse read the staging bucket itself, which is faster than relaying the rows. It requires Unity Catalog — the legacy hive_metastore catalog cannot govern GCS access this way — so set the connection’s Catalog to a Unity Catalog catalog.

  1. A Unity Catalog storage credential for GCS. A Databricks admin creates this in Unity Catalog. When created, Databricks generates a managed Google service-account email that looks like db-uc-credential-xxxxx@<region>.iam.gserviceaccount.com. Paste that email into the GCS Service Account field of the Zeotap Databricks connection. This is the service account Databricks impersonates to read GCS — it is not a key you download.

  2. An external location for the staging bucket, bound to the storage credential above. Zeotap auto-creates this external location for each staging bucket at load time if the connecting principal (the token or service principal used in the connection) has the CREATE EXTERNAL LOCATION privilege on the metastore. If it does not, a Databricks admin must pre-create the external location manually under Catalog → External Data → External Locations, pointing it at the staging bucket URL (gs://<bucket>) and using the storage credential above.

  3. Bucket read access for the credential’s service account. The storage credential’s service account needs read access on the staging bucket. Zeotap grants this automatically.

Leaving the GCS Service Account field empty is always safe: loading falls back to the relayed path rather than failing.

Overriding the choice

The Bulk Load Mode connection field pins the path instead of inferring it:

ValueBehaviour
auto (default)Chooses from the workspace hostname and whether a GCS service account is set
external_locationAlways have the warehouse read gs:// directly
insertAlways relay the rows through Zeotap

Set it only to work around a specific problem — for example pinning insert on a GCP workspace whose external location has been removed, or external_location on a workspace reached through a custom domain that Zeotap cannot classify.

Data Types

Databricks TypeZeotap Handling
STRINGMapped as text
INT, BIGINT, DOUBLE, DECIMALMapped as number
BOOLEANMapped as boolean
TIMESTAMP, DATEMapped as date/datetime
STRUCTSupported in queries; flattened for sync
ARRAYSupported in queries; flattened for sync
MAPSupported in queries; serialized as JSON for sync

Delta Lake Features

Databricks tables use Delta Lake format, which provides:

  • Time travel — Query historical versions of tables in your model SQL
  • Schema evolution — Tables can change schema over time; Zeotap detects column changes
  • ACID transactions — Consistent reads even during concurrent writes
-- Time travel: query data as of a specific timestamp SELECT * FROM customer_data.users TIMESTAMP AS OF '2025-01-01' -- Time travel: query a specific version SELECT * FROM customer_data.users VERSION AS OF 42

Example Configuration

curl -X POST https://agentic.zeotap.com/api/v1/sources \ -H "Authorization: Bearer $API_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "name": "Production Databricks", "type": "databricks", "config": { "host": "mycompany.cloud.databricks.com", "http_path": "/sql/1.0/warehouses/abc123def456", "token": "dapi1234567890abcdef", "catalog": "main", "schema": "customer_data" } }'

Network Configuration

If your Databricks workspace uses private networking (Private Link, VNet injection), ensure that Zeotap can reach the workspace’s public or private endpoint. Options include:

  • Public endpoint with IP allowlisting — Add Zeotap IPs to the workspace’s IP access list
  • Private connectivity — Contact Zeotap support for private link options

All Zeotap connections to your workspace originate from these static egress IPs:

Egress IPRegion
34.76.7.172Europe (europe-west1)
34.22.225.249Europe (europe-west1)

To configure IP access lists in Databricks:

  1. Go to Admin Console > Workspace Settings
  2. Enable IP Access Lists
  3. Add Zeotap’s egress IPs (34.76.7.172, 34.22.225.249) to the allowlist

These addresses are stable — Zeotap does not rotate them. If the list ever changes, this page is updated first.

Troubleshooting

IssueSolution
”Invalid access token”Verify the token is correct and has not expired; generate a new one if needed
”SQL warehouse is not running”Start the SQL warehouse or enable auto-start in its configuration
”Catalog ‘X’ not found”Verify the catalog name; use hive_metastore if not using Unity Catalog
”Connection timed out”Check network access — ensure Zeotap IPs are allowed and the workspace is reachable
”Insufficient privileges”Verify USE CATALOG, USE SCHEMA, and SELECT permissions are granted
”Schema not found” or a permission error on the first orchestration, identity resolution, or loader runThe feature’s platform schema does not exist and the principal cannot create it — pre-create the schema or grant CREATE SCHEMA on the catalog. See Platform Schemas
”HTTP Path is invalid”Verify the HTTP path from the warehouse/cluster Connection Details tab
”Journey feed unavailable” / “enabling Change Data Feed needs MODIFY” in a journey run logExpected, and harmless: the principal cannot enable Delta Change Data Feed on that event table, so the journey keeps its ordinary entry scan and late-arriving events enter on the next reconciliation pass. Grant MODIFY on the table to enable it — see Optional: MODIFY on event tables. After granting it, allow up to 24 hours before the feed is retried: the refusal is recorded with a one-day backoff so a missing grant does not cost a metadata statement on every single wake. Re-saving the journey does not reset that timer — the verdict lives on the journey row and journey edits deliberately do not touch it — and neither does toggling the workspace setting off and on. Waiting is the supported answer
”Error getting access token from metadata server at http://169.254.169.254/...” during a loadNo external location governs the staging bucket, so Databricks fell back to the cluster’s default service account (which has no access). This only affects connections using direct GCS access: either fix the Unity Catalog storage credential and external location, or clear the GCS Service Account field (or set Bulk Load Mode to insert) to have Zeotap relay the rows instead — see Bulk Loading

Next Steps

Last updated on