Spark Lakehouse
This guide covers how to connect your own Spark + Iceberg lakehouse to Zeotap. Unlike the managed warehouses, a Spark lakehouse is infrastructure you run: Zeotap submits Spark SQL to a REST gateway you expose — Kyuubi or Livy — and reads results back from the catalog and object storage you already own.
Prerequisites
- A Spark engine reachable over HTTP through a Kyuubi REST gateway (default) or a Livy REST gateway
- An Iceberg REST catalog the engine is configured against
- A catalog name Zeotap can write to, and object storage backing that catalog (S3, MinIO, or GCS)
- Network access from Zeotap to the gateway endpoint
- For a Livy gateway: object-storage credentials, which are required rather than optional — see Staging
Connection Configuration
Required Fields
| Field | Description | Example |
|---|---|---|
| Kyuubi / Livy REST endpoint | Base URL of the gateway. The label follows the transport you pick | https://kyuubi.your-domain:10099 |
| Catalog | The top of the catalog.schema.table namespace | prod |
Everything else has a working default.
Optional Fields
| Field | Description | Default |
|---|---|---|
| Transport | Kyuubi or Livy | Kyuubi |
| Default schema | Schema used for browsing and models | — |
| Catalog type | Catalog protocol the engine is configured with. Informational | Iceberg REST |
| Table format | Format for the tables Zeotap creates | Iceberg |
| Resource pool | The engine scheduler pool or queue Zeotap’s work is submitted to | — |
| Auth method | None or Basic | None |
Zeotap keeps its own tables in a dedicated namespace named zeotap inside your catalog, so its working tables never mix with yours.
Transport
Kyuubi is the default and the recommended gateway. It pages large result sets natively, so a Kyuubi source needs nothing beyond the endpoint and catalog to run a full-sized sync.
Livy truncates large interactive result sets. A Livy source therefore always unloads query results to object storage as Parquet and reads them back, which is why staging credentials are required for Livy rather than optional. Choose Livy only if Livy is already what fronts your Spark.
The endpoint you enter should be the gateway’s base URL, including the scheme and port:
# Kyuubi
https://kyuubi.your-domain:10099
# Livy
https://livy.your-domain:8998Table Format
Iceberg is the only supported table format. Leaving the field blank selects Iceberg.
Two other formats are rejected at save time, each with a specific reason:
| Value | Result |
|---|---|
iceberg | Supported. Every table Zeotap creates is an Iceberg table |
delta | Rejected — every write path currently creates Iceberg tables. Delta support is a planned follow-up |
hudi | Rejected — Hudi’s Spark SQL row-level DML is too limited for orchestrations and identity resolution |
Reading from Delta or Hudi tables your engine already exposes is a separate question from the format Zeotap writes; this setting governs only what Zeotap creates.
Authentication
None
The default. Use it when the gateway is only reachable from inside a trusted network and performs no authentication of its own.
Basic
Sends a username and password with each request. Use HTTPS for the endpoint so credentials are encrypted in transit.
| Field | Description |
|---|---|
| Username | Gateway username |
| Password | Gateway password. Stored encrypted and never returned by the API |
Staging
Staging is how Zeotap reads large result sets back: the engine writes Parquet to object storage and Zeotap reads the part files directly.
- Kyuubi — optional. Kyuubi pages results natively.
- Livy — required. Without staging credentials a Livy source silently caps out at Livy’s interactive result limit.
Which object store your lakehouse sits on selects the credential fields.
S3 or MinIO
| Field | Description |
|---|---|
| Access key | S3 access key ID |
| Secret key | S3 secret access key |
| S3 endpoint | Endpoint override. Leave blank for AWS S3; set it for MinIO or another non-AWS S3 implementation |
| Path-style addressing | Enable for MinIO and most non-AWS implementations, which do not support virtual-host-style bucket addressing |
GCS
| Field | Description |
|---|---|
| Service account JSON | Service-account key with read access to the objects the engine writes |
Spark-Specific Notes
- Spark SQL is the dialect. Models and audience filters are written in Spark SQL, which for most purposes is the same dialect Databricks accepts.
- Custom-code transforms are not available. JavaScript and Python field transforms are unsupported on a Spark lakehouse, as they are on ClickHouse and Redshift. Express the transformation in SQL instead.
- Spark Lakehouse also works as a destination. You can write model or audience data back into your lakehouse the same way you would to Snowflake, BigQuery, Databricks, or ClickHouse.
- You own the capacity. Zeotap submits work to an engine you run, so its footprint is bounded by whatever the resource pool you name allows.
Example Configuration
A Kyuubi-fronted lakehouse on MinIO:
Transport: Kyuubi
Endpoint: https://kyuubi.internal.mycompany.com:10099
Catalog: prod
Default schema: cdp
Catalog type: Iceberg REST
Table format: Iceberg
Auth method: Basic
Username: zeotap
Object store: S3
S3 endpoint: https://minio.internal.mycompany.com:9000
Path-style: EnabledTroubleshooting
| Issue | Solution |
|---|---|
| ”endpoint is required” | Enter the gateway base URL including scheme and port |
| ”catalog is required” | Enter the catalog name your engine is configured with — the top of catalog.schema.table |
table_format 'delta' is not supported yet | Use Iceberg. Delta writes are a planned follow-up |
table_format 'hudi' is not supported | Use Iceberg. Hudi’s row-level DML cannot support orchestrations or identity resolution |
| Syncs return far fewer rows than expected on Livy | Staging is not configured. Livy truncates interactive results — add object-storage credentials |
| ”Connection timed out” | Confirm the gateway is reachable from Zeotap and the port is open |
| Object-storage reads fail against MinIO | Set the S3 endpoint override and enable path-style addressing |
| Custom-code transform rejected | JavaScript and Python transforms are not supported here — rewrite the transform as SQL |
Next Steps
- Create a model using your Spark lakehouse
- Use the SQL editor to write Spark SQL queries
- Set up a Reverse ETL sync to activate your data