Skip to Content
WarehousesSpark Lakehouse

Spark Lakehouse

This guide covers how to connect your own Spark + Iceberg lakehouse to Zeotap. Unlike the managed warehouses, a Spark lakehouse is infrastructure you run: Zeotap submits Spark SQL to a REST gateway you expose — Kyuubi or Livy — and reads results back from the catalog and object storage you already own.

Prerequisites

  • A Spark engine reachable over HTTP through a Kyuubi REST gateway (default) or a Livy REST gateway
  • An Iceberg REST catalog the engine is configured against
  • A catalog name Zeotap can write to, and object storage backing that catalog (S3, MinIO, or GCS)
  • Network access from Zeotap to the gateway endpoint
  • For a Livy gateway: object-storage credentials, which are required rather than optional — see Staging

Connection Configuration

Required Fields

FieldDescriptionExample
Kyuubi / Livy REST endpointBase URL of the gateway. The label follows the transport you pickhttps://kyuubi.your-domain:10099
CatalogThe top of the catalog.schema.table namespaceprod

Everything else has a working default.

Optional Fields

FieldDescriptionDefault
TransportKyuubi or LivyKyuubi
Default schemaSchema used for browsing and models—
Catalog typeCatalog protocol the engine is configured with. InformationalIceberg REST
Table formatFormat for the tables Zeotap createsIceberg
Resource poolThe engine scheduler pool or queue Zeotap’s work is submitted to—
Auth methodNone or BasicNone

Zeotap keeps its own tables in a dedicated namespace named zeotap inside your catalog, so its working tables never mix with yours.

Transport

Kyuubi is the default and the recommended gateway. It pages large result sets natively, so a Kyuubi source needs nothing beyond the endpoint and catalog to run a full-sized sync.

Livy truncates large interactive result sets. A Livy source therefore always unloads query results to object storage as Parquet and reads them back, which is why staging credentials are required for Livy rather than optional. Choose Livy only if Livy is already what fronts your Spark.

The endpoint you enter should be the gateway’s base URL, including the scheme and port:

# Kyuubi https://kyuubi.your-domain:10099 # Livy https://livy.your-domain:8998

Table Format

Iceberg is the only supported table format. Leaving the field blank selects Iceberg.

Two other formats are rejected at save time, each with a specific reason:

ValueResult
icebergSupported. Every table Zeotap creates is an Iceberg table
deltaRejected — every write path currently creates Iceberg tables. Delta support is a planned follow-up
hudiRejected — Hudi’s Spark SQL row-level DML is too limited for orchestrations and identity resolution

Reading from Delta or Hudi tables your engine already exposes is a separate question from the format Zeotap writes; this setting governs only what Zeotap creates.

Authentication

None

The default. Use it when the gateway is only reachable from inside a trusted network and performs no authentication of its own.

Basic

Sends a username and password with each request. Use HTTPS for the endpoint so credentials are encrypted in transit.

FieldDescription
UsernameGateway username
PasswordGateway password. Stored encrypted and never returned by the API

Staging

Staging is how Zeotap reads large result sets back: the engine writes Parquet to object storage and Zeotap reads the part files directly.

  • Kyuubi — optional. Kyuubi pages results natively.
  • Livy — required. Without staging credentials a Livy source silently caps out at Livy’s interactive result limit.

Which object store your lakehouse sits on selects the credential fields.

S3 or MinIO

FieldDescription
Access keyS3 access key ID
Secret keyS3 secret access key
S3 endpointEndpoint override. Leave blank for AWS S3; set it for MinIO or another non-AWS S3 implementation
Path-style addressingEnable for MinIO and most non-AWS implementations, which do not support virtual-host-style bucket addressing

GCS

FieldDescription
Service account JSONService-account key with read access to the objects the engine writes

Spark-Specific Notes

  • Spark SQL is the dialect. Models and audience filters are written in Spark SQL, which for most purposes is the same dialect Databricks accepts.
  • Custom-code transforms are not available. JavaScript and Python field transforms are unsupported on a Spark lakehouse, as they are on ClickHouse and Redshift. Express the transformation in SQL instead.
  • Spark Lakehouse also works as a destination. You can write model or audience data back into your lakehouse the same way you would to Snowflake, BigQuery, Databricks, or ClickHouse.
  • You own the capacity. Zeotap submits work to an engine you run, so its footprint is bounded by whatever the resource pool you name allows.

Example Configuration

A Kyuubi-fronted lakehouse on MinIO:

Transport: Kyuubi Endpoint: https://kyuubi.internal.mycompany.com:10099 Catalog: prod Default schema: cdp Catalog type: Iceberg REST Table format: Iceberg Auth method: Basic Username: zeotap Object store: S3 S3 endpoint: https://minio.internal.mycompany.com:9000 Path-style: Enabled

Troubleshooting

IssueSolution
”endpoint is required”Enter the gateway base URL including scheme and port
”catalog is required”Enter the catalog name your engine is configured with — the top of catalog.schema.table
table_format 'delta' is not supported yetUse Iceberg. Delta writes are a planned follow-up
table_format 'hudi' is not supportedUse Iceberg. Hudi’s row-level DML cannot support orchestrations or identity resolution
Syncs return far fewer rows than expected on LivyStaging is not configured. Livy truncates interactive results — add object-storage credentials
”Connection timed out”Confirm the gateway is reachable from Zeotap and the port is open
Object-storage reads fail against MinIOSet the S3 endpoint override and enable path-style addressing
Custom-code transform rejectedJavaScript and Python transforms are not supported here — rewrite the transform as SQL

Next Steps

Last updated on