Amazon S3
The Amazon S3 loader imports data files from an S3 bucket into your warehouse. Zeotap lists the objects under a bucket prefix, parses each one, and loads the rows into a table you can model, segment, and activate.
Prerequisites
- An Amazon S3 bucket containing your data files.
- An AWS access key ID and secret access key for a principal that can list and read those objects — the loader calls
ListObjectsV2on the bucket andGetObjecton each file, so it needss3:ListBucketon the bucket ands3:GetObjecton the keys under your prefix. - The AWS region the bucket lives in.
- A connected warehouse to load the data into.
Authentication
Zeotap authenticates with an access key pair.
- In the AWS console, create (or select) an IAM user or role for Zeotap.
- Attach a policy granting
s3:ListBucketon the bucket ands3:GetObjecton the objects under the prefix you want to read. - Create an access key for it.
- In Zeotap, add a new Amazon S3 loader and enter the Access Key ID and Secret Access Key.
The secret access key is stored encrypted and is only used to read objects from the configured bucket.
Configuration
| Field | Type | Required | Description |
|---|---|---|---|
| Access Key ID | Text | Yes | The AWS access key ID (e.g. AKIAIOSFODNN7EXAMPLE). |
| Secret Access Key | Secret | Yes | The matching secret access key. |
| Bucket | Text | Yes | The S3 bucket name (without the s3:// prefix). |
| Prefix | Text | No | Object key prefix to filter which files are read (e.g. data/exports/). Leave blank to read the whole bucket. |
| File Format | Select | Yes | The format of the files: CSV or JSON (newline-delimited). Defaults to CSV. |
| AWS Region | Select | Yes | The region the bucket is in. Defaults to US East (N. Virginia) (us-east-1). |
| Incremental Cursor | Select | No | For incremental syncs, how new files are detected: File name (default) or Last modified time. Ignored for full-refresh syncs. See Sync Behavior. |
All files under the prefix that match the selected format are loaded into a single table.
Supported Formats
| Format | Notes |
|---|---|
| CSV | The first row is treated as the header. Values are loaded as text. Matches .csv and .csv.gz. |
| JSON | Newline-delimited JSON (one JSON object per line, also called JSONL/NDJSON). Matches .json, .jsonl, .ndjson and .json.gz. A single large JSON array is not supported — use newline-delimited objects. |
Column types are loaded as text and the schema is inferred from the first file read. Cast or transform values in a downstream model.
File Layout
- Place files for one logical dataset under a shared prefix (e.g.
data/exports/orders/) and point the loader at that prefix. - Every matching file under the prefix is combined into one table. The table’s columns are determined by the first file read — all files under the prefix are assumed to share the same schema. Columns present only in later files are dropped, and columns missing from a later file are loaded as
NULL. - Each loaded row includes an
_s3_filecolumn recording the source object key, so you can trace rows back to their file.
Sync Behavior
The loader supports both full refresh and incremental syncs.
- Full refresh — each run truncates the target table and reads every matching file under the prefix. Use it when files are rewritten in place or the dataset is small.
- Incremental — each run loads only files that are new since the previous run and appends them. Pick how “new” is decided with the Incremental Cursor setting:
| Cursor | How it detects new files | Best for | Trade-offs |
|---|---|---|---|
| File name (default) | Objects whose key sorts lexicographically after the last file loaded. Already-processed keys are skipped server-side (via the ListObjectsV2 StartAfter parameter), so listing stays fast even as history grows. | Date- or sequence-partitioned exports where keys always increase, e.g. exports/2026/07/16/part-000.json. | Requires fixed-width, always-increasing keys. A counter that overflows its padding (part-999 → part-1000) sorts before the cursor and is silently skipped. Files written with an earlier-sorting key after a sync are missed. |
| Last modified time | Objects whose last-modified timestamp is after the newest one loaded. | Any layout — catches renamed, backfilled, or edited files regardless of key, as long as their last-modified time is newer than the previous sync. | No naming assumptions, but S3 offers no server-side time filter, so the loader lists the whole prefix each run and filters client-side (cost grows with total object count). A backfilled file written with an older timestamp is at or before the watermark and is skipped. |
Each incremental run appends; it does not truncate. Switching a loader between the two cursor strategies (or from full refresh to incremental) triggers one reprocessing pass — the stored watermark from the other strategy is not reused — which may duplicate rows already loaded.
To keep the warehouse table current, schedule the loader and write new data to the prefix.
Troubleshooting
| Issue | Resolution |
|---|---|
| ”bucket is required” / “region is required” | Both are mandatory. Enter the bucket name without s3://, and pick the region the bucket was created in. |
| Access denied | The key pair cannot list or read the objects. Confirm the policy grants s3:ListBucket on the bucket and s3:GetObject on the keys under your prefix. |
| No files found with prefix | No objects under the prefix match the selected format. Check the prefix value and confirm the files carry a matching extension (.csv, .csv.gz, .json, .jsonl, .ndjson, .json.gz). |
| Wrong region | A bucket reached with the wrong region fails to list. Set AWS Region to the region the bucket was created in. |
| Wrong or missing columns (CSV) | The first row must be a header row. Confirm the file is delimited correctly and the header matches the data. |
| JSON rows skipped | Lines that are not valid JSON objects are skipped. Confirm the file is newline-delimited JSON (one object per line), not a single JSON array. |
| Type mismatches | CSV and JSON values are loaded as text. Cast or transform them in a downstream model. |
Next Steps
- Create a Model over the loaded table.
- Build an Audience from the modeled data.