Skip to Content
LoadersAmazon S3

Amazon S3

The Amazon S3 loader imports data files from an S3 bucket into your warehouse. Zeotap lists the objects under a bucket prefix, parses each one, and loads the rows into a table you can model, segment, and activate.

Prerequisites

  • An Amazon S3 bucket containing your data files.
  • An AWS access key ID and secret access key for a principal that can list and read those objects — the loader calls ListObjectsV2 on the bucket and GetObject on each file, so it needs s3:ListBucket on the bucket and s3:GetObject on the keys under your prefix.
  • The AWS region the bucket lives in.
  • A connected warehouse to load the data into.

Authentication

Zeotap authenticates with an access key pair.

  1. In the AWS console, create (or select) an IAM user or role for Zeotap.
  2. Attach a policy granting s3:ListBucket on the bucket and s3:GetObject on the objects under the prefix you want to read.
  3. Create an access key for it.
  4. In Zeotap, add a new Amazon S3 loader and enter the Access Key ID and Secret Access Key.

The secret access key is stored encrypted and is only used to read objects from the configured bucket.

Configuration

FieldTypeRequiredDescription
Access Key IDTextYesThe AWS access key ID (e.g. AKIAIOSFODNN7EXAMPLE).
Secret Access KeySecretYesThe matching secret access key.
BucketTextYesThe S3 bucket name (without the s3:// prefix).
PrefixTextNoObject key prefix to filter which files are read (e.g. data/exports/). Leave blank to read the whole bucket.
File FormatSelectYesThe format of the files: CSV or JSON (newline-delimited). Defaults to CSV.
AWS RegionSelectYesThe region the bucket is in. Defaults to US East (N. Virginia) (us-east-1).
Incremental CursorSelectNoFor incremental syncs, how new files are detected: File name (default) or Last modified time. Ignored for full-refresh syncs. See Sync Behavior.

All files under the prefix that match the selected format are loaded into a single table.

Supported Formats

FormatNotes
CSVThe first row is treated as the header. Values are loaded as text. Matches .csv and .csv.gz.
JSONNewline-delimited JSON (one JSON object per line, also called JSONL/NDJSON). Matches .json, .jsonl, .ndjson and .json.gz. A single large JSON array is not supported — use newline-delimited objects.

Column types are loaded as text and the schema is inferred from the first file read. Cast or transform values in a downstream model.

File Layout

  • Place files for one logical dataset under a shared prefix (e.g. data/exports/orders/) and point the loader at that prefix.
  • Every matching file under the prefix is combined into one table. The table’s columns are determined by the first file read — all files under the prefix are assumed to share the same schema. Columns present only in later files are dropped, and columns missing from a later file are loaded as NULL.
  • Each loaded row includes an _s3_file column recording the source object key, so you can trace rows back to their file.

Sync Behavior

The loader supports both full refresh and incremental syncs.

  • Full refresh — each run truncates the target table and reads every matching file under the prefix. Use it when files are rewritten in place or the dataset is small.
  • Incremental — each run loads only files that are new since the previous run and appends them. Pick how “new” is decided with the Incremental Cursor setting:
CursorHow it detects new filesBest forTrade-offs
File name (default)Objects whose key sorts lexicographically after the last file loaded. Already-processed keys are skipped server-side (via the ListObjectsV2 StartAfter parameter), so listing stays fast even as history grows.Date- or sequence-partitioned exports where keys always increase, e.g. exports/2026/07/16/part-000.json.Requires fixed-width, always-increasing keys. A counter that overflows its padding (part-999 → part-1000) sorts before the cursor and is silently skipped. Files written with an earlier-sorting key after a sync are missed.
Last modified timeObjects whose last-modified timestamp is after the newest one loaded.Any layout — catches renamed, backfilled, or edited files regardless of key, as long as their last-modified time is newer than the previous sync.No naming assumptions, but S3 offers no server-side time filter, so the loader lists the whole prefix each run and filters client-side (cost grows with total object count). A backfilled file written with an older timestamp is at or before the watermark and is skipped.

Each incremental run appends; it does not truncate. Switching a loader between the two cursor strategies (or from full refresh to incremental) triggers one reprocessing pass — the stored watermark from the other strategy is not reused — which may duplicate rows already loaded.

To keep the warehouse table current, schedule the loader and write new data to the prefix.

Troubleshooting

IssueResolution
”bucket is required” / “region is required”Both are mandatory. Enter the bucket name without s3://, and pick the region the bucket was created in.
Access deniedThe key pair cannot list or read the objects. Confirm the policy grants s3:ListBucket on the bucket and s3:GetObject on the keys under your prefix.
No files found with prefixNo objects under the prefix match the selected format. Check the prefix value and confirm the files carry a matching extension (.csv, .csv.gz, .json, .jsonl, .ndjson, .json.gz).
Wrong regionA bucket reached with the wrong region fails to list. Set AWS Region to the region the bucket was created in.
Wrong or missing columns (CSV)The first row must be a header row. Confirm the file is delimited correctly and the header matches the data.
JSON rows skippedLines that are not valid JSON objects are skipped. Confirm the file is newline-delimited JSON (one object per line), not a single JSON array.
Type mismatchesCSV and JSON values are loaded as text. Cast or transform them in a downstream model.

Next Steps

Last updated on