Data Prep with Zeotap Agent
Data Prep was built to be usable by Zeotap Agent as well as by hand. The agent can profile a raw table, propose a plan, write the recipe, validate it, preview it, build it and create the model over it — producing exactly the same object you would have made in the editor, which you can then open and change.
What you can ask for
- “What’s actually in
cdp_raw.contacts?” — it profiles the table and describes what it found: the types, the keys, the JSON, how much of it is castable, which column looks like a watermark. - “Clean this up so I can model it.” — it proposes a plan in plain language (which tables, which steps, how it would be built and on what schedule, and what a build would cost), waits for your go-ahead, and then builds it.
- “We just connected Shopify.” — it checks for a connector blueprint first, because one is faster and better tested than a recipe written from a profile.
- “Why is this table stale?” / “Did that build work?” — it reads the build back and reports the mode, the rows, any warnings and any failed data test.
- “Is my incremental table correct?” — it can start a reconcile and report the drift with its numbers.
How it behaves
It asks before it spends. The sequence is always: validate, read the cost estimate, tell you what a build would scan, and only then build. Where the estimate cannot be known — most warehouses cannot price a query before running it, and in practice only BigQuery can — it says that it could not price it rather than implying the build is small.
A very large build needs a workspace admin’s approval. If the agent asks for a build whose estimated scan is above the workspace’s threshold, it is held rather than run: the agent tells you, gives you the approval to look at, and the build runs automatically once an admin approves it. This is the same approval flow the agent’s other expensive actions use.
The approved build runs as the person who asked for it, not as the person who approved it. The requester’s own permissions are re-checked at that moment, so approving a request can never grant somebody access they do not have — an admin approving a build is agreeing to the cost, not lending their role.
Personal data is masked before the agent receives it. A column a profile measured as mostly email addresses or phone numbers, or that a model over the same table marks sensitive, comes back redacted — in profiles and in previews alike. Where Zeotap cannot prove which input a column came from, which is what happens when a recipe contains a hand-written SQL step, it masks rather than assumes the column is clean.
On top of that, a preview measures its own rows: a text column that lineage says nothing about, in which the email or phone pattern matches at least half of the values on screen, is redacted before the agent sees it — with no profile needed. Lineage still wins where it speaks, non-text columns are never caught this way, and the measurement is never stored.
The agent still profiles a table before previewing a recipe over it, because that is what finds every other kind of personal data — names, postal addresses, dates of birth have no pattern reliable enough to fire on. See Permissions & Grants.
It prefers typed steps. They give exact column lineage, a real incremental-safety check and portability across warehouses. When it does reach for a SQL step it will say so, and say what that costs.
It relays refusals verbatim. Zeotap’s refusals are written to state the fix — “this step
keeps the latest row per contact_id, which needs the merge strategy…” — so the agent quotes them
rather than paraphrasing.
When Data Prep is not enabled
The agent says so, says that a platform administrator enables it, and falls back to ordinary modelling. It does not call a Data Prep tool “to check”.
The onboarding assistant
A workspace being set up for the first time gets an agent that leads with connect → profile → prepare → model → relate, rather than starting at models. That order exists because a workspace pointing at a warehouse full of raw landings needs prepared tables before it needs models — and starting at models is how cleaning ends up pasted into one.
What you should still do yourself
- Read the plan before you approve it. The agent names the tables, the steps, the materialization and the schedule; that is the moment to disagree.
- Check a cast’s success rate. The agent reports it; a rate just under the threshold means real rows will become NULL.
- Decide on data tests. The agent offers the tests a table’s own keys imply and adds only what
you accept — an
errortest can stop a table updating, and that is your call.