Apify Dataset data extractor

Import a new customer’s Apify Dataset data — dataset, dataset collection, item collection, and item collection website content crawler — into your product as normalised, relational tables. Vern handles Apify Dataset’s auth, pagination and rate limits, so onboarding a customer off Apify Dataset doesn’t mean building and maintaining that plumbing yourself.

Objects extracted
4

Every queryable object type

Incremental

Full refresh on each run

Authentication
API key or access token

Stored encrypted, never in the transcript

Export coverage

What Vern pulls from Apify Dataset

Know exactly what comes across before you migrate. Every object below is normalised into typed tables with referential keys preserved, so the relationships between a customer's records survive the import.

  • Dataset
  • Dataset collection
  • Item collection
  • Item collection website content crawler

Managed by Vern

The parts you'd otherwise build yourself

Authentication

Your customer pastes their key once when they start the migration. Vern stores it encrypted, signs each request and handles retries — the raw credential never reaches the agent transcript.

Pagination

Vern detects and handles Apify Dataset's paging model — cursor, page number or offset — internally, and returns one consistent, de-duplicated result set. You import a customer's complete data without ever managing a page token.

Rate limits

API sources rarely publish their real rate ceiling until you hit a 429 mid-import, and many silently cap page size. Vern queues and throttles requests to stay under the limit, so a partial pull never lands a customer in your product with missing data.

Output shape

Normalised JSON per object with referential keys preserved and timestamps in ISO 8601. Delivered as JSON or CSV, pushed to your API or webhook, or loaded straight into your application database.

Schema drift

If Apify Dataset changes what it returns between migrations, the agent notices the difference against what it recorded last time and raises it, rather than importing a customer with a column quietly missing.

What it learns

Everything Vern works out about Apify Dataset's quirks — field formats, enum values, ID and date patterns — is written down against the source and applied automatically to every customer you migrate off it afterwards.

What your customer provides

  • Dataset ID
  • API token

Collected once, stored encrypted in a vault, and attached to every request. Credentials are never written into the agent transcript.

API quickstart

Run a Apify Dataset migration from your own product

Create the migration, let the agent extract and map, then pull the result out — without your customer ever seeing a Vern dashboard.

app.vern.so/api/v1
API=https://app.vern.so/api/v1
KEY="x-api-key: $VERN_API_KEY"
JSON="Content-Type: application/json"

# 1 — create a migration against Apify Dataset
ID=$(curl -sX POST $API/migrations -H "$KEY" -H "$JSON" \
  -d '{"name":"Acme onboarding","source":"Apify Dataset"}' \
  | jq -r .migration.id)

# 2 — hand over the customer's credentials for this run
curl -X POST $API/migrations/$ID/source-connection -H "$KEY" -H "$JSON" \
  -d '{"credentials": { ... }}'

# 3 — the agent extracts, maps and previews, stopping before it writes
curl -X POST $API/migrations/$ID/runs -H "$KEY" -H "$JSON" \
  -d '{"kind":"generate"}'

curl $API/migrations/$ID/preview -H "$KEY"

# 4 — approve, then pull the normalised rows back out
curl -X POST $API/migrations/$ID/runs -H "$KEY" -H "$JSON" \
  -d '{"kind":"execute"}'

curl -X POST $API/migrations/$ID/exports -H "$KEY" -H "$JSON" \
  -d '{"slugs":["customers","invoices"]}'

The Migration API runs the whole lifecycle over HTTP with an organisation-scoped key. The agent conversation can be streamed into your own UI, and the preview lets you check exactly what an import will produce before a single row is written.

Read the API reference

Onboarding a customer off Apify Dataset?

Bring a real Apify Dataset export and we'll show you what lands in your product.