Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Last updated: 2026-07-09.

Getting Started

Rivet exports tables from PostgreSQL, MySQL, and SQL Server (and collections from MongoDB) to Parquet (or CSV) files — locally, to S3, GCS, or Azure Blob Storage. Point it at a database, scaffold a config from your real tables, then run. (MongoDB has its own reference: reference/mongodb.md.)

brew install panchenkoai/rivet/rivet
export DATABASE_URL='postgresql://user:pass@localhost:5432/mydb'
# `orders` is a placeholder — use one of YOUR tables, or omit --table to scan the whole schema
rivet init --source-env DATABASE_URL --table orders -o rivet.yaml
rivet run -c rivet.yaml --validate

That’s the whole flow. The four steps below explain each command, expected output, and where to go from each. Read time: ~3 minutes.

Already running it locally? Jump to §3 Preflight & run. If you’re evaluating it for production, finish this page first, then continue with docs/pilot/.


1 · Install

# macOS / Linux — Homebrew (recommended)
brew install panchenkoai/rivet/rivet
rivet --version
# Docker — try without installing anything
docker run --rm ghcr.io/panchenkoai/rivet:latest --version

Pre-built binaries are published for Linux and macOS (x86-64 + arm64). On Windows, install from source with cargo install rivet-cli (a native binary is not currently published). Other install paths — cargo install rivet-cli, build from source, plus the full Docker recipe with database-on-host pointers (host.docker.internal vs --network host) — live in the project README § Installation. Shell completions: rivet completions bash|zsh|fish.

Try it in 60 seconds — no database of your own

Spin up a throwaway PostgreSQL, seed one table, and export it — nothing external to configure:

docker run -d --name rivet-demo -e POSTGRES_PASSWORD=demo -p 5432:5432 postgres:16
sleep 3
docker exec -i rivet-demo psql -U postgres <<'SQL'
CREATE TABLE orders (id serial PRIMARY KEY, name text, price numeric(10,2),
                     updated_at timestamptz DEFAULT now());
INSERT INTO orders (name, price)
  SELECT 'order-'||g, (random()*500)::numeric(10,2) FROM generate_series(1,500) g;
SQL

export DATABASE_URL='postgresql://postgres:demo@localhost:5432/postgres'
rivet init --source-env DATABASE_URL --table orders -o rivet.yaml
rivet run -c rivet.yaml --validate
# → 500 rows of typed Parquet in ./output/orders/.  Clean up: docker rm -f rivet-demo

That is the whole flow against a real (throwaway) database. Then jump to §4 Inspect & iterate, or read on to point Rivet at your own database.

2 · Connect & scaffold a config

Recommended pattern: put the connection URL in an environment variable and reference it from the config so credentials never enter the file or shell history.

export DATABASE_URL='postgresql://user:pass@localhost:5432/mydb'
# MySQL: same flag, just a mysql:// URL
# export DATABASE_URL='mysql://user:pass@localhost:3306/mydb'

rivet init --source-env DATABASE_URL --table orders -o rivet.yaml
# `orders` is a placeholder — use one of YOUR tables, or omit --table to scan the whole schema

rivet init connects once, reads the column list + a rough row estimate from the live database, and writes a YAML file with url_env: DATABASE_URL and a sensible default mode. For a large table with a single-column primary key it picks keyset (chunk_by_key) — seek paging that stays flat-memory and is immune to sparse/gappy keys; keyset is scaffolded sequential (add parallel: N yourself to fan it into row-percentile ranges). A large table with no single-column PK gets a range chunk_column with a row-scaled parallel: (1 / 2 / 4) out of the box — measured ~1.4× faster on a wide 2 M-row table and ~4× on a narrow one. You can also point it at a whole schema (--schema public) or emit a richer JSON discovery artifact instead (--discover -o discovery.json).

Full flag reference: reference/init.md. For a manually-authored YAML instead of rivet init, see reference/config.md.

State file. Rivet creates .rivet_state.db next to the config (cursors, chunk checkpoints, run history). Add it to .gitignore if the folder is version-controlled — see SECURITY.md § Sensitive local artifacts.

3 · Preflight & run

rivet doctor -c rivet.yaml   # verify source + destination auth
rivet check  -c rivet.yaml   # dry-run analysis per export
rivet run    -c rivet.yaml --validate --reconcile

The full basic workflow (init → doctor → check → run → state) recorded as a single terminal cast:

What each step does:

  • rivet doctor — connects to the source and writes a tiny probe object (.rivet_doctor_probe) to every destination prefix — removed afterwards on local destinations, while on S3 / GCS / Azure it stays at the prefix (the destination seam has no delete) and is filtered out of manifest and validate listings; fixes nothing, fails loudly on any auth / network issue.

  • rivet check — runs EXPLAIN against your queries, estimates row counts, detects whether your cursor / chunk columns are indexed, and emits a verdict + concrete suggestion. Verdicts are EFFICIENT · ACCEPTABLE · DEGRADED · UNSAFE; on the SQL engines the last two carry a mode-aware Suggestion: line (MongoDB is full-scan-only, so its verdicts omit the mode suggestion).

  • rivet run --validate --reconcile — extracts. --validate reads each output file back and verifies its row count; --reconcile runs SELECT COUNT(*) on the source query and compares with what was exported.

Example summary card after a successful run:

── orders ──
  run_id:      orders_20260519T120000.123
  status:      success
  tuning:      profile=balanced (default), batch_size=10,000 (batch_size_memory_mb=32MiB → effective FETCH in logs)
  rows:        5,432
  files:       1
  output:      file://./output
  bytes read:    1.2 MB
  bytes written: 847.0 KB
  duration:    1.2s
  peak RSS:    15 MB (sampled during run)
  validated:   pass
  schema:      unchanged
  reconcile:   MATCH (5,432/5,432)

4 · Inspect & iterate

rivet state show   -c rivet.yaml             # cursors (incremental exports)
rivet metrics      -c rivet.yaml --last 10   # per-run history
rivet state files  -c rivet.yaml             # files actually written
rivet journal      -c rivet.yaml --export orders   # per-run events / retries / quality issues

To make the second run only export rows that changed, switch the export to incremental mode with a cursor_column: (must be monotonically increasing — usually updated_at or a sequence id):

exports:
  - name: orders
    query: "SELECT id, name, updated_at FROM orders"
    mode: incremental
    cursor_column: updated_at
    format: parquet
    skip_empty: true            # a run with no new rows reports `skipped`
    destination:
      type: local
      path: ./output

Subsequent rivet run invocations will only fetch rows with updated_at > the stored cursor. For tables larger than ~5 M rows, switch to mode: chunked instead — see modes/chunked.md.


5 · Many tables: plan once, apply by waves

When a config has several exports, rivet plan assigns each one a wave — a priority band derived from its size, chunking strategy, and risk (ADR-0006). By default rivet plan is read-only: it prints the schedule for you to review but does not touch the config. Add --annotate-waves to write the wave: / parallel_safe: fields back into the config, where you can see and hand-edit them:

rivet plan -c rivet.yaml                   # review the schedule (read-only)
rivet plan -c rivet.yaml --annotate-waves  # write `wave: N` onto every export, in place
exports:
  - name: users
    wave: 1        # small / cheap → runs first
    # …
  - name: events
    wave: 3        # large → runs later
    # …

rivet apply then runs the whole config wave by wave, lowest first, with a barrier between waves — every export in wave 1 finishes before wave 2 starts. Exports with no wave: run last:

rivet apply rivet.yaml          # a .yaml path → wave-ordered execution

(A .json path still means the sealed single-artifact replay — see reference/cli.md § rivet apply.) The plan suggests the waves; you stay in control — hand-edit wave: and apply respects your order.

Parallel within a wave — only where it’s safe

Add parallel_export_processes: true (or pass rivet apply --parallel-export-processes) and, within each wave, the cheap exports — the ones rivet plan marked parallel_safe: true (cost class Low, under ~100K rows) — run concurrently as separate processes. A heavier export already chunk-parallelizes its own ranges internally, so it runs alone in its wave: two big tables at once would multiply the load on the source. The wave stays bounded because only the cheap parallel_safe exports run concurrently, and each child honors its own batch/memory caps (the adaptive back-pressure governor is a separate opt-in: tuning.adaptive: true with parallel > 1).

parallel_export_processes: true   # top-level: parallelize the cheap (parallel_safe) exports within each wave

Load into BigQuery or Snowflake (optional)

Rivet stops at typed Parquet by default. To load it into a warehouse, add a top-level load: block to the same config and run rivet load — the target table, column types, and source files are all derived from the export (nothing hand-typed):

# rivet.yaml — the export above, plus a load target
load:
  target: bigquery          # or: snowflake (+ connection / warehouse / database / schema / storage_integration)
  project: my-gcp-project
  dataset: analytics
  cleanup_source: true      # wipe the staged Parquet once the load is row-count-verified
rivet run  -c rivet.yaml    # extract → GCS
rivet load -c rivet.yaml    # load → warehouse (native types; count-gated before any cleanup)

The load follows the export’s mode: — full overwrites the latest snapshot; incremental / cdc append to <table>__changes and expose a current-state dedup view keyed on the source primary key rivet run recorded (set pk: [id] in the load: block for a query: export or to override it). Recipes: snowflake-load.md · cdc-bigquery-load.md.


When something is wrong

Rivet tries to fail early and say exactly what to fix — most mistakes are caught at check / doctor time, before a single row is read.

A query that references a table (or column) that doesn’t exist is caught by rivet check — it exits non-zero with the offending name and SQLSTATE, instead of passing through to a half-finished run:

A typo in a config field is caught at parse time with a Did you mean …? suggestion that names the line:

An unreachable database — down, wrong host/port, or a tunnel that isn’t up — is reported by rivet doctor with a reachability hint before you waste a run:

More failure modes (retries, schema drift, crash/resume) and exactly what rivet does for each: semantics.md.


Next steps

When you need to …Go to
Pick the right export mode for each tablemodes/ — full · incremental · chunked · time_window · cdc
Configure S3 / GCS / Azure / stdout destinationsdestinations/
Load exports into BigQuery / Snowflakerecipes/snowflake-load.md · cdc-bigquery-load.md
Look up a YAML field or a CLI flagreference/config.md · reference/cli.md
Understand run_id / cursor / chunk / manifest / journalconcepts.md
Tune for memory, throughput, source pressurereference/tuning.md · best-practices/
Take it to production (read replicas, poolers, monitoring)pilot/production-checklist.md
Run a serious pilot (chunked + reconcile + repair on your data)pilot/pilot-walkthrough.md
See exactly what happens under retry / crash / resumesemantics.md
Auditable plan/apply workflow for CI/CDreference/cli.md § rivet plan · ADR-0005