Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Full Export Mode

When to use

Use mode: full when you want a complete snapshot of the query result set every time. Each run re-exports all rows from scratch. Best for:

  • Small-to-medium tables (up to a few million rows)
  • Reference/dimension tables that need a fresh copy each day
  • One-time data migrations

MongoDB. full is MongoDB’s primary batch mode — the SQL-runner modes (incremental / chunked / time_window) do not apply to a document store. Within full, MongoDB adds keyset (seek) paging, parallel: N _id-range fan-out, and resume — all keyed on _id (source.mongo.page_size / parallel). See ../reference/mongodb.md.

Minimal config

source:
  type: postgres                                    # postgres, mysql, mssql, or mongo
  url: "postgresql://user:pass@host:5432/dbname"

exports:
  - name: users_daily                               # unique export name
    query: "SELECT id, name, email, created_at FROM users"
    mode: full                                      # re-export everything each run
    format: parquet                                 # parquet or csv
    destination:
      type: local
      path: ./output                                # directory for output files

Output file: ./output/users_daily_20260406_120000_123.parquet (the timestamp ends with a 3-digit millisecond field, e.g. _123, so rapid re-runs never collide)

Run it

# 1. Verify config and connectivity
rivet check --config users.yaml
rivet doctor --config users.yaml

# 2. Run with validation
rivet run --config users.yaml --validate --reconcile

# 3. Check results
rivet metrics --config users.yaml --last 5

What happens

  1. Rivet connects to the source database
  2. Executes SELECT id, name, email, created_at FROM users
  3. Fetches rows in batches controlled by tuning.batch_size (default 10,000 with balanced profile)
  4. Writes a timestamped output file
  5. Records metrics in the state database

No cursor is stored. Each run produces a new file with all rows.

batch_size directly controls memory usage and source load. For wide tables (many columns, TEXT/JSONB fields), reduce it to 1,000-5,000. See reference/tuning.md.

Common options

exports:
  - name: users_daily
    query: "SELECT * FROM users"
    mode: full
    format: parquet
    compression: zstd           # zstd (default), snappy, gzip, lz4, none
    skip_empty: true            # a 0-row run reports `skipped`, not `success`
    max_file_size: "512MB"      # split into multiple files if output exceeds this
    meta_columns:
      exported_at: true         # add _rivet_exported_at column
    destination:
      type: local
      path: ./output
    tuning:
      profile: safe             # safe/balanced/fast — controls batch size, timeouts

Troubleshooting

Export is slow on a large table – Switch to mode: chunked with parallel: 4 for tables over 1M rows. See chunked.md.

Output file is too large – Add max_file_size: "256MB" to split into parts.

A run that read 0 rows reports success – Add skip_empty: true to record it as skipped. No file is written for 0 rows either way, and a skipped run leaves the prefix describing the last run that delivered — so a later full rivet load keeps the previous data rather than emptying the table.