Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Best Practices

Practical guidance for using Rivet’s resource-aware extraction capabilities. These guides go beyond the reference documentation to explain why settings matter and when to use them.

The tuning and compression settings shown here apply to every Rivet source (PostgreSQL, MySQL, SQL Server, MongoDB) and every mode (full, incremental, chunked, time_window, cdc) — the quick-start examples below use PostgreSQL + incremental only for concreteness. Quality checks are the exception: on the multi-part runners (chunked, keyset, parallel-Mongo) only row_count bounds are enforced; null_ratio_max and unique_columns are single-runner only (each part processes independently). See quality-checks.md.

GuideWhat it covers
Resource-aware extractionMemory budgets, batch cap policies (warn/fail/auto_shrink), RSS formula
Parquet tuningRow group strategies, target sizes, downstream read implications
Compression profilesProfile-to-codec mapping, CPU/size trade-offs, when to use each
Quality checksRow count gates, null ratio, uniqueness tracking, unique_max_entries cap
Low-memory runnersSettings for 512 MB–4 GB hosts; auto_shrink guarantees and caveats
Gentle SQL Server extractionEasy on the source DB and the worker; why chunk_size (not chunk_size_memory_mb) on MSSQL — config: rivet_mssql_gentle.yaml
Recovery and resume--resume semantics, crash recovery, state inspection
Benchmark methodologyHow to run E2E and Criterion benchmarks, interpret results, compare versions

Quick-start recipes

Safe production export

source:
  type: postgres
  url_env: DATABASE_URL
  tuning:
    profile: balanced

exports:
  - name: orders
    query: "SELECT * FROM orders"
    mode: incremental
    cursor_column: updated_at
    format: parquet
    compression_profile: balanced
    destination:
      type: local
      path: ./out
    parquet:
      row_group_strategy: auto
      target_row_group_mb: 128
    quality:
      row_count_min: 1
      unique_columns: [id]
      unique_max_entries: 1000000
    tuning:
      max_batch_memory_mb: 256
      on_batch_memory_exceeded: warn

Low-memory runner (≤ 512 MB RAM)

tuning:
  profile: safe
  max_batch_memory_mb: 64
  on_batch_memory_exceeded: auto_shrink
parquet:
  row_group_strategy: auto
  target_row_group_mb: 32
  max_row_group_mb: 64
compression_profile: fast

CI strict mode

tuning:
  max_batch_memory_mb: 128
  on_batch_memory_exceeded: fail
quality:
  row_count_min: 100
  unique_columns: [id]
  unique_max_entries: 500000