Keyboard shortcuts

Press ← or β†’ to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

YAML Config Guide

πŸ“Œ The exhaustive, always-current field reference is generated from the code: config-reference.md β€” rendered from rivet schema config (the schemars-derived JSON Schema), so it cannot drift and needs no manual verification. This page is the guide: the same options with defaults, rationale, and worked examples. For the guaranteed-current field/type/enum list, trust the generated reference.

The most-used options, grouped by section, with the why and examples.


Root

FieldTypeRequiredDescription
sourceobjectyesDatabase connection and global tuning
exportslistyesOne or more export definitions
notificationsobjectnoSlack / webhook notification settings

source

FieldTypeRequiredDefaultDescription
typepostgres | mysql | mssql | mongo | oracleyesβ€”Database type. mssql = SQL Server (URL scheme sqlserver://); mongo = MongoDB (URL scheme mongodb://, see mongodb.md); oracle = Oracle Database (URL scheme oracle://…/SERVICE, see oracle.md).
urlstringone of url/url_env/url_file or structuredβ€”Full connection URL (postgresql:// / mysql:// / sqlserver:// / mongodb:// / oracle://)
url_envstringβ€”Env var name containing the URL
url_filestringβ€”Path to file containing the URL
hoststringfor structuredβ€”Database hostname
portintegerno5432 (PG) / 3306 (MySQL) / 1433 (MSSQL) / 27017 (MongoDB) / 1521 (Oracle)Database port
userstringfor structuredβ€”Database user
passwordstringnoβ€”Not recommended β€” plaintext; see Credentials & plan artifacts below
password_envstringnoβ€”Env var name containing the password (recommended)
databasestringfor structuredβ€”Database name
tuningobjectnoβ€”Global tuning (see tuning.md)
tlsobjectnoβ€”Transport security (see TLS below). Omit β†’ plaintext + WARN log.

Connection approaches (mutually exclusive):

  1. URL-based: provide exactly one of url, url_env, or url_file
  2. Structured: provide host, user, database (+ optional port, password/password_env)

TLS

FieldTypeDefaultDescription
modedisable | require | verify-ca | verify-fullverify-fullEnforcement level (mirrors libpq sslmode semantics)
ca_filestringβ€”PEM-encoded CA certificate for private trust stores; required for verify-ca/verify-full against custom CAs
accept_invalid_certsbooleanfalseDangerous β€” disables certificate verification. Only honored when explicitly true.
accept_invalid_hostnamesbooleanfalseDangerous β€” disables hostname (SAN/CN) verification. Only honored when explicitly true.

Example (production):

source:
  type: postgres
  url_env: DATABASE_URL
  tls:
    mode: verify-full
    ca_file: /etc/ssl/certs/rds-ca-2019-root.pem

Example (local dev only β€” no TLS):

source:
  type: mysql
  host: 127.0.0.1
  port: 3306
  user: dev
  password_env: DEV_PWD
  database: rivet
  tls: { mode: disable }       # explicit opt-out β€” silences the plaintext WARN

Example (SQL Server β€” sqlserver:// scheme, port 1433):

source:
  type: mssql
  url_env: MSSQL_URL           # sqlserver://user:pass@host:1433/database
  tls:
    ca_file: /etc/ssl/certs/your-sql-server-ca.pem   # private CA, or:
    # accept_invalid_certs: true                      # self-signed dev cert

SQL Server always encrypts the login handshake, so TLS is on regardless; the tls: block only controls how the server certificate is trusted. Supported export modes and types are listed in compatibility.md.

When tls: is omitted entirely, Rivet connects without TLS and emits a WARN so you notice. See reference/compatibility.md for which servers ship TLS-ready and Rivet’s dev-environment defaults.

Credentials & plan artifacts

A PlanArtifact (produced by rivet plan) is designed to be committed / reviewed; it must not carry plaintext credentials. Rivet enforces ADR-0005 PA9 (SourceConfig::redact_for_artifact):

  • password: field β†’ always stripped from the artifact (set to None).
  • url: containing scheme://user:pass@… β†’ userinfo rewritten to REDACTED.
  • password_env / url_env / url_file β†’ preserved as references so apply-time can re-resolve against the apply-environment.

When redaction runs, rivet plan logs:

WARN plan 'orders': plaintext credentials stripped from artifact β€”
     apply time must have equivalent env/file-based auth available

Recommendation: use password_env (or url_env) everywhere; only use plaintext password: for one-off local scripts. See ADR-0005 PA9.


exports[]

Each entry in the exports list defines one export job.

FieldTypeRequiredDefaultDescription
namestringyesβ€”Unique identifier for this export
querystringone of query/query_file/table/tablesβ€”Inline SQL SELECT query
query_filestringβ€”Path to .sql file (relative to config dir)
tablestringβ€”Whole-table shortcut (name or schema.table) β€” enables PK auto-chunking; required for chunk_by_key / chunk_size_memory_mb
tableslistβ€”CDC only (mode: cdc): capture several tables through one change stream (one slot/binlog connection); rejected at config load for batch exports (batch is one query/table per export). Mutually exclusive with table:; not supported for SQL Server
modefull | incremental | chunked | time_window | cdcnofullExport mode. cdc = log-based change data capture (cdc.md). MongoDB supports full + cdc only (a document store has no chunked/incremental/time_window).
formatparquet | csvyesβ€”Output format
compressionzstd | snappy | gzip | lz4 | nonenozstdCompression codec (low-level; prefer compression_profile)
compression_levelintegernocodec defaultCompression level (low-level; prefer compression_profile)
compression_profilenone | fast | balanced | compactnoβ€”High-level preset β€” overrides compression and compression_level. See Compression profiles below.
destinationobjectyesβ€”Where to write output (see below)
verifysize | contentnosizeIntegrity depth required of --validate. content checks every part’s MD5 against the store’s listing (no download) and fails validation for any part only size-verified β€” e.g. a part too large to upload as a single PUT (lower max_file_size so it fits) or a backend that exposes no checksum (local FS, streamed multipart). See Verification depth below.
skip_emptybooleannofalseRecord a 0-row batch run as skipped instead of success (no file is written for 0 rows either way; a full load then keeps the previous data). Not read by mode: cdc
max_file_sizestringnoβ€”Split output: "256MB", "1GB", etc.
waveintegernoβ€”Advisory execution wave (1 = highest priority, runs first). Written by rivet plan from the source-aware prioritization score (ADR-0006); consumed by rivet apply <config>, which runs exports wave-by-wave in ascending order (no wave: runs last). Hand-editable; a later rivet plan refreshes it.
parallel_safebooleannoβ€”Whether this export is cheap enough (cost class Low, < ~100K rows, and not isolate_on_source) to run concurrently with its wave-mates under rivet apply --parallel-export-processes. Written by rivet plan; a heavier export runs alone in its wave (it already chunk-parallelizes internally). Hand-editable.
meta_columnsobjectnoβ€”Extra columns added to output
qualityobjectnoβ€”Data quality checks
tuningobjectnoβ€”Per-export tuning overrides
source_groupstringnoβ€”Logical group for shared source capacity (replica, host). Drives campaign-level warnings in rivet plan (advisory only β€” ADR-0006)
reconcile_requiredbooleannofalseAdvisory hint: treat this export as reconcile-sensitive in planning, independent of the --reconcile CLI flag (ADR-0006, Epic C)
columnsmapnoβ€”Per-column type overrides (see below)
on_schema_driftwarn|continue|failnowarnPolicy when structural schema drift is detected (see below)
shape_drift_warn_factorfloatno2.0Warn when a string/binary column’s max byte length grows beyond N Γ— stored_max. Set to 0 to disable shape tracking.
parquetobjectnoβ€”Parquet row group tuning (Parquet format only). See Parquet row group tuning below.

Compression profiles

compression_profile is the recommended way to pick a codec. It maps to a (codec, level) pair and takes precedence over any compression / compression_level fields.

ProfileCodecLevelBest for
noneno compressionβ€”Debug, local scratch, fast iteration
fastsnappyβ€”Backfills, pilot runs, low-CPU environments
balancedzstd3Default for production β€” good ratio, moderate CPU
compactzstd9Storage- or network-cost-sensitive pipelines
exports:
  - name: events
    format: parquet
    compression_profile: balanced    # zstd level 3
    destination: { type: local, path: ./out }

If you need a specific codec that is not covered by the presets, use compression + compression_level directly and omit compression_profile.


Verification depth

verify controls how thoroughly --validate (and rivet validate) checks each part at the destination:

  • size (default) β€” confirm each part exists at its recorded size_bytes, plus manifest self-consistency and _SUCCESS. Content is also MD5-checked for free whenever the store surfaces a checksum in its listing, but a part without one is accepted as size-only.
  • content β€” require every part’s content MD5 to match the store’s listing checksum (no download). Any part that could only be size-verified fails validation with an actionable message.

How content verification works: Rivet computes each part’s MD5 before upload and records it in the manifest; GCS and Azure compute their own for a part uploaded as a single PUT and return it in object listings. --validate compares the two with no download. Parts large enough to stream as multipart / block-list get no checksum, and neither does S3 (its ETag is not an MD5 under SSE-KMS / SSE-C, so rivet does not trust it) or local FS. Under verify: content on GCS / Azure, set destination.oneshot_budget_mb (default 64 MB) comfortably above your part size: the budget is shared by concurrent uploads, so a part one-shots only if it fits what is free at that moment. S3 cannot meet verify: content.

The run report and rivet validate show coverage explicitly, e.g. 3 verified (2 md5, 1 size-only).


Parquet row group tuning

Parquet row groups affect memory usage during write, compression ratio, and downstream query performance (predicate pushdown, column skipping). When parquet: is omitted, Rivet uses the library default of 1,048,576 rows per group, which is optimal for narrow tables but can be large for wide tables.

exports:
  - name: events
    format: parquet
    parquet:
      row_group_strategy: auto          # auto | fixed_rows | fixed_memory
      target_row_group_mb: 128          # target Arrow buffer size per group (auto + fixed_memory)
      max_row_group_mb: 256             # optional upper bound (all strategies)
FieldTypeDefaultDescription
row_group_strategyauto | fixed_rows | fixed_memoryautoHow to determine row group size
row_group_rowsintegerβ€”Exact rows per group; used with fixed_rows only
target_row_group_mbinteger128Target Arrow buffer per group in MB; used with auto and fixed_memory
max_row_group_mbintegerβ€”Hard upper bound on group memory in MB (all strategies)
StrategyBehavior
autoEstimates row width from schema column types, computes rows-per-group to hit target_row_group_mb. Narrow tables get large groups; wide tables get smaller groups.
fixed_rowsUse row_group_rows exactly. Simple and deterministic, but does not adapt to row width.
fixed_memorySame math as auto (target / estimated row bytes), but the strategy name is explicit in logs.

Examples:

# Auto-tune for a wide JSON table β€” groups sized to ~64 MB
parquet:
  row_group_strategy: auto
  target_row_group_mb: 64
  max_row_group_mb: 128

# Fixed row count β€” useful when downstream tooling requires exact group sizes
parquet:
  row_group_strategy: fixed_rows
  row_group_rows: 500000

Note: rivet plan shows the selected strategy and target in the Format section when parquet: is configured.

rivet init auto-generates this block for chunked exports and large full-mode tables, pre-selecting target_row_group_mb: 64 for wide schemas (β‰₯ 5 text/JSON/bytea columns) and 128 for narrow ones.


exports[].on_schema_drift β€” schema drift policy

Controls what Rivet does when it detects a structural change in the output schema (column added, removed, or retyped) compared to the snapshot stored from the previous run.

ValueBehavior
warn(default) Log a warning, store the new schema fingerprint, and continue the run.
continueSilently accept β€” store the new schema, no log output.
failAbort the run with exit code 4 (the schema-drift exit class). The schema store is not updated, so the next run will detect the same change again.

fail is useful in CI pipelines where schema changes must be reviewed before the new shape is exported downstream.

exports:
  - name: orders
    on_schema_drift: fail

When fail triggers, behavior depends on the runner: in single, keyset, and parallel-Mongo modes the schema check runs post-extraction, so the output file has already been written to the destination (but no cursor advance or manifest commit occurs). In chunked mode the check runs pre-chunk from a scan-free type probe, so the run aborts before any chunk is written. Re-run after confirming the schema change is intentional, or switch to warn to accept it.


exports[].columns β€” per-column type overrides

Override the Arrow type Rivet infers for a specific column. Useful when:

  • a NUMERIC / DECIMAL column has no explicit precision/scale in the source schema (beyond rivet init’s default decimal(38,18) placeholder), or
  • you need a narrower precision for BigQuery NUMERIC compatibility.
columns:
  <column_name>: <type>          # applies to every captured table with the column
  "<table>.<column>": <type>     # applies to ONE table, wins over the bare key

Supported override types: decimal(p,s) / numeric(p,s) (both precision and scale required; precision > 38 produces Decimal256), integer widths (smallint/int/bigint/int16/int32/int64), floats (real/float4/double/float8), bool, text/string/varchar, uuid, json/jsonb, date, the timestamp* family (naive / tz / _ns variants), and binary (bytea/binary/varbinary/blob).

Overrides apply to batch and CDC identically (the same resolution surface). Key shapes: a bare column name applies to every captured table that has the column β€” on a multi-table CDC export that means ALL of them; a qualified "table.column" key targets one table and wins over the bare key there. A qualified key naming a table the export does not capture β€” or used on a query-shaped export β€” is a config error at load.

Example:

exports:
  - name: orders
    query: "SELECT id, amount, fee FROM orders"
    format: parquet
    destination:
      type: local
      path: ./out
    columns:
      amount: decimal(18,2)
      fee: decimal(18,6)

rivet init generates these automatically. When introspecting a table, rivet init reads numeric_precision and numeric_scale from information_schema.columns. If both are present, it emits a concrete override (decimal(p,s)). If the column is unbounded (NUMERIC without explicit precision), rivet init emits a working default decimal(38,18) plus a # REVIEW: YAML comment β€” the config header adds a # NOTE: line, and rivet init -o … prints a stderr reminder so you tighten precision when you know the real domain rules:

    columns:
      price: decimal(38,18)  # REVIEW: DDL has no numeric(p,s); edit to the real decimal(p,s) …

Type overrides are applied at export time and are reflected in rivet check --type-report output.


Some PostgreSQL types have no Arrow representation and cannot be exported directly. Rivet will report an error listing all unmappable columns before the run starts.

PostgreSQL typeReasonWorkaround
geometry (PostGIS)No Arrow equivalentCast to text: ST_AsText(col) AS col in your query
geography (PostGIS)No Arrow equivalentCast to text: ST_AsText(col) AS col
hstoreNo Arrow equivalentCast to JSON text: hstore_to_json(col)::text AS col
tsvector, tsqueryNo Arrow equivalentCast to text: col::text AS col
point, line, polygon, etc.No Arrow equivalentCast to text: col::text AS col

Use a SQL expression in your query field to work around any unsupported type:

exports:
  - name: locations
    query: >
      SELECT id, name, ST_AsText(geom) AS geom_wkt
      FROM locations
    format: parquet
    destination:
      type: local
      path: ./out

Rivet exports the WKT text as a Utf8 (string) column. Downstream tools (DuckDB, GeoPandas, QGIS) can reconstruct geometry from WKT.


Mode-specific fields

Incremental (mode: incremental):

FieldTypeRequiredDefaultDescription
cursor_columnstringyesβ€”Primary progression column. Must be strictly per-row-distinct and monotonically increasing β€” resume uses WHERE cursor > last_value, so rows that tie on the high-watermark value and become visible after it is passed are skipped. A low-resolution updated_at (second granularity) can tie; prefer a sequence/identity id or a sub-value-unique timestamp. See semantics.md β†’ Known non-guarantees.
cursor_fallback_columnstringwhen coalesceβ€”Fallback column used when primary is NULL. Only valid with incremental_cursor_mode: coalesce
incremental_cursor_modesingle_column | coalescenosingle_columncoalesce progresses on COALESCE(primary, fallback). See modes/incremental-coalesce.md and ADR-0007.
settleobjectnoβ€”Hold rows back until they stop changing: { after: 1h } ages the cursor itself, { after: 1h, column: server_time } ages another date/timestamp column. A row exports only once it is older than after (s/m/h/d) by the source clock. See modes/incremental.md Β§ Settle window.

Chunked (mode: chunked):

FieldTypeRequiredDefaultDescription
chunk_columnstringyes*β€”Numeric or date/timestamp column to partition by. *Required unless chunk_by_key is set (mutually exclusive).
chunk_by_keystringyes*β€”Single index-backed UNIQUE NOT NULL column for keyset (seek) pagination β€” the source-safe shape for tables with no single-integer PK (UUID / string / composite). Requires the table: shortcut; mutually exclusive with chunk_column. See chunked modes and ADR-0020.
chunk_sizeintegerno100000Rows per chunk (numeric mode), or page size for keyset. Ignored when chunk_count is set.
chunk_size_memory_mbintegernoβ€”Target memory budget per chunk in MB; chunk_size is derived from a per-engine row-size estimate, clamped to [10000, 5000000] rows. Works on PostgreSQL, MySQL and SQL Server (PG: pg_relation_size / reltuples; MySQL: information_schema AVG_ROW_LENGTH with InnoDB overflow correction; SQL Server: no estimate, falls back to 512 B/row with a warning). Requires the table: shortcut, mutually exclusive with an explicit non-default chunk_size:.
chunk_countintegernoβ€”Divide the column range into exactly this many equal chunks. chunk_size is computed dynamically from min/max. Must be β‰₯ 1. Mutually exclusive with chunk_by_days.
chunk_by_daysintegernoβ€”Enable date chunking: window size in days. Mutually exclusive with chunk_count.
parallelintegerno1Concurrent chunk workers
chunk_densebooleannofalseRemoved. true is refused at config load (it skipped or duplicated rows under concurrent writes); use chunk_by_key or chunk_column.
chunk_checkpointbooleannofalsePersist per-chunk progress for resume
chunk_max_attemptsintegernoβ€”Max retry attempts per chunk

Time-window (mode: time_window):

FieldTypeRequiredDefaultDescription
time_columnstringyesβ€”Timestamp column to filter on
time_column_typetimestamp | unixnotimestampColumn type
days_windowintegeryesβ€”Rolling window size in days

exports[] β€” value-based partitioning

Splits a full, chunked or incremental export’s rows into Hive-style col=value/ destination sub-folders by a date column. See partitioning.md.

FieldTypeRequiredDefaultDescription
partition_bystringnoβ€”Date/timestamp column to bucket rows by. Requires a {partition} token in destination.path/prefix. NULLs β†’ col=__HIVE_DEFAULT_PARTITION__/. Not compatible with mode: time_window, mode: cdc, chunk_by_key, a load: block (per-export or top-level), or a MongoDB source.
partition_granularityday | month | yearnodayBucket width.

exports[].meta_columns

FieldTypeDefaultDescription
exported_atbooleanfalseAdd _rivet_exported_at column (Timestamp UTC; one value captured at sink construction and shared by every batch/row that sink writes β€” effectively one value per export run in single mode, per chunk (or keyset page) in multi-part modes)
row_hashboolean or list of column namesfalseAdd _rivet_row_hash column β€” lower 64 bits of xxHash3-128, written as Int64 for fast PARTITION BY / JOIN. true hashes every column; a list (row_hash: [id, status, updated_at]) hashes exactly those columns in that order and records the covered set in the run manifest. Deterministic across runs; distinguishes NULL from empty string.

exports[].quality

FieldTypeDescription
row_count_minintegerFail if fewer rows exported
row_count_maxintegerFail if more rows exported
null_ratio_maxmap (column β†’ float)Fail if null ratio exceeds threshold
unique_columnslist of stringsFail if values are not unique
unique_max_entriesintegerCap on distinct values tracked per column during uniqueness checks. When reached, a Warn is emitted and checking stops for that column; duplicates already found before the cap still fail the run. Prevents unbounded memory growth on high-cardinality columns (UUIDs, email addresses, event IDs).

Uniqueness tracking uses typed xxHash3-64 internally β€” numeric and binary columns are hashed directly from raw bytes without string formatting. unique_max_entries is the primary knob to control memory on very large tables.

Example:

quality:
  row_count_min: 100
  null_ratio_max:
    email: 0.05          # email must be <5% null
  unique_columns:
    - id
    - email
  unique_max_entries: 1000000   # stop after 1M unique values; warn if limit hit

Without unique_max_entries β€” tracking is unbounded. Safe for tables with hundreds of thousands of rows; may use significant RAM on tables with tens or hundreds of millions of distinct values.

With unique_max_entries β€” tracking stops at the limit and the run summary shows a warning. Duplicates found before the limit still fail the run; ones past it go unseen. Use when you want a best-effort uniqueness check without memory risk.


exports[].destination

The complete per-backend field list (local / s3 / gcs / azure / stdout) is in the generated config-reference.md (section exports[].destination). Per-backend setup, auth flows, and permissions: destinations/ β€” local Β· s3 Β· gcs Β· azure Β· stdout, plus the cloud auth matrix.

Path and prefix placeholders

The path (local) and prefix (S3 / GCS) fields support template placeholders, substituted at plan-build time:

PlaceholderValue
{date}UTC date as YYYY-MM-DD
{export}Export name from config
{table}Alias for {export}
{run_id}The run’s unique id β€” substituted only by rivet validate --run-id, which re-targets validation at that run’s prefix. run and apply resolve destinations without a run id, so the token is always left verbatim there and the destination open fails fast rather than aliasing to an unintended prefix (and rivet load refuses a {run_id} prefix outright β€” it cannot know which run’s output to load).
destination:
  type: s3
  bucket: my-data
  prefix: exports/{date}/{export}/
  region: us-east-1

With an export named orders running on 2026-05-14, this resolves to exports/2026-05-14/orders/.


notifications

FieldTypeDescription
slackobjectSlack notification config

notifications.slack

FieldTypeDescription
webhook_urlstringSlack incoming webhook URL
webhook_url_envstringEnv var containing webhook URL
onlistEvents to notify on: failure, schema_change, degraded

Example:

notifications:
  slack:
    webhook_url_env: SLACK_WEBHOOK
    on: [failure, schema_change]

Environment variable interpolation

Any string value can reference environment variables:

source:
  url: "postgresql://${DB_USER}:${DB_PASS}@${DB_HOST}:5432/mydb"

Query parameters

Queries can use ${key} placeholders filled by --param key=value:

exports:
  - name: filtered
    query: "SELECT * FROM orders WHERE region = '${region}'"
rivet run --config export.yaml --param region=us-east