YAML Config Guide
π The exhaustive, always-current field reference is generated from the code: config-reference.md β rendered from
rivet schema config(the schemars-derived JSON Schema), so it cannot drift and needs no manual verification. This page is the guide: the same options with defaults, rationale, and worked examples. For the guaranteed-current field/type/enum list, trust the generated reference.
The most-used options, grouped by section, with the why and examples.
Root
| Field | Type | Required | Description |
|---|---|---|---|
source | object | yes | Database connection and global tuning |
exports | list | yes | One or more export definitions |
notifications | object | no | Slack / webhook notification settings |
source
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
type | postgres | mysql | mssql | mongo | oracle | yes | β | Database type. mssql = SQL Server (URL scheme sqlserver://); mongo = MongoDB (URL scheme mongodb://, see mongodb.md); oracle = Oracle Database (URL scheme oracle://β¦/SERVICE, see oracle.md). |
url | string | one of url/url_env/url_file or structured | β | Full connection URL (postgresql:// / mysql:// / sqlserver:// / mongodb:// / oracle://) |
url_env | string | β | Env var name containing the URL | |
url_file | string | β | Path to file containing the URL | |
host | string | for structured | β | Database hostname |
port | integer | no | 5432 (PG) / 3306 (MySQL) / 1433 (MSSQL) / 27017 (MongoDB) / 1521 (Oracle) | Database port |
user | string | for structured | β | Database user |
password | string | no | β | Not recommended β plaintext; see Credentials & plan artifacts below |
password_env | string | no | β | Env var name containing the password (recommended) |
database | string | for structured | β | Database name |
tuning | object | no | β | Global tuning (see tuning.md) |
tls | object | no | β | Transport security (see TLS below). Omit β plaintext + WARN log. |
Connection approaches (mutually exclusive):
- URL-based: provide exactly one of
url,url_env, orurl_file - Structured: provide
host,user,database(+ optionalport,password/password_env)
TLS
| Field | Type | Default | Description |
|---|---|---|---|
mode | disable | require | verify-ca | verify-full | verify-full | Enforcement level (mirrors libpq sslmode semantics) |
ca_file | string | β | PEM-encoded CA certificate for private trust stores; required for verify-ca/verify-full against custom CAs |
accept_invalid_certs | boolean | false | Dangerous β disables certificate verification. Only honored when explicitly true. |
accept_invalid_hostnames | boolean | false | Dangerous β disables hostname (SAN/CN) verification. Only honored when explicitly true. |
Example (production):
source:
type: postgres
url_env: DATABASE_URL
tls:
mode: verify-full
ca_file: /etc/ssl/certs/rds-ca-2019-root.pem
Example (local dev only β no TLS):
source:
type: mysql
host: 127.0.0.1
port: 3306
user: dev
password_env: DEV_PWD
database: rivet
tls: { mode: disable } # explicit opt-out β silences the plaintext WARN
Example (SQL Server β sqlserver:// scheme, port 1433):
source:
type: mssql
url_env: MSSQL_URL # sqlserver://user:pass@host:1433/database
tls:
ca_file: /etc/ssl/certs/your-sql-server-ca.pem # private CA, or:
# accept_invalid_certs: true # self-signed dev cert
SQL Server always encrypts the login handshake, so TLS is on regardless; the
tls: block only controls how the server certificate is trusted. Supported
export modes and types are listed in compatibility.md.
When tls: is omitted entirely, Rivet connects without TLS and emits a WARN so you notice. See reference/compatibility.md for which servers ship TLS-ready and Rivetβs dev-environment defaults.
Credentials & plan artifacts
A PlanArtifact (produced by rivet plan) is designed to be committed / reviewed; it must not carry plaintext credentials. Rivet enforces ADR-0005 PA9 (SourceConfig::redact_for_artifact):
password:field β always stripped from the artifact (set toNone).url:containingscheme://user:pass@β¦β userinfo rewritten toREDACTED.password_env/url_env/url_fileβ preserved as references so apply-time can re-resolve against the apply-environment.
When redaction runs, rivet plan logs:
WARN plan 'orders': plaintext credentials stripped from artifact β
apply time must have equivalent env/file-based auth available
Recommendation: use password_env (or url_env) everywhere; only use plaintext password: for one-off local scripts. See ADR-0005 PA9.
exports[]
Each entry in the exports list defines one export job.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
name | string | yes | β | Unique identifier for this export |
query | string | one of query/query_file/table/tables | β | Inline SQL SELECT query |
query_file | string | β | Path to .sql file (relative to config dir) | |
table | string | β | Whole-table shortcut (name or schema.table) β enables PK auto-chunking; required for chunk_by_key / chunk_size_memory_mb | |
tables | list | β | CDC only (mode: cdc): capture several tables through one change stream (one slot/binlog connection); rejected at config load for batch exports (batch is one query/table per export). Mutually exclusive with table:; not supported for SQL Server | |
mode | full | incremental | chunked | time_window | cdc | no | full | Export mode. cdc = log-based change data capture (cdc.md). MongoDB supports full + cdc only (a document store has no chunked/incremental/time_window). |
format | parquet | csv | yes | β | Output format |
compression | zstd | snappy | gzip | lz4 | none | no | zstd | Compression codec (low-level; prefer compression_profile) |
compression_level | integer | no | codec default | Compression level (low-level; prefer compression_profile) |
compression_profile | none | fast | balanced | compact | no | β | High-level preset β overrides compression and compression_level. See Compression profiles below. |
destination | object | yes | β | Where to write output (see below) |
verify | size | content | no | size | Integrity depth required of --validate. content checks every partβs MD5 against the storeβs listing (no download) and fails validation for any part only size-verified β e.g. a part too large to upload as a single PUT (lower max_file_size so it fits) or a backend that exposes no checksum (local FS, streamed multipart). See Verification depth below. |
skip_empty | boolean | no | false | Record a 0-row batch run as skipped instead of success (no file is written for 0 rows either way; a full load then keeps the previous data). Not read by mode: cdc |
max_file_size | string | no | β | Split output: "256MB", "1GB", etc. |
wave | integer | no | β | Advisory execution wave (1 = highest priority, runs first). Written by rivet plan from the source-aware prioritization score (ADR-0006); consumed by rivet apply <config>, which runs exports wave-by-wave in ascending order (no wave: runs last). Hand-editable; a later rivet plan refreshes it. |
parallel_safe | boolean | no | β | Whether this export is cheap enough (cost class Low, < ~100K rows, and not isolate_on_source) to run concurrently with its wave-mates under rivet apply --parallel-export-processes. Written by rivet plan; a heavier export runs alone in its wave (it already chunk-parallelizes internally). Hand-editable. |
meta_columns | object | no | β | Extra columns added to output |
quality | object | no | β | Data quality checks |
tuning | object | no | β | Per-export tuning overrides |
source_group | string | no | β | Logical group for shared source capacity (replica, host). Drives campaign-level warnings in rivet plan (advisory only β ADR-0006) |
reconcile_required | boolean | no | false | Advisory hint: treat this export as reconcile-sensitive in planning, independent of the --reconcile CLI flag (ADR-0006, Epic C) |
columns | map | no | β | Per-column type overrides (see below) |
on_schema_drift | warn|continue|fail | no | warn | Policy when structural schema drift is detected (see below) |
shape_drift_warn_factor | float | no | 2.0 | Warn when a string/binary columnβs max byte length grows beyond N Γ stored_max. Set to 0 to disable shape tracking. |
parquet | object | no | β | Parquet row group tuning (Parquet format only). See Parquet row group tuning below. |
Compression profiles
compression_profile is the recommended way to pick a codec. It maps to a (codec, level) pair and takes precedence over any compression / compression_level fields.
| Profile | Codec | Level | Best for |
|---|---|---|---|
none | no compression | β | Debug, local scratch, fast iteration |
fast | snappy | β | Backfills, pilot runs, low-CPU environments |
balanced | zstd | 3 | Default for production β good ratio, moderate CPU |
compact | zstd | 9 | Storage- or network-cost-sensitive pipelines |
exports:
- name: events
format: parquet
compression_profile: balanced # zstd level 3
destination: { type: local, path: ./out }
If you need a specific codec that is not covered by the presets, use compression + compression_level directly and omit compression_profile.
Verification depth
verify controls how thoroughly --validate (and rivet validate) checks each
part at the destination:
size(default) β confirm each part exists at its recordedsize_bytes, plus manifest self-consistency and_SUCCESS. Content is also MD5-checked for free whenever the store surfaces a checksum in its listing, but a part without one is accepted as size-only.contentβ require every partβs content MD5 to match the storeβs listing checksum (no download). Any part that could only be size-verified fails validation with an actionable message.
How content verification works: Rivet computes each partβs MD5 before upload and
records it in the manifest; GCS and Azure compute their own for a part uploaded
as a single PUT and return it in object listings. --validate compares the two
with no download. Parts large enough to stream as multipart / block-list get
no checksum, and neither does S3 (its ETag is not an MD5 under SSE-KMS / SSE-C, so
rivet does not trust it) or local FS. Under verify: content on GCS / Azure, set
destination.oneshot_budget_mb (default 64 MB) comfortably above your part size:
the budget is shared by concurrent uploads, so a part one-shots only if it fits
what is free at that moment. S3 cannot meet verify: content.
The run report and rivet validate show coverage explicitly, e.g.
3 verified (2 md5, 1 size-only).
Parquet row group tuning
Parquet row groups affect memory usage during write, compression ratio, and downstream query performance (predicate pushdown, column skipping). When parquet: is omitted, Rivet uses the library default of 1,048,576 rows per group, which is optimal for narrow tables but can be large for wide tables.
exports:
- name: events
format: parquet
parquet:
row_group_strategy: auto # auto | fixed_rows | fixed_memory
target_row_group_mb: 128 # target Arrow buffer size per group (auto + fixed_memory)
max_row_group_mb: 256 # optional upper bound (all strategies)
| Field | Type | Default | Description |
|---|---|---|---|
row_group_strategy | auto | fixed_rows | fixed_memory | auto | How to determine row group size |
row_group_rows | integer | β | Exact rows per group; used with fixed_rows only |
target_row_group_mb | integer | 128 | Target Arrow buffer per group in MB; used with auto and fixed_memory |
max_row_group_mb | integer | β | Hard upper bound on group memory in MB (all strategies) |
| Strategy | Behavior |
|---|---|
auto | Estimates row width from schema column types, computes rows-per-group to hit target_row_group_mb. Narrow tables get large groups; wide tables get smaller groups. |
fixed_rows | Use row_group_rows exactly. Simple and deterministic, but does not adapt to row width. |
fixed_memory | Same math as auto (target / estimated row bytes), but the strategy name is explicit in logs. |
Examples:
# Auto-tune for a wide JSON table β groups sized to ~64 MB
parquet:
row_group_strategy: auto
target_row_group_mb: 64
max_row_group_mb: 128
# Fixed row count β useful when downstream tooling requires exact group sizes
parquet:
row_group_strategy: fixed_rows
row_group_rows: 500000
Note:
rivet planshows the selected strategy and target in the Format section whenparquet:is configured.
rivet initauto-generates this block for chunked exports and large full-mode tables, pre-selectingtarget_row_group_mb: 64for wide schemas (β₯ 5 text/JSON/bytea columns) and128for narrow ones.
exports[].on_schema_drift β schema drift policy
Controls what Rivet does when it detects a structural change in the output schema (column added, removed, or retyped) compared to the snapshot stored from the previous run.
| Value | Behavior |
|---|---|
warn | (default) Log a warning, store the new schema fingerprint, and continue the run. |
continue | Silently accept β store the new schema, no log output. |
fail | Abort the run with exit code 4 (the schema-drift exit class). The schema store is not updated, so the next run will detect the same change again. |
fail is useful in CI pipelines where schema changes must be reviewed before the new shape is exported downstream.
exports:
- name: orders
on_schema_drift: fail
When fail triggers, behavior depends on the runner: in single, keyset, and parallel-Mongo modes the schema check runs post-extraction, so the output file has already been written to the destination (but no cursor advance or manifest commit occurs). In chunked mode the check runs pre-chunk from a scan-free type probe, so the run aborts before any chunk is written. Re-run after confirming the schema change is intentional, or switch to warn to accept it.
exports[].columns β per-column type overrides
Override the Arrow type Rivet infers for a specific column. Useful when:
- a
NUMERIC/DECIMALcolumn has no explicit precision/scale in the source schema (beyondrivet initβs defaultdecimal(38,18)placeholder), or - you need a narrower precision for BigQuery NUMERIC compatibility.
columns:
<column_name>: <type> # applies to every captured table with the column
"<table>.<column>": <type> # applies to ONE table, wins over the bare key
Supported override types: decimal(p,s) / numeric(p,s) (both precision and
scale required; precision > 38 produces Decimal256), integer widths
(smallint/int/bigint/int16/int32/int64), floats
(real/float4/double/float8), bool, text/string/varchar,
uuid, json/jsonb, date, the timestamp* family (naive / tz /
_ns variants), and binary (bytea/binary/varbinary/blob).
Overrides apply to batch and CDC identically (the same resolution
surface). Key shapes: a bare column name applies to every captured table that
has the column β on a multi-table CDC export that means ALL of them; a
qualified "table.column" key targets one table and wins over the bare
key there. A qualified key naming a table the export does not capture β or
used on a query-shaped export β is a config error at load.
Example:
exports:
- name: orders
query: "SELECT id, amount, fee FROM orders"
format: parquet
destination:
type: local
path: ./out
columns:
amount: decimal(18,2)
fee: decimal(18,6)
rivet init generates these automatically. When introspecting a table, rivet init reads numeric_precision and numeric_scale from information_schema.columns. If both are present, it emits a concrete override (decimal(p,s)). If the column is unbounded (NUMERIC without explicit precision), rivet init emits a working default decimal(38,18) plus a # REVIEW: YAML comment β the config header adds a # NOTE: line, and rivet init -o β¦ prints a stderr reminder so you tighten precision when you know the real domain rules:
columns:
price: decimal(38,18) # REVIEW: DDL has no numeric(p,s); edit to the real decimal(p,s) β¦
Type overrides are applied at export time and are reflected in rivet check --type-report output.
Some PostgreSQL types have no Arrow representation and cannot be exported directly. Rivet will report an error listing all unmappable columns before the run starts.
| PostgreSQL type | Reason | Workaround |
|---|---|---|
geometry (PostGIS) | No Arrow equivalent | Cast to text: ST_AsText(col) AS col in your query |
geography (PostGIS) | No Arrow equivalent | Cast to text: ST_AsText(col) AS col |
hstore | No Arrow equivalent | Cast to JSON text: hstore_to_json(col)::text AS col |
tsvector, tsquery | No Arrow equivalent | Cast to text: col::text AS col |
point, line, polygon, etc. | No Arrow equivalent | Cast to text: col::text AS col |
Use a SQL expression in your query field to work around any unsupported type:
exports:
- name: locations
query: >
SELECT id, name, ST_AsText(geom) AS geom_wkt
FROM locations
format: parquet
destination:
type: local
path: ./out
Rivet exports the WKT text as a Utf8 (string) column. Downstream tools (DuckDB, GeoPandas, QGIS) can reconstruct geometry from WKT.
Mode-specific fields
Incremental (mode: incremental):
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
cursor_column | string | yes | β | Primary progression column. Must be strictly per-row-distinct and monotonically increasing β resume uses WHERE cursor > last_value, so rows that tie on the high-watermark value and become visible after it is passed are skipped. A low-resolution updated_at (second granularity) can tie; prefer a sequence/identity id or a sub-value-unique timestamp. See semantics.md β Known non-guarantees. |
cursor_fallback_column | string | when coalesce | β | Fallback column used when primary is NULL. Only valid with incremental_cursor_mode: coalesce |
incremental_cursor_mode | single_column | coalesce | no | single_column | coalesce progresses on COALESCE(primary, fallback). See modes/incremental-coalesce.md and ADR-0007. |
settle | object | no | β | Hold rows back until they stop changing: { after: 1h } ages the cursor itself, { after: 1h, column: server_time } ages another date/timestamp column. A row exports only once it is older than after (s/m/h/d) by the source clock. See modes/incremental.md Β§ Settle window. |
Chunked (mode: chunked):
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
chunk_column | string | yes* | β | Numeric or date/timestamp column to partition by. *Required unless chunk_by_key is set (mutually exclusive). |
chunk_by_key | string | yes* | β | Single index-backed UNIQUE NOT NULL column for keyset (seek) pagination β the source-safe shape for tables with no single-integer PK (UUID / string / composite). Requires the table: shortcut; mutually exclusive with chunk_column. See chunked modes and ADR-0020. |
chunk_size | integer | no | 100000 | Rows per chunk (numeric mode), or page size for keyset. Ignored when chunk_count is set. |
chunk_size_memory_mb | integer | no | β | Target memory budget per chunk in MB; chunk_size is derived from a per-engine row-size estimate, clamped to [10000, 5000000] rows. Works on PostgreSQL, MySQL and SQL Server (PG: pg_relation_size / reltuples; MySQL: information_schema AVG_ROW_LENGTH with InnoDB overflow correction; SQL Server: no estimate, falls back to 512 B/row with a warning). Requires the table: shortcut, mutually exclusive with an explicit non-default chunk_size:. |
chunk_count | integer | no | β | Divide the column range into exactly this many equal chunks. chunk_size is computed dynamically from min/max. Must be β₯ 1. Mutually exclusive with chunk_by_days. |
chunk_by_days | integer | no | β | Enable date chunking: window size in days. Mutually exclusive with chunk_count. |
parallel | integer | no | 1 | Concurrent chunk workers |
chunk_dense | boolean | no | false | Removed. true is refused at config load (it skipped or duplicated rows under concurrent writes); use chunk_by_key or chunk_column. |
chunk_checkpoint | boolean | no | false | Persist per-chunk progress for resume |
chunk_max_attempts | integer | no | β | Max retry attempts per chunk |
Time-window (mode: time_window):
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
time_column | string | yes | β | Timestamp column to filter on |
time_column_type | timestamp | unix | no | timestamp | Column type |
days_window | integer | yes | β | Rolling window size in days |
exports[] β value-based partitioning
Splits a full, chunked or incremental exportβs rows into Hive-style
col=value/ destination sub-folders by a date column. See partitioning.md.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
partition_by | string | no | β | Date/timestamp column to bucket rows by. Requires a {partition} token in destination.path/prefix. NULLs β col=__HIVE_DEFAULT_PARTITION__/. Not compatible with mode: time_window, mode: cdc, chunk_by_key, a load: block (per-export or top-level), or a MongoDB source. |
partition_granularity | day | month | year | no | day | Bucket width. |
exports[].meta_columns
| Field | Type | Default | Description |
|---|---|---|---|
exported_at | boolean | false | Add _rivet_exported_at column (Timestamp UTC; one value captured at sink construction and shared by every batch/row that sink writes β effectively one value per export run in single mode, per chunk (or keyset page) in multi-part modes) |
row_hash | boolean or list of column names | false | Add _rivet_row_hash column β lower 64 bits of xxHash3-128, written as Int64 for fast PARTITION BY / JOIN. true hashes every column; a list (row_hash: [id, status, updated_at]) hashes exactly those columns in that order and records the covered set in the run manifest. Deterministic across runs; distinguishes NULL from empty string. |
exports[].quality
| Field | Type | Description |
|---|---|---|
row_count_min | integer | Fail if fewer rows exported |
row_count_max | integer | Fail if more rows exported |
null_ratio_max | map (column β float) | Fail if null ratio exceeds threshold |
unique_columns | list of strings | Fail if values are not unique |
unique_max_entries | integer | Cap on distinct values tracked per column during uniqueness checks. When reached, a Warn is emitted and checking stops for that column; duplicates already found before the cap still fail the run. Prevents unbounded memory growth on high-cardinality columns (UUIDs, email addresses, event IDs). |
Uniqueness tracking uses typed xxHash3-64 internally β numeric and binary columns are hashed directly from raw bytes without string formatting. unique_max_entries is the primary knob to control memory on very large tables.
Example:
quality:
row_count_min: 100
null_ratio_max:
email: 0.05 # email must be <5% null
unique_columns:
- id
- email
unique_max_entries: 1000000 # stop after 1M unique values; warn if limit hit
Without unique_max_entries β tracking is unbounded. Safe for tables with hundreds of thousands of rows; may use significant RAM on tables with tens or hundreds of millions of distinct values.
With unique_max_entries β tracking stops at the limit and the run summary shows a warning. Duplicates found before the limit still fail the run; ones past it go unseen. Use when you want a best-effort uniqueness check without memory risk.
exports[].destination
The complete per-backend field list (local / s3 / gcs / azure / stdout)
is in the generated config-reference.md (section
exports[].destination). Per-backend setup, auth flows, and permissions:
destinations/ β local Β·
s3 Β· gcs Β·
azure Β· stdout, plus the
cloud auth matrix.
Path and prefix placeholders
The path (local) and prefix (S3 / GCS) fields support template placeholders, substituted at plan-build time:
| Placeholder | Value |
|---|---|
{date} | UTC date as YYYY-MM-DD |
{export} | Export name from config |
{table} | Alias for {export} |
{run_id} | The runβs unique id β substituted only by rivet validate --run-id, which re-targets validation at that runβs prefix. run and apply resolve destinations without a run id, so the token is always left verbatim there and the destination open fails fast rather than aliasing to an unintended prefix (and rivet load refuses a {run_id} prefix outright β it cannot know which runβs output to load). |
destination:
type: s3
bucket: my-data
prefix: exports/{date}/{export}/
region: us-east-1
With an export named orders running on 2026-05-14, this resolves to exports/2026-05-14/orders/.
notifications
| Field | Type | Description |
|---|---|---|
slack | object | Slack notification config |
notifications.slack
| Field | Type | Description |
|---|---|---|
webhook_url | string | Slack incoming webhook URL |
webhook_url_env | string | Env var containing webhook URL |
on | list | Events to notify on: failure, schema_change, degraded |
Example:
notifications:
slack:
webhook_url_env: SLACK_WEBHOOK
on: [failure, schema_change]
Environment variable interpolation
Any string value can reference environment variables:
source:
url: "postgresql://${DB_USER}:${DB_PASS}@${DB_HOST}:5432/mydb"
Query parameters
Queries can use ${key} placeholders filled by --param key=value:
exports:
- name: filtered
query: "SELECT * FROM orders WHERE region = '${region}'"
rivet run --config export.yaml --param region=us-east