Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Config reference (generated — rivet-cli 0.30.0)

Rendered from the JSON Schema rivet schema config emits (schemars ← the Rust Config types). It cannot drift from the code. Hand-written guidance lives in the surrounding config guide; this table set is generated — edit the Rust structs, not this block.

Top level (rivet.yaml)

FieldTypeRequiredDescription
sourceSourceConfigyes
exportsarray of ExportConfigyes
notificationsNotificationsConfig
parallel_exportsbooleanSame as rivet run --parallel-exports: the exports run concurrently, at most 16 at once; a CDC export run alone also takes its pending baseline snapshots at most 16 at once.
parallel_export_processesboolean
loadLoadSectionThe warehouse load target — consumed by rivet load, so ONE config drives both the export and the downstream load. The extraction commands validate it (a malformed block fails rivet check before an extract runs) and otherwise ignore it: it shapes the load, not the extract.

source

FieldTypeRequiredDescription
typepostgres | mysql | mssql | oracle | mongoyes
urlstring
url_envstring
url_filestring
hoststring
portinteger
userstring
passwordstring
password_envstring
databasestring
environmentlocal | replica | productionOperational profile of the source database. Selects the default tuning profile when none is explicitly set in source.tuning.profile or export.tuning.profile:
tuningTuningConfig
tlsTlsConfigTransport security settings (ADR: SecOps). When absent, Rivet connects without TLS — a warning is emitted so operators are aware. See [TlsConfig].
mongoMongoConfigMongoDB-specific read options (source.mongo:). Honoured only when type: mongo; ignored by the SQL engines. See [MongoConfig].

source.mongo (MongoDB read options)

FieldTypeRequiredDescription
jsonrelaxed | canonicalJSON rendering of the document column. relaxed (default) keeps common scalars native (42, "x"); canonical wraps every number ({"$numberLong":"…"}) so Int64/Double round-trip losslessly through a JSON-number parser that would otherwise clamp values beyond 2^53.
read_concernserver | snapshotRead concern for the collection scan. snapshot gives a point-in-time consistent full export (no doc missed/double-read under concurrent writes) — requires MongoDB 5.0+ on a replica set; a standalone rejects it. Default (server) uses the server’s default read concern.
no_cursor_timeoutbooleanKeep the scan cursor alive past the server’s idle timeout (default 10 min) so a slow destination cannot let the server reap the cursor mid-scan and silently drop the tail of a large collection. Default: true.
page_sizeintegerWhen set, read the collection with keyset (seek) pagination on _id instead of one long-held cursor: each page is a bounded find({_id: {$gt: last}}).sort({_id: 1}).limit(page_size) — an indexed range scan that becomes one output part file. Bounds longest-query time (no 35-minute cursor to hit a timeout / snapshot window) and is the base for parallel _id-range reads. Works with any uniform _id type (ObjectId — the default — integer, string, date, …); a collection mixing _id type brackets errors with a clear message pointing at the full ordered scan (Mongo’s $gt compares only within a type bracket, so a mixed key would silently drop every bracket but one). Unset ⇒ the single-cursor full scan.
resumebooleanWith keyset paging (page_size), persist the last committed _id and resume from it next run — a crashed export continues where it left off, and a re-run captures only documents inserted since (ObjectId _id is time-ordered). Default false re-reads the whole collection each run (plain mode: full semantics). No effect without page_size.

source.tls

FieldTypeRequiredDescription
modedisable | require | verify-ca | verify-fullEnforcement level. See [TlsMode].
ca_filestringPEM-encoded CA certificate to trust for server verification. Required for [TlsMode::VerifyCa] and [TlsMode::VerifyFull] against a private CA.
accept_invalid_certsbooleanAccept certificates not chained to a trusted CA. Dangerous — disables server authentication — and only honored when explicitly true.
accept_invalid_hostnamesbooleanAccept certificates whose subjectAltName does not match the connection hostname. Dangerous — disables hostname verification.

exports[]

FieldTypeRequiredDescription
namestringyes
querystring
query_filestring
tablestringShortcut for query: "SELECT * FROM <schema>.<table>". Accepts table or schema.table with ASCII-only identifiers ([A-Za-z_][A-Za-z0-9_]*). Generates an unquoted single-table query so the Postgres NUMERIC catalog-hint resolver recognises it and auto-types numeric(p,s) columns without manual overrides. Mutually exclusive with query and query_file.
tablesarray of stringCDC only: capture several tables through ONE change stream (one PostgreSQL slot / one MySQL binlog connection) instead of one export — and one slot — per table. Each table’s parts land under <destination>/<table>/ with their own manifest.json + _SUCCESS; the checkpoint (stream position) is shared. Mutually exclusive with table:. Not yet supported for SQL Server (capture instances are per-table).
modefull | incremental | chunked | time_window | cdc
cdcCdcExportConfigChange-data-capture settings, required when mode: cdc. Reuses the export’s table, destination, and format; carries only the CDC-specific knobs (resume checkpoint, per-engine stream params).
cursor_columnstring
cursor_fallback_columnstringSecondary column for [IncrementalCursorMode::Coalesce] only (see ADR-0007).
incremental_cursor_modesingle_column | coalesceHow primary (and optional fallback) columns drive incremental progression.
settleSettleConfigIncremental only: export a row once it is older than settle.after (source clock).
chunk_columnstring
chunk_densebooleanRemoved. Kept only so a config that still sets chunk_dense: true is refused at load.
chunk_sizeinteger
chunk_size_memory_mbintegerTarget memory budget per chunk in MB. When set, chunk_size is derived from this budget at plan-build time using the engine’s row-size estimate (PostgreSQL pg_relation_size / reltuples; MySQL information_schema average row length; a defensive 512 B/row default when no estimate exists, e.g. SQL Server), clamped to [10_000, 5_000_000] rows. Mutually exclusive with an explicit non-default chunk_size:. Requires mode: chunked and the table: shortcut (the row-size probe needs a known relation); any SQL engine works. yaml exports: - name: page_views table: public.page_views mode: chunked chunk_size_memory_mb: 256
chunk_countintegerDivide the column range into exactly this many equal chunks. Mutually exclusive with chunk_by_days. When set, chunk_size is computed dynamically from min/max.
chunk_by_daysinteger
chunk_by_keystringKeyset (seek) pagination on this single index-backed unique key — the source-safe shape for tables without a single-integer PK (OPT-4). The column MUST be backed by a usable index (PK or unique); the planner refuses a non-indexed key rather than emit a full-scan + filesort query.
parallelintegerConcurrent chunk/page workers (default 1). On a RANGE chunk (chunk_column) or KEYSET (chunk_by_key) export, parallel: N fans the table into N ROW-percentile ranges that seek concurrently over separate connections — the half-open intervals partition the key, so the union reads every row exactly once (structural parity, all engines). Extraction is I/O-bound, so the win plateaus early (~3x at N=4, little beyond). SWEET SPOT: indexed tables up to ~10M rows at parallel: 4. rivet init scaffolds a row-scaled value (<=500K -> 1, <5M -> 2, >=5M -> 4); a preflight warns past ~5M rows (peak RSS ~= N x chunk_size). Beyond ~10M the KEYSET boundary sampler (an index OFFSET skip) grows costly at setup — prefer a range chunk_column there.
waveintegerAdvisory execution wave (1 = highest priority, run first). Written by rivet plan from the source-aware prioritization score (see ADR-0006) and consumed by rivet apply, which runs exports wave-by-wave in ascending order. None = unscheduled (apply treats it as the last wave). Operators may hand-edit it; a later rivet plan refreshes it in place.
parallel_safebooleanWhether this export is cheap enough to run concurrently with its wave-mates under rivet apply --parallel-export-processes. Written by rivet plan (true when the source-aware cost class is Low, i.e. < ~100K rows); a heavier table already chunk-parallelizes internally, so two of them at once would overload the source. None/false → the export runs alone within its wave. Operators may hand-edit it; a later rivet plan refreshes it in place.
time_columnstring
time_column_typetimestamp | unix
days_windowinteger
partition_bystringDate/time output partitioning: split this export’s rows into one destination sub-prefix per calendar bucket of this DATE or TIMESTAMP column, bucketed by partition_granularity (day / month / year), in a Hive-style col=value/ layout (created_at=2023-01-01/, created_at=2023-01/, created_at=2023/). Requires a {partition} token in destination.path / destination.prefix. This is not arbitrary value partitioning: the column’s min/max is read and parsed as a date to generate contiguous calendar buckets, so a non-temporal column (e.g. partition_by: status) fails at run time with “could not parse partition min <value> from column <col> as a date”. To split by a categorical column, write one export per value with a WHERE filter instead. Applies to full, chunked and incremental exports on a SQL source: each partition runs the export’s own mode, so mode: chunked chunks within a day. Rows whose partition column is NULL land in col=__HIVE_DEFAULT_PARTITION__/ (Hive default partition) so no row is silently dropped. Not compatible with mode: time_window, mode: cdc, chunk_by_key, a load: block (per-export or top-level), or a MongoDB source — each is refused when the config loads. yaml exports: - name: events table: events partition_by: created_at # must be a DATE or TIMESTAMP column partition_granularity: day destination: type: s3 bucket: my-bucket prefix: "events/{partition}/" # → events/created_at=2023-01-01/
partition_granularityday | month | yearCalendar bucket width for partition_by: day (default), month, or year. Determines how the partition column’s date/timestamp range is split into contiguous Hive buckets (col=2023-01-01/ / col=2023-01/ / col=2023/). Has no effect unless partition_by is set.
formatparquet | csvyes
compressionzstd | snappy | gzip | lz4 | none
compression_levelinteger
compression_profilenone | fast | balanced | compact
skip_emptybooleanRecord a batch run that delivers 0 rows as skipped (with a reason) instead of success, on every batch runner. No file is written for 0 rows either way, and a skipped run leaves the prefix describing the last run that delivered, so a full load keeps the previous data. mode: cdc does not read it.
destinationDestinationConfigyes
verifysize | contentIntegrity depth required of --validate for this export’s parts. size (default) accepts size-only verification; content requires every part’s content MD5 to be checked against the store’s listing (no download) and fails validation for any part that could only be size-verified — a part too large for a single PUT (on GCS / Azure, raise destination.oneshot_budget_mb above the part size), or a backend that exposes no trusted checksum (S3, local FS).
meta_columnsMetaColumns
qualityQualityConfig
max_file_sizestringRotate to a new part when the current file reaches this size. Accepts B/KB/MB/GB (case-insensitive) or a bare byte count; a fractional value is allowed (1.5GB). Units are binary (IEC-style): KB = 1024 bytes, MB = 1024 KB, GB = 1024 MB. Example: 256MB. Parquet row groups are capped at a quarter of it, so a part stays within about one row group of the size whatever parquet.row_group_strategy says.
chunk_checkpointbooleanPersist per-chunk / per-page progress so a crashed run resumes from the last durably committed point instead of re-reading from the start. This is pure crash-recovery: a clean re-run (the prior run finished) still does a full pass — it never silently skips already-exported rows. Safe to enable on any table; rivet init defaults it on for chunked and keyset exports.
keyset_incrementalbooleanKeyset only (chunk_by_key): on a clean re-run, continue from the last exported key — pull ONLY rows with a key past the high-water mark. This is incremental-by-key, correct ONLY for APPEND-ONLY tables (a mutable row whose key already passed is silently never re-read). Opt-in and off by default; crash-recovery does not need it (that is chunk_checkpoint). For a mutable table use mode: incremental on a timestamp cursor instead.
chunk_max_attemptsinteger
tuningTuningConfig
source_groupstringOptional logical group for shared source capacity (replica, host). Advisory prioritization only.
reconcile_requiredbooleanHint (Epic C / ADR-0006) that this export should always be treated as reconcile-heavy by planning, independent of the --reconcile CLI flag. Advisory only.
columnsobjectPer-column type overrides (roadmap §8). Keys are column names; values are short type strings such as decimal(18,2), timestamp_tz, json. yaml exports: - name: payments columns: amount: decimal(18,2) fee: decimal(18,6) created_at: timestamp_tz Overrides take priority over autodetection and are validated at plan time — an invalid type string fails before the export runs.
targetstringDownstream warehouse this export targets (bigquery / bq, duckdb). When set, rivet check --type-report resolves each column against it (native type, honest autoload type, recovery hint) without needing --target on the CLI — the CLI flag still wins when both are present. The Parquet interchange stays target-neutral (ADR-0014 T2); target: only drives guidance and the future load-schema artifact. yaml exports: - name: payments target: bigquery
loadLoadOverridePer-export overrides for the top-level load: block (pk, cleanup_source, gc_orphans, cluster_by, partition, allow_source_drift); any field omitted here inherits the top-level value. The warehouse target is shared and stays in the top-level load: — it cannot be overridden per export. yaml load: { target: bigquery, project: p, dataset: d } # shared default exports: - name: orders table: orders mode: cdc load: pk: [id] # this table's pk partition: { column: created_at, granularity: day, expiration_days: 400 }
on_schema_driftwarn | continue | failPolicy applied when structural schema drift is detected (column added, removed, or retyped). Defaults to warn: log a warning and continue.
shape_drift_warn_factornumberGrowth-factor threshold for data shape drift warnings (Epic 8). When a string/binary column’s max observed byte length in the current run exceeds stored_max * shape_drift_warn_factor, Rivet logs a warning. None uses the default of 2.0. Set to 0.0 to disable shape tracking. Applies to every batch mode — multi-part runs compare the largest value any chunk, page or worker saw. mode: cdc does not check it.
parquetParquetConfigParquet row group tuning. Only meaningful when format: parquet. When absent, the parquet library default (1,048,576 rows/group) is used.

exports[].cdc (mode: cdc)

FieldTypeRequiredDescription
initialsnapshotFirst-run behaviour: snapshot = anchor → full snapshot → drain (see [CdcInitialMode]). Omitted ⇒ capture changes only, with no anchor step (the default; the operator owns the initial load). snapshot anchors, and on engines with no server-side anchor (MySQL, SQL Server) that makes checkpoint: mandatory — the checkpoint file IS the anchor there. This doc line is what the generated config reference renders, so every accepted value must be explained HERE: the reference lists the variants from the enum but describes only this sentence, so an explanation left on a variant alone documents a value the reader is told exists and never told the meaning of.
checkpointstringPersist/resume the source log position to this file. Omit to tail from the current position without checkpointing.
until_currentbooleanCatch up to the source’s current end and exit (a bounded run), instead of streaming indefinitely — ideal for a scheduler. For MySQL this is a non-blocking binlog dump; PostgreSQL / SQL Server already drain-and-exit. Defaults to true (bounded): the OSS model is scheduler-driven, and omitting this must NOT silently start a never-terminating stream. Setting false opts into the continuous model, which is engine-specific: a true daemon on MySQL (blocking binlog dump) and MongoDB (the change stream blocks awaiting events; ends only if the stream is invalidated/closed); PostgreSQL / SQL Server still exit on catch-up — one unbounded pass, run it under a supervisor. Oracle refuses false: LogMiner is always a bounded drain to the SCN current at open.
max_eventsintegerStop at the first COMMIT BOUNDARY once N change events have been captured (default: until end of stream / interrupted). A soft cap, like rollover: a transaction is never split, so the run may overshoot N by the remainder of the transaction the cap landed in — a hard per-event stop cut transactions mid-flight and left the stream unable to advance past them.
rolloverintegerRows per output part file (default 100000). A part also rolls at a transaction boundary, so it never splits a transaction. Larger ⇒ fewer, bigger files but more drain memory — the PostgreSQL peek reads a part’s worth per batch, so drain RSS is O(rollover). Tune per workload: raise it to cut file count, lower it to cap memory on a small extractor.
rollover_memory_mbintegerRoll a part once its buffered changes reach this many MB, whichever comes first with rollover. Caps the in-memory buffer and the part file size by bytes instead of a fixed row count — predictable for tables with wide (large JSON / blob) rows, mirroring the batch path’s batch_size_memory_mb. Defaults to 256 (MiB): the row count alone is a budget for one row width, and absence must not mean “no byte budget”. The bytes are the buffered changes’ RESIDENT cost (struct + commit position + values), not the part file’s size on disk. It bounds the buffer, not the process: a roll encodes up to 16 tables’ parts at once, each a columnar copy of its buffered rows, so peak RSS sits above the budget (measured +31% on a 60-table stream).
server_idintegerMySQL replica server-id for the binlog connection (default 4271; must be distinct from the source’s and any other replica).
slotstringPostgreSQL logical replication slot name (default rivet_slot).
capture_instancestringSQL Server CDC capture instance, e.g. dbo_orders — required for sqlserver:// sources.
backfillautoWhich EXPORTS supply the baseline read (see [CdcBackfill]). Absent ⇒ no baseline: the stream captures changes only, and the operator owns the initial load.

exports[].tuning

FieldTypeRequiredDescription
profilefast | balanced | safe
batch_sizeinteger
batch_size_memory_mbintegerTarget memory per batch in MB. Mutually exclusive with batch_size.
throttle_msinteger
statement_timeout_sinteger
max_retriesinteger
retry_backoff_msinteger
lock_timeout_sinteger
memory_threshold_mbinteger
max_batch_memory_mbintegerHard cap on Arrow batch memory in MB. When a batch exceeds this limit, on_batch_memory_exceeded determines the response.
on_batch_memory_exceededwarn | fail | auto_shrinkPolicy applied when a batch exceeds max_batch_memory_mb. Default: warn.
adaptivebooleanEnable real-time batch size adaptation based on DB pressure metrics. The batch loop samples the export’s OWN extraction pressure: Postgres pg_stat_bgwriter checkpoint pressure; MySQL the read-spill pair Created_tmp_disk_tables and Innodb_buffer_pool_wait_free. SQL Server takes no batch sample — its batch size comes from the memory cap alone. It also arms the OPT-2 concurrency governor when parallel > 1. The governor samples a DIFFERENT, write-driven signal on its own monitoring connection — one a read-only export cannot inflate, so it can never shed its own workers over its own reads: Postgres checkpoints_req, MySQL Innodb_log_waits, SQL Server Log Flush Waits/sec (_Total).
min_parallelintegerFloor for the concurrency governor (lowest parallelism under pressure). Default 1. Ceiling is the export’s parallel.
max_value_mbintegerHard per-value size ceiling in MB. A single text/JSON/blob cell larger than this aborts the run with RIVET_VALUE_TOO_LARGE. 0 disables the guard. Default: 256.

exports[].destination

FieldTypeRequiredDescription
typelocal | s3 | gcs | azure | stdoutyes
bucketstring
prefixstring
pathstring
regionstring
endpointstring
credentials_filestring
access_key_envstring
secret_key_envstring
session_token_envstringName of an env var holding an AWS STS session token, for use with short-lived credentials issued by AWS IAM Identity Center / SSO, aws sts assume-role, MFA-protected sessions, EKS IAM Roles for Service Accounts, etc. Pair with access_key_env + secret_key_env. See docs/cloud-auth.md for the AWS auth-flow matrix.
aws_profilestring
account_namestringAzure storage account name (the prefix in <account>.blob.core.windows.net). Plain string — not a secret. Pair with account_key_env. See docs/cloud-auth.md for the Azure auth-flow matrix.
account_key_envstringName of an env var holding the Azure Storage account key. Treated as a credential and wiped from heap on drop — same SecOps treatment as access_key_env. Pair with account_name. Mutually exclusive with sas_token_env.
sas_token_envstringName of an env var holding an Azure Storage SAS token — typically a short-lived, scope-limited credential issued out-of-band (Azure portal / az storage container generate-sas / Azure SDK). Use this instead of account_key_env when the operator does not have the long-lived account key or wants per-job scoped access. Pair with account_name. Mutually exclusive with account_key_env. The token value is wiped from heap on drop via the same Zeroizing<String> wrapper as account_key_env. Leading ? is trimmed transparently so the operator can paste either the full ?sv=…&sig=… query string or the raw token body.
allow_anonymousboolean
oneshot_budget_mbintegerCap on the RAM one-shot (single-PUT) upload buffers may hold, in MB (default 64; cloud destinations only). A one-shot PUT buffers the whole part; on GCS and Azure the store then records a Content-MD5 that validate checks, and on every store it is one request instead of a sequential multipart (S3 verifies size-only either way). A part that does not fit the remaining budget streams instead (memory-bounded). 0 streams every non-empty part. Each distinct value is one pool per rivet process, shared by every destination configured with it (including each table of a CDC export). Different values are separate pools, so worst-case one-shot RAM is the sum of the distinct values in use; under parallel_export_processes every child has its own.

exports[].quality

FieldTypeRequiredDescription
row_count_mininteger
row_count_maxinteger
null_ratio_maxobject
unique_columnsarray of string
unique_max_entriesintegerCap on the number of distinct values tracked per column during uniqueness checks. When the limit is hit, a Warn issue is emitted and tracking stops for that column. Prevents unbounded HashSet growth on high-cardinality columns.

exports[].parquet

FieldTypeRequiredDescription
row_group_strategyauto | fixed_rows | fixed_memoryHow to determine the row group size. Default: auto.
row_group_rowsintegerExact number of rows per group (fixed_rows only).
target_row_group_mbintegerTarget Arrow buffer memory per row group in MB (auto and fixed_memory). Default: 128.
max_row_group_mbintegerHard upper bound on row group memory in MB. When set, further reduces computed row count.

load (the warehouse target, consumed by rivet load)

FieldTypeRequiredDescription
targetbigquery | snowflake | clickhouseyesThe warehouse: bigquery, snowflake or clickhouse.
projectstringBigQuery: the project the dataset lives in.
datasetstringBigQuery: the dataset the tables are created in.
connectionstringSnowflake: the snow CLI connection name.
warehousestringSnowflake: the virtual warehouse the load runs on.
databasestringSnowflake / ClickHouse: the database the tables are created in.
schemastringSnowflake: the schema the tables are created in.
storage_integrationstringSnowflake: a pre-created GCS STORAGE INTEGRATION.
urlstringClickHouse: the HTTP endpoint, e.g. http://localhost:8123.
userstringClickHouse: the user the load authenticates as.
password_envstringClickHouse: the env var holding that user’s password.
named_collectionstringClickHouse: a server-side named collection holding the bucket’s URL and HMAC keys; ClickHouse then reads the Parquet itself instead of rivet sending it.
cleanup_sourcebooleanAfter a successful load, delete the staged Parquet under the export prefix.
pkauto | noneDedup key of the incremental/CDC current-state view: auto (the source primary key rivet run recorded), none, or explicit columns; ignored for full.
layoutlog_view | base_bufferlog_view or base_buffer — where the current state lives. Absent derives it from the mode: a CDC stream with a backfill: is base+buffer, the rest changelog+view. base_buffer needs target: bigquery — rivet compact is what merges the buffer into the base, and it is BigQuery-only.
deleted_flagbooleanWhether the base carries a __is_deleted column. Absent derives it from the mode: a CDC stream expresses deletes and gets the flag, a query-based export cannot express one and does not — an extra column per row otherwise.
allow_source_driftbooleanLoad even when a run manifest’s source count disagrees with what it extracted (source→file drift): warn instead of blocking.
gc_orphansbooleanAfter a successful load, delete staged Parquet under the export prefix that no Success manifest references — crash leftovers. Only when no extract writes the prefix concurrently.
cluster_byauto | noneCLUSTER BY of the table the load writes: auto (the primary key), none, or explicit columns (at most 4 on BigQuery).
partitionnoneHow the table the load writes is partitioned: none (default), or exactly one of column (+ granularity), an integer range, or ingestion time.

exports[].load and exports[].load.tables.<table>

FieldTypeRequiredDescription
pkauto | noneDedup key of this table’s current-state view.
cleanup_sourceboolean
gc_orphansboolean
cluster_byauto | noneCLUSTER BY of this table.
allow_source_driftboolean
layoutlog_view | base_bufferWhere this table’s current state lives; inherits when absent.
deleted_flagbooleanWhether this table’s base carries __is_deleted; inherits when absent.
partitionnoneThis table’s partitioning; none clears an inherited one.
tablesobjectOn a multiplex tables: CDC export: the override for ONE captured table, keyed by its name, layered over this block — six tables through one stream rarely share a partition column or a key. Every name must be one of the export’s tables:; a nested tables: is refused.