Command-Line Help for rivet
This document contains the help content for the rivet command-line program.
Command Overview:
rivet↴rivet run↴rivet check↴rivet doctor↴rivet cdc↴rivet load↴rivet compact↴rivet state↴rivet state show↴rivet state reset↴rivet state files↴rivet state reset-chunks↴rivet state chunks↴rivet state progression↴rivet state runs↴rivet state finish-run↴rivet state loads↴rivet completions↴rivet init↴rivet plan↴rivet apply↴rivet repair↴rivet validate↴rivet reconcile↴rivet metrics↴rivet schema↴rivet schema config↴rivet schema cli↴rivet schema errors↴rivet journal↴
rivet
Export data from databases to files
Usage: rivet [OPTIONS] <COMMAND>
Getting started (the happy path):
- rivet init scaffold a config from your database
- rivet doctor test source + destination auth
- rivet check column-type & schema report
- rivet run export your data
Docs: https://github.com/panchenkoai/rivet/blob/main/docs/getting-started.md
Subcommands:
run— Run export jobs defined in configcheck— Column-type & schema report for each export (needs a working connection; rundoctorfirst if it can’t connect)doctor— Verify source + destination auth/connectivity (run this first)cdc— Stream change data capture (CDC) from a source’s transaction logload— Load an export’s Parquet into a warehouse (BigQuery / Snowflake)compact— Merge each base-and-buffer CDC table’s<table>__changesbuffer into its base table (MERGEby primary key: updates, inserts, deletes flagged as__is_deleted) and drop the buffer — the billed step of the cyclerun → load → compact, labelledrivet_op:mergeper tablestate— Manage export statecompletions— Generate shell completionsinit— Generate a config scaffold from a live database (connect + introspect)plan— Generate an execution plan artifact (no data exported)apply— Execute a sealed plan artifact, or run a config’s exports wave-by-waverepair— Targeted repair of chunks flagged by reconcile: emit a repair plan, or re-export only mismatched rangesvalidate— Re-run manifest-aware verification against an existing destination, no extractionreconcile— Partition/window reconciliation: re-count per-partition on source and report mismatches. Requires a chunked export previously run withchunk_checkpoint: true. Exits non-zero when a mismatch is detected, so CI / orchestrators can gate on it (anunknownpartition warns but does not fail)metrics— Show export metrics historyschema— Emit machine-readable schemas for Rivet’s data contractsjournal— Inspect structured run journal (events, files, retries, quality issues)
Options:
--json-errors— Output errors as {“error”:“…”} JSON to stderr; useful for machine-readable orchestration
rivet run
Run export jobs defined in config
Usage: rivet run [OPTIONS] --config <CONFIG>
Options:
-
-c,--config <CONFIG>— Path to YAML config file -
-e,--export <EXPORT>— Run only a specific export by name -
--validate— Validate output files after writing -
--reconcile— Row-count audit: run COUNT(*) on the source and compare with the exported row count; a mismatch fails the run. Implies--validate(also verifies the output file manifest) -
--resume— Resume a chunked export withchunk_checkpoint: true(same query/chunk_column/chunk_size) -
--force— Override safety gates that would otherwise refuse the run.Today: with
--resume, allows starting against a destination prefix whose_SUCCESSmarker is already present. Without--force, resume against an already-complete run refuses, so an operator cannot accidentally re-export over a verified dataset. -
--parallel-exports— Run the config’s exports concurrently, at most 16 at once (needs 2+ exports); a CDC export run alone also takes its pending baseline snapshots at most 16 at once -
--parallel-export-processes— Run each export as a separaterivetchild process (parallel; true per-export peak RSS; more overhead than threads) -
--summary-output <PATH>— Write the run aggregate summary as JSON to this file (in addition to .rivet_state.db) -
--json— Print the run aggregate summary as JSON to stdout at the end of the run -
-p,--param <KEY=VALUE>— Query parameter: key=value (repeatable, substitutes ${key} in queries)
rivet check
Column-type & schema report for each export (needs a working connection; run doctor first if it can’t connect)
Usage: rivet check [OPTIONS] --config <CONFIG>
Options:
-c,--config <CONFIG>— Path to YAML config file-e,--export <EXPORT>— Check only a specific export by name-p,--param <KEY=VALUE>— Query parameter: key=value (repeatable, substitutes ${key} in queries)--type-report— Show per-column type fidelity report (source type → Rivet type → Arrow type)--strict— Fail with non-zero exit code if any column has an unsafe type mapping--json— Output type report as JSON (implies –type-report)--target <TARGET>— Check compatibility against a target warehouse (e.g. bigquery)
rivet doctor
Verify source + destination auth/connectivity (run this first)
Usage: rivet doctor [OPTIONS] --config <CONFIG>
Options:
-c,--config <CONFIG>— Path to YAML config file--json— Emit the probe results as a JSON object ({config_path, all_ok, checks: [{name, ok, detail?, hint?}]}) instead of the text report
rivet cdc
Stream change data capture (CDC) from a source’s transaction log.
The engine is chosen from the URL scheme: mysql:// (binlog), postgresql:// (logical slot), sqlserver:// (change tables), or mongodb:// (change stream). Emits one JSON object per row change to stdout (NDJSON) and, with --checkpoint, persists a resume position; --output writes typed Parquet/CSV instead. Per-engine prerequisites (ROW binlog + REPLICATION grant, wal_level=logical, enabled CDC, a replica set) are in docs/reference/cdc.md. The fuller, config-driven path is rivet run with mode: cdc.
Usage: rivet cdc [OPTIONS] <--source <SOURCE>|--source-env <ENV_VAR>|--source-file <PATH>>
Options:
-
--source <SOURCE>— Database URL —postgresql://,mysql://,sqlserver://, ormongodb://(engine chosen from the scheme). Visible inps; prefer--source-env/--source-fileoutside local dev -
--source-env <ENV_VAR>— Name of an environment variable holding the database URL -
--source-file <PATH>— Path to a file containing just the database URL (one line) -
--server-id <SERVER_ID>— Replica server-id for the binlog connection (must be distinct from the source’s and any other replica)Default value:
4271 -
--checkpoint <PATH>— Persist/resume the engine’s log position to this file (MySQL binlog coordinates / PostgreSQL slot-resume marker / SQL Server from-LSN / MongoDB resume token). If omitted, each engine falls back to its own anchor: MySQL and MongoDB start at the source’s CURRENT position (nothing written before now is captured), PostgreSQL resumes from the slot itself (server-side — a slot created here pins at the current WAL position), and SQL Server starts at the capture instance’sfn_cdc_get_min_lsn(it over-reads the retained backlog rather than skipping) -
--table <TABLE>— Only emit changes for this table (repeatable; default: all tables) -
--max-events <N>— Stop at the first COMMIT BOUNDARY once N change events have been emitted — a soft cap, so the run may overshoot N by the remainder of the transaction the cap lands in. A hard per-event stop cannot checkpoint inside a transaction, so a transaction longer than N left the run re-reading the same position on every restart. Without it the default bounded run drains to the log end as of open and exits; streaming until interrupted needs--stream -
--output <DIR>— Write typed Parquet/CSV files to this directory (the upsert/after-image shape) instead of NDJSON to stdout. Requires exactly one--table— its schema is resolved from the source -
--format <FORMAT>— Output file format when--outputis set:parquet(default) orcsvDefault value:
parquet -
--rollover <N>— Rows per output file (rollover) when--outputis set. Larger ⇒ fewer, bigger files but more drain memory (the PostgreSQL peek reads a part’s worth per batch: memory is O(rollover)). Turn it up/down per workloadDefault value:
100000 -
--slot <NAME>— PostgreSQL logical slot name (CDC; created if absent)Default value:
rivet_slot -
--capture-instance <INSTANCE>— SQL Server CDC capture instance, e.g.dbo_orders— required forsqlserver://sources -
--stream— Stream continuously instead of the DEFAULT bounded “read to the log end and exit” drain. What “continuously” means is per engine: MySQL (a blocking binlog dump) and MongoDB (a change stream that blocks awaiting events) stay up until stopped; PostgreSQL and SQL Server are poll adapters that STILL EXIT ON CATCH-UP — there this is one unbounded pass, not a daemon, so run it under a supervisor that restarts it. Omit it for the scheduler-friendly bounded run (the default). For MySQL the bounded run is a non-blocking binlog dump; PostgreSQL / SQL Server drain their backlog and exit
rivet load
Load an export’s Parquet into a warehouse (BigQuery / Snowflake)
The native column schema, target table, partition, and source URIs are all derived from the config’s top-level load: block — nothing is hand-typed. A multi-table config loads its exports into the shared target on a POOL of up to 16 worker threads, capped at the number of tables; --pool 1 is the strictly sequential pass. Column types come from the state DB, recorded by each export’s last successful rivet run; the load never connects to the source.
Usage: rivet load [OPTIONS] --config <CONFIG>
Options:
-c,--config <CONFIG>— Path to YAML config file — extraction PLUS a top-levelload:block. ONE file drives both the export and the load: the mode (full/incremental/cdc),pk:,cleanup_source:,gc_orphans:andallow_source_drift:all live in the config, not on the CLI--run-id <RUN_ID>— Correlation id stamped on every warehouse job/query of this load run (BigQueryrivet_runlabel / SnowflakeQUERY_TAG), so cost slices per run as well as per table. Defaults to a generated id--rebuild-changelog— Rebuild a<table>__changeswhose partitioning differs from the config’sload.partition— a billed query copying every row — and swap it in. Without this flag such a load is refused naming the difference; a rebuild is never a side effect of a scheduled load--pool <N>— Load the config’s tables on N worker threads instead of one after another: every freeing worker takes the next table, so a slow table no longer blocks the ones queued behind it. A failing table still isolates to itself and the rest keep loading, and the per-table lease is unchanged —rivet loadandrivet compactstill refuse a table the other holds. Each worker opens its own ledger connection, so N is also N connections to the state backend; a worker that cannot reopen the ledger takes no table, and the other workers load the queue. Defaults to 16 — the ceiling — capped at the number of tables. Pass--pool 1for the strictly sequential pass
rivet compact
Merge each base-and-buffer CDC table’s <table>__changes buffer into its base table (MERGE by primary key: updates, inserts, deletes flagged as __is_deleted) and drop the buffer — the billed step of the cycle run → load → compact, labelled rivet_op:merge per table
Usage: rivet compact [OPTIONS] --config <CONFIG>
Options:
-c,--config <CONFIG>— Path to YAML config file — the same onerivet loadreads--run-id <RUN_ID>— Correlation id stamped on every warehouse job of this compaction (BigQueryrivet_runlabel). Defaults to a generated id--pool <N>— Merge the config’s tables on N worker threads instead of one after another: every freeing worker takes the next table. A failing table still isolates to itself, and the per-table lease is unchanged — a tablerivet loadholds is still refused. Each worker opens its own ledger connection, so N is also N connections to the state backend; a worker that cannot reopen the ledger takes no table, and the other workers compact the queue. Defaults to 16 — the ceiling — capped at the number of tables. Pass--pool 1for the sequential pass
rivet state
Manage export state
Usage: rivet state <COMMAND>
Subcommands:
show— Show current state for all exportsreset— Reset state for an exportfiles— Show file manifest (files produced by exports)reset-chunks— Clear persisted chunk checkpoint rows (chunk_run/chunk_task)chunks— Show chunk checkpoint status for an exportprogression— Show committed / verified export boundaries (the last fully-exported cursor position)runs— Show the run-status ledger (extraction-run lifecycle rows gc/cleanup read)finish-run— Terminal-stamp a run-status row you KNOW is dead (hard crash, no successful successor) — the escape hatch for a prefix frozen by a stalerunningrowloads— Show the load ledger (rivet loadruns recorded in the state DB)
rivet state show
Show current state for all exports
Usage: rivet state show [OPTIONS] --config <CONFIG>
Options:
-c,--config <CONFIG>--json— Emit the incremental-cursor state as a JSON array to stdout instead of the text table. Empty →[]
rivet state reset
Reset state for an export
Usage: rivet state reset --config <CONFIG> --export <EXPORT>
Options:
-c,--config <CONFIG>-e,--export <EXPORT>— Export name to reset
rivet state files
Show file manifest (files produced by exports)
Usage: rivet state files [OPTIONS] --config <CONFIG>
Options:
-
-c,--config <CONFIG> -
-e,--export <EXPORT>— Show files for a specific export -
-l,--last <LAST>— Number of recent files to showDefault value:
50 -
--json— Emit the file list as a JSON array to stdout (CI completeness checks) instead of the text table. Empty →[]
rivet state reset-chunks
Clear persisted chunk checkpoint rows (chunk_run / chunk_task)
Usage: rivet state reset-chunks --config <CONFIG> <--export <EXPORT>|--stuck-checkpoints>
Options:
-
-c,--config <CONFIG> -
-e,--export <EXPORT>— Export whose chunk checkpoints should be cleared (same aschunk_checkpointruns) -
--stuck-checkpoints[alias:failed] — Reset checkpoints for every export named in this config that currently haschunk_run.status = 'in_progress'(crash, SIGKILL, stale concurrent worker).Ignores exports whose latest chunk run already finished (
completed). Runs listed in the database but removed from the YAML are skipped with a printed note.Alias
--failedrefers to “checkpoint state stuck”, not HTTP-style failures or metric rows.
rivet state chunks
Show chunk checkpoint status for an export
Usage: rivet state chunks [OPTIONS] --config <CONFIG> --export <EXPORT>
Options:
-c,--config <CONFIG>-e,--export <EXPORT>--json— Emit the checkpoint (run header + per-chunk tasks) as a JSON object to stdout instead of the text table. No checkpoint →null
rivet state progression
Show committed / verified export boundaries (the last fully-exported cursor position)
Usage: rivet state progression [OPTIONS] --config <CONFIG>
Options:
-c,--config <CONFIG>-e,--export <EXPORT>— Show progression for a specific export
rivet state runs
Show the run-status ledger (extraction-run lifecycle rows gc/cleanup read)
Usage: rivet state runs [OPTIONS] --config <CONFIG>
Options:
-
-c,--config <CONFIG> -
--running— Show onlyrunningrows — the ones that can freeze a prefix -
-l,--last <LAST>— Number of recent rows to showDefault value:
50 -
--json— Emit the rows as a JSON array to stdout instead of the text table. Empty →[]
rivet state finish-run
Terminal-stamp a run-status row you KNOW is dead (hard crash, no successful successor) — the escape hatch for a prefix frozen by a stale running row
Usage: rivet state finish-run --config <CONFIG> --run-id <RUN_ID>
Options:
-c,--config <CONFIG>--run-id <RUN_ID>— The run id to close (find it withrivet state runs -c <config> --running)
rivet state loads
Show the load ledger (rivet load runs recorded in the state DB)
Usage: rivet state loads [OPTIONS] --config <CONFIG>
Options:
-
-c,--config <CONFIG> -
-t,--target <TARGET>— Show only loads into this fully-qualified target (proj.ds.table) -
-l,--last <LAST>— Number of recent loads to showDefault value:
50
rivet completions
Generate shell completions
Usage: rivet completions <SHELL>
Arguments:
-
<SHELL>— Shell to generate completions forPossible values:
bash,elvish,fish,powershell,zsh
rivet init
Generate a config scaffold from a live database (connect + introspect)
Usage: rivet init [OPTIONS] <--source <SOURCE>|--source-env <ENV_VAR>|--source-file <PATH>>
Options:
-
--source <SOURCE>— Database URL (postgresql://, mysql://, sqlserver://, mongodb://, or oracle://). Visible in shell history /ps; prefer--source-envor--source-filefor anything other than local dev -
--source-env <ENV_VAR>— Name of an environment variable holding the database URL (e.g. DATABASE_URL). The URL never touches the command line -
--source-file <PATH>— Path to a file containing just the database URL (one line). Credentials stay on disk instead of entering the process command line -
--table <TABLE>— Single table, optionally schema-qualified (e.g. public.orders, dbo.orders). Omit to emit all tables/views in a Postgres/SQL Server schema or MySQL database -
--schema <SCHEMA>— PostgreSQL: schema to export (default public). SQL Server: schema (default dbo). MySQL: database name when the URL omits it (a –schema naming a DIFFERENT database than the URL is refused — put the database in the URL) -
--include <GLOB>— Whole-schema only: keep only tables/views matching these globs (*/?) — several after one flag (--include orders users) or the flag repeated; a table is kept if it matches any. No--include= keep all -
--exclude <GLOB>— Whole-schema only: drop tables/views matching these globs (*/?) — several after one flag or the flag repeated;--excludewins over--include -
-o,--output <OUTPUT>— Write output to this file instead of stdout -
--discover— Emit a machine-readable JSON discovery artifact instead of a YAML scaffold. Includes row estimates, size bytes, ranked cursor candidates, chunk candidates, and advisory notes. Mutually exclusive with the YAML-only--gcs-bucket/--s3-bucketflags -
--mode <MODE>— Override the suggested extraction mode for every scaffolded export.cdcscaffolds a change-data-capture export (mode: cdc + a cdc: block with engine-specific stream params) instead of a batch query; on MySQL, and on PostgreSQL when every table is inpublic, over two or more tables it writes one batch recipe per table plus onetables:stream withbackfill: auto(one export per table otherwise). Other values (full / incremental / chunked / time_window) just override the auto-suggested mode -
--gcs-bucket <NAME>— Scaffolddestination: type: gcswith this bucket (each export getsprefix: exports/<table>/). Incompatible with--s3-bucketand--discover -
--gcs-credentials-file <PATH>— Optional path forcredentials_file:on GCS scaffolds. Omit entirely to use ADC (gcloud auth application-default login) orGOOGLE_APPLICATION_CREDENTIALS— no key in YAML -
--s3-bucket <NAME>— Scaffolddestination: type: s3with this bucket (each export getsprefix: exports/<table>/). Incompatible with--gcs-bucketand--discover -
--s3-region <REGION>— Optional AWS region for S3 scaffolds (when using--s3-bucket) -
--bigquery-project <PROJECT>— Scaffold aload:block for this BigQuery project. With--bigquery-datasetthe generated config carries the warehouse target, a per-table partition guess and the base+buffer layout, sorivet loadandrivet compactwork from it after a review. Needs--gcs-bucket: the load reads GCS only, so a local or S3 scaffold with aload:block is a configrivet loadrefuses -
--bigquery-dataset <DATASET>— The dataset the load creates its tables in (with--bigquery-project) -
--clickhouse-url <URL>— Scaffold aload:block for this ClickHouse HTTP endpoint, e.g.http://localhost:8123. Needs--clickhouse-databaseand a bucket the export stages in (--gcs-bucketor--s3-bucket) -
--clickhouse-database <DATABASE>— The ClickHouse database the load creates its tables in (with--clickhouse-url) -
--clickhouse-user <USER>— The ClickHouse user the load authenticates as (with--clickhouse-url)Default value:
default -
--tls <MODE>— TLS posture for BOTH the introspection connection init opens AND thesource.tls:block written into the scaffold. Required (ordisable, explicitly) for any non-loopback host — without it the TLS gate refuses before connecting, and at init time there is no config file to add atls:block to yetPossible values:
disable: Plaintext. Use only inside trusted networks (loopback, cgroup-private)require: Require a TLS handshake; accept the server certificate without verifying issuer or hostname. Protects against passive sniffing, not MITMverify-ca: TLS + verify certificate chains to the configured / system trust store. Does not check hostname (useful for IP-addressed or internal names)verify-full: TLS + verify chain and hostname against the server cert’s SAN/CN. Recommended default for production
-
--tls-ca <PATH>— PEM CA certificate for--tls verify-ca/verify-fullagainst a private CA; written into the scaffold asca_file:. Refused withdisable/require, where it would be silently meaningless
rivet plan
Generate an execution plan artifact (no data exported)
Usage: rivet plan [OPTIONS] --config <CONFIG>
Options:
-
-c,--config <CONFIG>— Path to YAML config file -
-e,--export <EXPORT>— Plan only a specific export by name -
-p,--param <KEY=VALUE>— Query parameter: key=value (repeatable) -
-o,--output <OUTPUT>— Write plan JSON to this file (default: print summary to stdout) -
--annotate-waves— Write this plan’swave:/parallel_safe:schedule into the config, (over)writing every export. WITHOUT this flagrivet planis READ-ONLY: it prints the schedule and the reviewable plan but never touches the config file — not even to fill in absent fields. This makes config mutation an explicit, opt-in act (a read-only-lookingrivet planonce turned a hand-tuned 5-per-wave split into one 76-export wave) -
--format <FORMAT>— Output format: “pretty” (human summary) or “json” (machine-readable)Default value:
prettyPossible values:
pretty: Human-readable summary printed to stdoutjson: Pretty-printed JSON (written to –output file or stdout)
rivet apply
Execute a sealed plan artifact, or run a config’s exports wave-by-wave
Usage: rivet apply [OPTIONS] <PLAN_FILE>
Arguments:
<PLAN_FILE>— A plan JSON artifact fromrivet plan(sealed single-export replay), OR a YAML config (.yaml/.yml) to run its exports wave-by-wave in ascendingwave:order — the wave each export was assigned byrivet plan
Options:
--parallel-export-processes— Run the cheap (low-cost) exports within each wave concurrently, as separate processes (same asparallel_export_processes: truein the config). Config-wave mode only; heavier exports — which already chunk-parallelize internally — still run one at a time--resume— Config-wave mode: skip exports a prior run already completed (_SUCCESSpresent) and resume incomplete chunked exports from their checkpoints, so a re-run after a partial failure does not redo finished tables. Independent tables are never re-exported--force— Override whichever safety gate refuses the run: in JSON-artifact mode the plan staleness check (> 24 h) and the incremental cursor-drift check (each bypass is recorded in the run’sapply_context); in YAML config mode, with--resume, the refusal to resume into a prefix whose_SUCCESSmarker is already present--pool <N>— Run the whole config as ONE bounded work-stealing pool of N export slots (config mode only, #166): exports start longest-first (LPT, by each export’s last measured duration) and every freeing slot pulls the next — no wave barriers, so the wall approachesmax(longest, total/N). Prioritywave:tiers are NOT honored (makespan mode); exports that are notparallel_safenever run concurrently with EACH OTHER (one heavy at a time; cheap exports backfill the remaining slots)--split— With--pool: when ONE export dominates the pool floor (its predicted duration ≫ the next-longest, #167), split it into N range sub-exports over its key span — separate scheduler units the pool places concurrently, so the giant stops being the makespan floor. The units share one destination prefix and fold to one family, so the load view reads them as a single logical table. Only full/chunked/keyset exports with achunk_by_key:/chunk_column:are split (never incremental/CDC). Off by default; ignored without--pool
rivet repair
Targeted repair of chunks flagged by reconcile: emit a repair plan, or re-export only mismatched ranges
Usage: rivet repair [OPTIONS] --config <CONFIG> --export <EXPORT>
Options:
-
-c,--config <CONFIG>— Path to YAML config file -
-e,--export <EXPORT>— Export name to repair (must bemode: chunked) -
--report <REPORT>— Path to a reconcile JSON report produced byrivet reconcile --format json. Omit to run reconcile in-process against the latest chunk run -
--execute— Actually re-export the affected chunks. Without this flag, the plan is printed and nothing is executed -
--format <FORMAT>— Output format for plan / reportDefault value:
prettyPossible values:
pretty,json -
-o,--output <OUTPUT>— Write plan / report JSON to this file (with--format json) -
-p,--param <KEY=VALUE>— Query parameter: key=value (repeatable)
rivet validate
Re-run manifest-aware verification against an existing destination, no extraction.
The same file-manifest checks rivet run --validate performs at end-of-run, exposed as a standalone command for between-run polling and triage. Reads manifest.json + _SUCCESS at the destination, head-checks every committed part for presence and recorded size_bytes. Source is not queried — use rivet reconcile for a source-vs-export row audit.
By default validate resolves the destination prefix the same way run does — {date} becomes today’s UTC date. Use --date, --run-id, or --prefix to point at a prior run instead of today.
Usage: rivet validate [OPTIONS] --config <CONFIG>
Options:
-
-c,--config <CONFIG>— Path to YAML config file -
-e,--export <EXPORT>— Validate only this export (default: every export in the config) -
--format <FORMAT>— Output format: “pretty” (human summary) or “json” (machine-readable)Default value:
prettyPossible values:
pretty,json -
--depth <DEPTH>— How deep to verify: “light” (manifest + _SUCCESS only, no prefix listing), “sample” (light + part reconcile + untracked surplus), or “full” (sample + the value-checksum re-read of every part; CSV parts carry no value checksum, so for CSV only each part’s row count is re-counted).fullis the default and matches the pre-graded behaviour. Uselightfor a fast “is this a complete, marked run?” poll, orsamplefor full structural verification without downloading parts.Default value:
fullPossible values:
light: Manifest read + self-consistency +_SUCCESSonly (no prefix listing)sample: Light + part reconcile + untracked surplus (onelist_prefix)full: Sample + the Form B value-checksum re-read (downloads parts; CSV: row counts only)
-
-o,--output <OUTPUT>— Write JSON report to this file (only with--format json) -
--date <YYYY-MM-DD>— Resolve{date}to this ISO-8601 day (e.g.2026-05-21) instead of today.Use when a run that landed on a prior day’s prefix needs to be re-verified — without this flag
validatelooks at today’s resolved prefix and reports “no manifest” for yesterday’s data. -
--run-id <RUN_ID>— Substitute{run_id}in the destination template with this value.Composes with
--date. Has no effect if the template does not contain{run_id}. -
--prefix <PREFIX>— Skip placeholder resolution entirely and verify exactly this prefix.Use when the resolved template no longer matches the physical layout (e.g. data was relocated, or the template changed since the run landed). The destination type still comes from config (
local,s3,gcs,azure); only the resolvedpath/prefixstring is overridden.
rivet reconcile
Partition/window reconciliation: re-count per-partition on source and report mismatches. Requires a chunked export previously run with chunk_checkpoint: true. Exits non-zero when a mismatch is detected, so CI / orchestrators can gate on it (an unknown partition warns but does not fail)
Usage: rivet reconcile [OPTIONS] --config <CONFIG> --export <EXPORT>
Options:
-
-c,--config <CONFIG>— Path to YAML config file -
-e,--export <EXPORT>— Export name to reconcile (must bemode: chunked) -
--format <FORMAT>— Output format: “pretty” (human summary) or “json” (machine-readable report)Default value:
prettyPossible values:
pretty,json -
-o,--output <OUTPUT>— Write report JSON to this file (only with--format json) -
-p,--param <KEY=VALUE>— Query parameter: key=value (repeatable)
rivet metrics
Show export metrics history
Usage: rivet metrics [OPTIONS] --config <CONFIG>
Options:
-
-c,--config <CONFIG>— Path to YAML config file -
-e,--export <EXPORT>— Show metrics for a specific export -
-l,--last <LAST>— Number of recent runs to showDefault value:
20 -
--json— Emit the metrics as a JSON array to stdout (for CI / dashboards) instead of the text table. Empty history prints[]
rivet schema
Emit machine-readable schemas for Rivet’s data contracts.
Today: rivet schema config prints the JSON Schema for the rivet.yaml config to stdout. Operators pipe this into a file and reference it via a # yaml-language-server: $schema=... header so VS Code / Neovim’s YAML language server highlights invalid keys, suggests enum values, and surfaces required fields as the YAML is edited. See docs/cloud-destinations.md for the broader contract.
Usage: rivet schema <COMMAND>
Subcommands:
config— Print the JSON Schema describingrivet.yamlto stdoutcli— Print a Markdown CLI reference (every command + flag) to stdout, generated from the clap definitions — the same source as--help, so it cannot drift from the actual commandserrors— Print the Markdown error-code reference (everyRIVET_*code, its kind, exit code and operator action) to stdout, generated from the code registry
rivet schema config
Print the JSON Schema describing rivet.yaml to stdout.
The schema is generated from the running binary’s Rust types, so it always matches the config grammar this version accepts. Pipe to a file and reference it via a # yaml-language-server: $schema=… header in your config:
rivet schema config > rivet.schema.json
Usage: rivet schema config
rivet schema cli
Print a Markdown CLI reference (every command + flag) to stdout, generated from the clap definitions — the same source as --help, so it cannot drift from the actual commands.
rivet schema cli > docs/reference/cli-reference.md
Usage: rivet schema cli
rivet schema errors
Print the Markdown error-code reference (every RIVET_* code, its kind, exit code and operator action) to stdout, generated from the code registry.
rivet schema errors > docs/reference/errors.md
Usage: rivet schema errors
rivet journal
Inspect structured run journal (events, files, retries, quality issues)
Usage: rivet journal [OPTIONS] --config <CONFIG> --export <EXPORT>
Options:
-
-c,--config <CONFIG>— Path to YAML config file -
-e,--export <EXPORT>— Export name to show journal for -
-l,--last <LAST>— Number of recent runs to show (newest first)Default value:
5 -
--run-id <RUN_ID>— Show journal for a specific run_id instead of recent runs
This document was generated automatically by
clap-markdown.