Who is Rivet for?
A short, honest fit-check. Rivet is intentionally narrow — the goal is to do one thing well, not to be the only data tool you need. This page exists so you can decide in 60 seconds whether to keep reading, or whether something else is a better fit for your problem.
If you came here from a search for “Postgres / MySQL / SQL Server / MongoDB → Parquet” or “extract a big SQL table without an OOM”, you are probably in the right place.
Yes, Rivet is probably a good fit if…
- You need to dump rows from PostgreSQL, MySQL, or SQL Server (or documents from MongoDB) into Parquet or CSV files — locally, on S3, GCS, or Azure Blob Storage.
- The source database is fragile, production-shared, or behind a
pooler (pgBouncer, ProxySQL, MaxScale) and a careless
SELECT *is going to hurt someone. - You want resumable extraction — the job can crash, the network can blip, and the next run continues from a checkpoint instead of starting over.
- You want a manifest +
_SUCCESStrust contract so a downstream loader can decide exactly which files belong to a given run. - You are happy operating Rivet from cron, Airflow, GitHub Actions, Argo, a one-off shell script, or a Kubernetes Job — anything that can invoke a single CLI binary with a YAML config.
- You can write a SQL query. Rivet does not abstract SQL away; it runs the query you give it.
No, Rivet is probably not the right tool if…
| You actually need… | Use this instead |
|---|---|
| Always-on live streaming — every insert/update/delete pushed continuously into Kafka or Kinesis as it happens | Debezium, Estuary, Materialize, or your cloud’s native CDC (AWS DMS, GCP Datastream). (Rivet does capture CDC — inserts/updates/deletes — but to files: mode: cdc, resumable, per-invocation, not a live stream. See semantics.md.) |
| A SaaS connector marketplace — pre-built connectors for Salesforce, Stripe, NetSuite, Shopify, Hubspot, etc. | Airbyte, Fivetran, Stitch |
A managed warehouse loader — a continuously-managed service that loads every warehouse (Redshift, Databricks, …) as one product. (Rivet does load BigQuery / Snowflake — rivet load, a discrete command you schedule, not a managed service; recipe.) | Fivetran, Airbyte (cloud), dlt (self-hosted with destinations), Sling |
| In-warehouse transformation — modeling, joins, materializations, lineage | dbt, SQLMesh |
| A data orchestrator — DAGs, retries-with-callbacks, schedule UI, lineage graphs | Airflow, Dagster, Prefect |
| A Kubernetes operator / Helm chart for an extraction platform | Rivet runs as a single binary in a Job or CronJob; a heavier platform like Airbyte on Kubernetes is a different architecture |
| Exactly-once delivery to the destination — no chance of a duplicate file under any failure mode | Rivet provides at-least-once file delivery + manifest; consumers deduplicate on the warehouse side (recipe). If you need exactly-once at the file layer, use a transactional sink (warehouse MERGE, lake table commit). |
| A data catalog / governance / PII detection layer | Amundsen, DataHub, Collibra |
| A query engine that reads from Postgres/MySQL and joins with other sources at query time | DuckDB (postgres_scanner/mysql_scanner), Trino, ClickHouse |
If you find yourself trying to bend Rivet into one of the above shapes, stop and pick the tool above instead. We are not trying to become any of those, and shoehorning will be painful.
Edge cases — Rivet can do this, but read first
-
Very large single-table dumps (100M+ rows). Yes, but read
docs/modes/chunked.mdfirst — chunked mode with the right cursor column is the difference between an export that finishes in 20 minutes and one that holds a single SQL statement open for 2 hours. -
Sources with weak or missing primary keys / cursor columns. Rivet has
incremental_cursor_mode: coalescefor nullable primaries (see composite cursor walkthrough) and keyset (chunk_by_key) for chunking without an integer key — but these surface tradeoffs inrivet check(sparse range warnings). Look at the warnings, do not ignore them. -
Read replicas with replication lag. On PostgreSQL, a full-mode export runs inside a single snapshot transaction, so the exported rows are point-in-time consistent as of the replica’s “now”. Chunked/keyset exports issue independent short queries (parallel workers each on their own connection), and other engines (e.g. MySQL) run plain autocommit SELECTs — there consistency is per-query, not per-run. Either way, a lagging replica gives you the replica’s “now”, not the primary’s. Operator’s responsibility to choose the right endpoint.
-
Sources behind SSH bastions / jump hosts. No native SSH tunnel support yet (tracked in
rivet_roadmap.mdEpic 13). Usessh -Lorautosshto forward the port and point Rivet atlocalhost:<forwarded>. -
Stateless / ephemeral runners (Kubernetes pods, Lambda, ECS tasks). Set
RIVET_STATE_URLto a PostgreSQL state backend so cursors and checkpoints survive pod death. See the README § Stateless deployment.
Decision shortcut
Need always-on live streaming (into Kafka)? → Debezium / Estuary
Need CDC captured to files (resumable)? → Rivet (mode: cdc)
Need a connector for a non-DB SaaS source? → Airbyte / Fivetran
Need a managed extract+load product? → Fivetran / Airbyte Cloud
Load extracted data into BigQuery / Snowflake? → Rivet (rivet load)
Need a SQL-based transformation framework? → dbt
Need an orchestrator? → Airflow / Dagster
Need to dump PG/MySQL to Parquet/CSV safely? → Rivet
If you stayed on the last line, the
Getting Started guide is ~3 minutes and ends
with one Parquet file you can arrow read or duckdb 'SELECT * FROM read_parquet(…)' against.