Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Source-safe under load

The first question a serious operator asks about an extraction tool is what does it do to my source database while it runs? A careless SELECT * holds a long-running transaction, pins a read snapshot, inflates temp space, and spikes p99 latency for every other query on the box. Rivet is built so that the honest answer is “almost nothing you’ll notice.”

It holds no long-running query

A batch export streams the source in bounded pages — one chunk / one page at a time — and flushes each to a Parquet part before asking for the next. The longest query Rivet ever holds open on the source is a single page, not a full-table scan. In the cross-tool benchmark (PostgreSQL → Parquet, measured under a concurrent OLTP workload), the longest single server-side query each tool held was:

ToolLongest source queryPeak RSS
rivet0.00 s57 MB
rivet (chunked)0.00 s57 MB
duckdb7.7 s2 067 MB
odbc2parquet40.0 s3 579 MB
clickhouse-local50.3 s820 MB
sling94.6 s129 MB

Rivet is the only tool in the field that never parks a long-running read on the source. Everything else holds one server-side query open for the length of a full scan — 8 to 95 seconds here, and proportionally longer on a real table. That is the query your DBA sees in pg_stat_activity blocking autovacuum, or the one a pooler’s statement timeout kills at the worst moment.

It stays gentle on concurrent traffic

Under the same concurrent OLTP workload, the p99 latency multiplier (how much the export slows other queries) is mid-field for Rivet — 5.1× single-stream, 3.3× chunked — comparable to the rest, but achieved at 57 MB of RAM and zero long queries instead of hundreds of megabytes to multiple gigabytes with a scan pinned open. Rivet trades a little client-side CPU for a source that never sees a heavy query.

If the source is fragile, production-shared, or behind a pooler (pgBouncer, ProxySQL, MaxScale), that trade is the whole point. See MSSQL gentle extraction and resource-aware extraction for the per-engine knobs.

CDC is a passive reader too

Change-data-capture reads the transaction log, not your tables. Running a looping CDC drain against a live writer, the writer’s throughput barely moves — except on SQL Server, and that cost is inherent to the engine, not to Rivet:

EngineWriter throughput under CDCWhy
PostgreSQL0.98×logical decoding is light
MySQL1.07×binlog dump is a passive reader
MongoDB1.01×change stream reads the oplog
SQL Server0.78×the capture Agent duplicates every change into change tables — a second write SQL Server itself performs

The full per-engine numbers, harnesses, and contracts live in the performance & harm ledger. The one operational hazard that is per-engine — a CDC reader pinning source log retention — is covered on its own page: CDC cost, per engine.