Benchmark · extraction to Parquet

rivet vs the field — across Postgres, MySQL, SQL Server & Mongo

Eight tools, four engines, every axis in one pass: throughput, peak memory, source-harm, and type fidelity. Same fixture, each competitor on its own best-effort (steelman) config. No axis hidden — rivet is not the fastest, and this says so.

rivet 0.18.0 hardware Apple M1 Pro · 10-core · 16 GB os macOS 26.2 (arm64) date 2026-07-09 source dev/bench/smoke.py + matrix.yaml

01 The one-line read

Peak memory

15–60× lower

57 MB where competitors take 0.8–3.6 GB. Streaming vs buffering — architectural, not a config trick.

Type fidelity

0 drift, everywhere

The only tool that never loses a type across all engines and tables — while also checksumming every value.

Throughput

mid-pack

Competitive, not the leader — ingestr & clickhouse beat its rows/s. Named, not hidden.

02 The picture · peak RSS

Postgres · content_items (2 M heavy-text rows) → Parquet

rivet
57 MB
sling
129 MB
clickhouse
820 MB
ingestr
1 288 MB
dlt
1 735 MB
duckdb
2 067 MB
odbc2parquet
3 579 MB

Bars are linear on the real numbers. rivet is the sliver by design — it streams the source through a server-side cursor and never holds the result set in memory.

03 Full matrix · Postgres

content_items, 2 M rows — all eight tools, steelman configs

toolrows/speak MBout MBfiles oltp p99×longq slockstype drift
rivet38 2045711.315.10.0030
rivet-chunked29 3075712.743.30.00260
duckdb31 6352 0679.317.17.7632
clickhouse39 16482038.712.750.335
odbc2parquet27 4473 57931.112.740.032
sling19 85312946.294.394.630
ingestr51 6771 28841.316.136.932
dlt7 2671 735114.411.81.5342
best / harmless notable costly rivet

longq is the source-safety headline: rivet holds no long-running query (a server-side cursor, not a single 40–95 s scan), while clickhouse/odbc/sling pin one query for the whole read. duckdb is fast but buys it with 2 GB and 63 locks; dlt spills 751 MB of temp on the source (not shown) and runs 7× slower.

04 Cross-engine · rivet holds; the field wobbles

Same tool, three engines — where the others break

engine · tablerivet rows/srivet MB rivet driftnotable competitor result
postgres · content_items 2M38 204570duckdb fast but 2 GB / 63 locks
mysql · content_items 2M31 8551430duckdb 3 180 rows/s — one query held 8.6 min
mssql · orders 1M387 710810duckdb & clickhouse: no native reader
mongo · content_items 200k29 795369n/arivet 369 MB vs ingestr 875 · sling 510

duckdb's mysql_scanner collapses

King of throughput on Postgres (627 k rows/s on page_views), duckdb falls to 3 180 rows/s on MySQL content_items — a single query held open for 8.6 minutes at 1.7 GB. Same tool, same data, one engine away. rivet's per-engine reader keeps it at 32 k / 143 MB.

05 Type fidelity · what the others drop

Source column → each tool's Parquet type (Postgres content_items + page_views)

source columntyperivetduckdb clickhouseodbcslingingestrdlt
metadatajsonbjsontexttexttextjsontexttext
created_attimestamptststz-shifttstststs
is_bouncebooleanboolboolinttextboolboolbool

Every competitor flattens jsonb → plain text (loses the JSON logical type a reader needs). clickhouse also promotes naive timestamps to timestamptz — a silent wall-clock shift — and renders booleans as ints; odbc renders booleans as text. rivet and sling keep JSON; only rivet keeps everything, across every engine.

06 How this stays honest

Steelman, applied to everyone. Each tool runs its lowest-memory config that still completes — the memory caps flatter competitors (a capped duckdb reports less RSS, not more). No arbitrary throttles; a self-audit removed a stray clickhouse thread cap.

Every axis reported side by side. rivet is beaten on rows/s by ingestr and clickhouse, and that is on the table — the transparency is the fairness. rivet's numbers already include its always-on per-value checksum (~7 %) that no competitor performs.

Measured, not theorised. rivet's steelman (profile: fast) came from a measured +24 % rows/s — dropping a 50 ms/batch throttle — not a guess; zstd was measured free vs snappy.

The honest position

rivet is the best where the thesis lives — memory footprint and data integrity — and the only tool that verifies every value while winning them. It is not the throughput leader, and this report does not pretend otherwise.

Single source of truth: docs/bench/matrix.yaml (metric catalog + steelman + seed) driving dev/bench/smoke.py, which guards against metric drift. Fixtures seeded into a dedicated rivet_bench per engine. MongoDB (JSON-blob, non-SQL) is covered — rivet / sling / ingestr, no type dimension. All figures from live runs, 2026-07-09.