Benchmark · extraction to Parquet
Eight tools, four engines, every axis in one pass: throughput, peak memory, source-harm, and type fidelity. Same fixture, each competitor on its own best-effort (steelman) config. No axis hidden — rivet is not the fastest, and this says so.
Peak memory
57 MB where competitors take 0.8–3.6 GB. Streaming vs buffering — architectural, not a config trick.
Type fidelity
The only tool that never loses a type across all engines and tables — while also checksumming every value.
Throughput
Competitive, not the leader — ingestr & clickhouse beat its rows/s. Named, not hidden.
Postgres · content_items (2 M heavy-text rows) → Parquet
Bars are linear on the real numbers. rivet is the sliver by design — it streams the source through a server-side cursor and never holds the result set in memory.
content_items, 2 M rows — all eight tools, steelman configs
| tool | rows/s | peak MB | out MB | files | oltp p99× | longq s | locks | type drift |
|---|---|---|---|---|---|---|---|---|
| rivet | 38 204 | 57 | 11.3 | 1 | 5.1 | 0.00 | 3 | 0 |
| rivet-chunked | 29 307 | 57 | 12.7 | 4 | 3.3 | 0.00 | 26 | 0 |
| duckdb | 31 635 | 2 067 | 9.3 | 1 | 7.1 | 7.7 | 63 | 2 |
| clickhouse | 39 164 | 820 | 38.7 | 1 | 2.7 | 50.3 | 3 | 5 |
| odbc2parquet | 27 447 | 3 579 | 31.1 | 1 | 2.7 | 40.0 | 3 | 2 |
| sling | 19 853 | 129 | 46.2 | 9 | 4.3 | 94.6 | 3 | 0 |
| ingestr | 51 677 | 1 288 | 41.3 | 1 | 6.1 | 36.9 | 3 | 2 |
| dlt | 7 267 | 1 735 | 114.4 | 1 | 1.8 | 1.5 | 34 | 2 |
longq is the source-safety headline: rivet holds no long-running query (a server-side cursor, not a single 40–95 s scan), while clickhouse/odbc/sling pin one query for the whole read. duckdb is fast but buys it with 2 GB and 63 locks; dlt spills 751 MB of temp on the source (not shown) and runs 7× slower.
Same tool, three engines — where the others break
| engine · table | rivet rows/s | rivet MB | rivet drift | notable competitor result |
|---|---|---|---|---|
| postgres · content_items 2M | 38 204 | 57 | 0 | duckdb fast but 2 GB / 63 locks |
| mysql · content_items 2M | 31 855 | 143 | 0 | duckdb 3 180 rows/s — one query held 8.6 min |
| mssql · orders 1M | 387 710 | 81 | 0 | duckdb & clickhouse: no native reader |
| mongo · content_items 200k | 29 795 | 369 | n/a | rivet 369 MB vs ingestr 875 · sling 510 |
duckdb's mysql_scanner collapses
King of throughput on Postgres (627 k rows/s on page_views), duckdb falls to 3 180 rows/s on MySQL content_items — a single query held open for 8.6 minutes at 1.7 GB. Same tool, same data, one engine away. rivet's per-engine reader keeps it at 32 k / 143 MB.
Source column → each tool's Parquet type (Postgres content_items + page_views)
| source column | type | rivet | duckdb | clickhouse | odbc | sling | ingestr | dlt |
|---|---|---|---|---|---|---|---|---|
| metadata | jsonb | json | text | text | text | json | text | text |
| created_at | timestamp | ts | ts | tz-shift | ts | ts | ts | ts |
| is_bounce | boolean | bool | bool | int | text | bool | bool | bool |
Every competitor flattens jsonb → plain text (loses the JSON logical type a reader needs). clickhouse also promotes naive timestamps to timestamptz — a silent wall-clock shift — and renders booleans as ints; odbc renders booleans as text. rivet and sling keep JSON; only rivet keeps everything, across every engine.
Steelman, applied to everyone. Each tool runs its lowest-memory config that still completes — the memory caps flatter competitors (a capped duckdb reports less RSS, not more). No arbitrary throttles; a self-audit removed a stray clickhouse thread cap.
Every axis reported side by side. rivet is beaten on rows/s by ingestr and clickhouse, and that is on the table — the transparency is the fairness. rivet's numbers already include its always-on per-value checksum (~7 %) that no competitor performs.
Measured, not theorised. rivet's steelman
(profile: fast) came from a measured +24 % rows/s — dropping
a 50 ms/batch throttle — not a guess; zstd was measured free vs snappy.
The honest position
rivet is the best where the thesis lives — memory footprint and data integrity — and the only tool that verifies every value while winning them. It is not the throughput leader, and this report does not pretend otherwise.
docs/bench/matrix.yaml (metric catalog +
steelman + seed) driving dev/bench/smoke.py, which guards against
metric drift. Fixtures seeded into a dedicated rivet_bench per
engine. MongoDB (JSON-blob, non-SQL) is covered — rivet / sling / ingestr, no
type dimension. All figures from live
runs, 2026-07-09.