Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Benchmark Methodology

How to run, interpret, and compare Rivet’s benchmark suites.

Rivet has two benchmark layers:

LayerToolPurpose
Micro-benchmarksCriterion (benches/)Hot-path throughput, compilation check, per-function regression gate
Cross-tool / cross-engine E2Edev/bench/smoke.py + docs/bench/matrix.yamlrivet vs 6 other tools on Postgres / MySQL / SQL Server / MongoDB — throughput, peak RSS, source-harm, type fidelity

Cross-tool / cross-engine E2E harness

The E2E harness compares rivet to duckdb, clickhouse-local, sling, ingestr, dlt, and odbc2parquet exporting the same fixture to Parquet, and captures what each tool does to the source (a co-running OLTP probe, longest query/txn, locks, native engine counters). It is a single source of truth: everything is driven from docs/bench/matrix.yaml by the runner dev/bench/smoke.py, which fails if the yaml declares a metric the code doesn’t capture. See docs/bench/README.md for prerequisites and docs/bench/report.html for the rendered results.

# system python has PyYAML + dlt; the homebrew pythons ship a broken pyexpat
/usr/bin/python3 dev/bench/smoke.py --engine postgres --table content_items
/usr/bin/python3 dev/bench/smoke.py --engine mysql --table content_items
/usr/bin/python3 dev/bench/smoke.py --engine mssql --table orders

Fixtures are seeded into a dedicated rivet_bench per engine (via the Rust seed tool; sizes in matrix.yaml) so the live-test fixtures are untouched. Three matrices print per run: benchmark, harm, type-loss.

Comparing rivet versions is a special case of the same harness — point RIVET_BIN (or $PATH rivet) at each build and re-run; the benchmark matrix’s rows_s / peak_mb columns are the comparison. rivet’s own steelman (mode: full, tuning.profile: fast, zstd) was chosen this way — a measured +24 % rows/s from dropping the balanced 50 ms/batch throttle.


Micro-benchmarks (Criterion)

Criterion benchmarks live in benches/ and measure specific hot paths in isolation.

# Run all benchmarks (full Criterion measurement)
cargo bench

# Run a specific group
cargo bench --bench hot_paths
cargo bench --bench resource_aware

# Compile and smoke-check (1 sample, no regression gate)
cargo bench --bench hot_paths -- --warm-up-time 1 --measurement-time 1 --sample-size 10

# Compare to a saved baseline
cargo bench --bench hot_paths -- --save-baseline main
# ... make changes ...
cargo bench --bench hot_paths -- --baseline main

Available benchmarks

BinaryGroupWhat it measures
hot_pathsparquet_write_batchParquet writer throughput for narrow / wide batches
hot_pathsquality_uniquenessQuality uniqueness tracking throughput
hot_pathscsv_write_batch, hash_column, column_scan, shape_tracking, mysql_parse_time, mysql_int_bytes, mysql_utf8_text_append, csv_binary_hex, csv_timestampRemaining hot-path groups (CSV writer, hashing, column scan, shape tracking, MySQL decode paths)
resource_awareauto_shrinkSplit overhead at different cap levels
resource_awarecompression_profilesPer-codec wall time for a 10,000-row batch
resource_awarerow_group_computationRow group target computation for narrow / wide schemas
resource_awarequality_uniqueness_capUniqueness tracking throughput with and without a cap

Criterion saves HTML reports to target/criterion/. Open target/criterion/index.html in a browser for violin plots and per-sample distribution.


CI integration

Only the Criterion micro-benchmark layer belongs in CI — it compiles and smoke-samples the Rust hot paths (a compile/panic check, not a regression gate). The cross-tool E2E harness is run manually: it needs four live database engines and per-engine vendor drivers, so it is not a CI job — reproduce it from docs/bench/README.md and publish docs/bench/report.html.

To add a micro-bench regression gate in the future, save a Criterion --baseline from a release tag and add a comparison step to ci.yml.


Interpreting results

Normal variance

E2E runs on shared CI (or a busy laptop) can show ±10–20% variance in wall time and ±5% in RSS depending on filesystem cache warmth and page cache pressure. Run each suite 3 times and take the median for a stable comparison.

When numbers look wrong

SymptomLikely cause
RSS much higher than expectedFilesystem cache not warm; first run always higher
Wall time much higher than expectedPostgres query planner chose a sequential scan; check indexes on bench tables
Files > 1 unexpectedlyFile splitting triggered; check max_file_size (export-level size string, e.g. “512MB”) in the config
Size(MB) unexpectedly largeWrong compression profile in the config template

Relating E2E numbers to rivet plan output

rivet plan shows a narrow–wide memory range for each export:

Batch memory : ~2 MB (narrow) – ~95 MB (wide)

The narrow bound assumes ~200 B/row; the wide bound assumes ~10 KB/row. The E2E benchmarks let you validate which bound your real table shape falls closer to by running with the same config and comparing the measured RSS(MB) against the plan estimate.


See also