Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Rivet v0.5.x Benchmark Report

⚠️ HISTORICAL — archived. These numbers are from v0.5.0 (2026-05-15), pre-dating the streaming read path that changed every RSS/throughput figure in the newer reports. Retained only as the pattern for config-tuning reports and as the cited evidence behind older best-practices claims. Pending a 0.18 re-measure — do not quote these figures as current.

Measured on: 2026-05-15
Binary: target/release/rivet (v0.5.0)
Host: macOS Darwin 25.4.0, Apple Silicon
Database: PostgreSQL 16 (local docker-compose)


§4.3 Compression Profiles

Dataset: bench_narrow — 500,000 rows, 5 numeric/timestamp columns, avg ~40 B/row

ProfileCodecWall (s)User (s)RSS (MB)Output (MB)
noneuncompressed0.930.1146.413
fastSnappy0.920.1149.510
balancedZstd-30.940.1450.25
compactZstd-91.040.2882.25

Key findings:

  • balanced (Zstd-3) compresses 2.6× better than none with identical wall time and only 4 MB more RSS. This is the recommended production default.
  • fast (Snappy) is slightly faster than balanced but produces 2× larger files. Use for large backfills where storage cost is secondary.
  • compact (Zstd-9) offers no additional compression benefit over balanced on numeric data, while using 64% more memory and 2× more CPU. Only use when storage/network cost dominates.
  • Wall time is dominated by source query and Parquet serialization, not compression. All profiles are within 12% of each other.

§4.4 Row Group Targets

Dataset: bench_wide — 100,000 rows, 10 TEXT columns, 200 chars each (~2 KB/row)

TargetRow group strategyWall (s)RSS (MB)Output (MB)
32 MBauto2.2675.91
64 MBauto2.2579.51
128 MBauto2.2682.71
256 MBauto2.2582.31

Key findings:

  • Smaller row group targets reduce peak RSS with no wall-time penalty — latency is identical across all targets.
  • 32 MB target saves ~7 MB RSS vs 256 MB on bench_wide (2 KB rows). On wider tables the difference is larger — see the content_items benchmark in low-memory-runners.md where it contributes to a 5.7× RSS reduction.
  • auto strategy chooses row count dynamically from Arrow schema widths. Users don’t need to tune a row count — they specify a memory budget and the writer adapts.
  • Output size is identical across targets (compression ratio is not affected by row group boundaries).

§4.5 Batch Memory Policies

Dataset: bench_wide — 100,000 rows, 10 TEXT columns, 200 chars each (~2 KB/row)
Cap: 64 MB, batch_size: 10,000 (estimated batch ~20 MB — cap does not trigger on this dataset)

PolicyWall (s)RSS (MB)Output (MB)
warn (cap=64 MB)2.2582.21
auto_shrink (cap=64 MB)2.2681.11
no cap (warn, 4 GB)2.2683.71

Key findings:

  • When batches stay under the cap, warn and auto_shrink add zero measurable overhead vs no cap. The policy check is a single memory comparison per batch.
  • All three policies produce identical output and identical wall time, confirming no correctness regression from the cap mechanism.
  • On bench_wide at batch_size: 10,000, each batch is ~20 MB in Arrow — below the 64 MB cap. To observe RSS reduction from auto_shrink, use wider tables or larger batch sizes.

Content_items reference (200,000 rows, avg ~3 KB/row, 12 columns including TEXT/JSONB):

Configbatch_sizePeak RSSWall (s)
No cap (batch_size: 25,000)25,000878 MB17.2
Safe baseline (max_batch_memory_mb: 64)~2,000154 MB16.6
Tight (batch_size: 500, cap 32 MB)500111 MB16.3

The 5.7× RSS reduction holds for wide real-world tables. See low-memory-runners.md for the full methodology.


§4.6 Quality Uniqueness

Dataset: bench_hc — 200,000 rows, UUID + email columns (high cardinality)

ConfigWall (s)RSS (MB)Output (MB)
unique_max_entries: 50,000 (capped)1.5441.05
no cap (200,000 unique values tracked)1.5441.85

Key findings:

  • At 200,000 rows, uncapped uniqueness tracking adds only ~0.8 MB RSS above the capped baseline. xxHash3-64 stores u64 hashes (8 bytes), so 200K × 2 columns × 8 bytes = ~3.2 MB of hash sets — negligible against total process RSS.
  • Wall time is identical: hash-based tracking is O(1) per row.
  • unique_max_entries is still recommended for high-cardinality tables as a defensive cap, not because the overhead is large. Without it, a runaway uniqueness tracking on a billion-row table could grow to hundreds of MB.
  • The hash-based approach (xxHash3-64, typed) is a quality signal, not an exact distinct count.

Summary

ClaimEvidence
balanced compression: same speed as none, 2.6× smaller files§4.3 ✓
compact adds no compression benefit over balanced on numeric data§4.3 ✓
Smaller row group targets reduce RSS with zero wall-time cost§4.4 ✓
Memory policies add zero overhead when cap doesn’t trigger§4.5 ✓
auto_shrink reduces RSS 5.7× on wide real-world tables§4.5 content_items ✓
xxHash3-64 uniqueness tracking overhead is < 1 MB on 200K rows§4.6 ✓
unique_max_entries cap recommended as defensive limit§4.6 ✓