Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

ADR-0012: Cloud Manifest Contract

Status: Accepted Date: 2026-05-21 (accepted; M1–M9 landed, incl. M8 chunked-resume executor and M9 best-effort quarantine move — see test-coverage table) Context: Rivet 0.7.0 introduces a public JSON manifest as the trust contract for cloud-output runs (local / S3 / GCS). The manifest is the operator-visible record of what was written, and the input to resume, validation, and reconciliation. Its invariants must be locked before the writer, the resume logic, and the verification extensions are coded — otherwise we will rewrite them.

This ADR defines those invariants. The shipping target is Rivet 0.7.0. Schema v8 has already reclaimed the manifest name by renaming the internal SQLite ledger to file_log (ADR refs: see CHANGELOG 0.6.1).


Goals

  1. Resume-aware cloud output: a re-run can decide, per part, whether to skip, rewrite, or quarantine.
  2. Trust verdict: --validate and --reconcile can give an unambiguous pass/fail by inspecting only the manifest and the destination (no local state required).
  3. Legacy compatibility: pre-0.7.0 runs work as-is, with no migration of in-flight state. See M6.
  4. Backend portability: identical semantics on local FS, S3-compatible, and GCS.

Non-goals

  1. Cross-engine column-level encryption metadata (a separate ADR if/when encryption ships).
  2. Schema evolution between successive runs of the same export (out of scope — schema_changed/fingerprint live in the run report, not in the manifest contract).
  3. A new rivet verify subcommand — explicitly rejected. The verdict is surfaced via the existing --validate / --reconcile / --report flags.

Artifacts

For every export run targeting a cloud or local-file destination, Rivet writes:

<destination_uri>/<export_layout>/
  part-000001.<format>
  part-000002.<format>
  ...
  manifest.json
  manifest-<run_id>.json       # immutable per-run copy of the manifest
  _SUCCESS                     # only if the run completed cleanly

<export_layout> is the destination-config-provided layout, typically <schema>.<table>/ and optionally namespaced by run_id. Layout policy is destination-config concern, not part of this ADR.

manifest.json is the authoritative record of the run for this export. Its schema is versioned (see “Versioning”) and stable across patch releases.

Every manifest write also leaves an immutable run-unique copy manifest-<sanitized-run_id>.json beside the canonical last-writer-wins manifest.json (src/manifest.rs::run_unique_manifest_name), so repeated runs into one prefix do not clobber prior runs’ records. The copies are Rivet-internal sidecars: resume, validate, and reconcile keep reading the canonical name, and the untracked-object scans exempt any manifest-*.json name (is_run_unique_manifest_name).

_SUCCESS is a single-line marker carrying the manifest fingerprint (xxh3:<16-hex>, src/manifest.rs::success_marker_body; see M2). Its only meaning is “the manifest at this prefix represents a fully-committed run”. Its existence implies the manifest exists, and every part the manifest references also exists at the recorded byte length.


Invariants

These extend ADR-0001 (I1–I7) into the cloud-destination plane.

M1 — Parts Before Manifest (PBM)

The manifest is written only after every part it references has been committed to the destination.

Rationale: A manifest pointing at a part that was never uploaded is a phantom record — worse than no manifest, because resume logic would skip work that wasn’t actually done.

Failure mode if violated: resume skips real work; --validate falsely reports completeness.

Recovery: a process killed between part-upload and manifest-write leaves the destination without a manifest. Resume detects “parts present, manifest absent” and re-derives the run state by listing parts (see M6 / M8).

M2 — Manifest Before SUCCESS (MBS)

_SUCCESS is written only after the manifest has been written and is readable at its destination URI.

Rationale: _SUCCESS is the single observable signal an external orchestrator (Airflow, Dagster, CI) can poll to decide “data is ready”. If _SUCCESS could appear before the manifest, downstream consumers reading the manifest would race.

Recovery: a process killed between manifest-write and _SUCCESS-write leaves the prefix with a manifest but no _SUCCESS. Resume treats this as “candidate complete; re-verify before finalizing”. Re-verification reads the manifest, checks each part, and writes _SUCCESS only if all checks pass.

_SUCCESS body: a single line xxh3:<16-hex>\n carrying the manifest’s content fingerprint (xxh3_64 over the exact bytes of manifest.json). This lets a polling consumer detect manifest changes (rerun, resume, repair) with a cheap GET _SUCCESS instead of re-reading the full manifest. The Hadoop empty-marker convention is not followed — Rivet does not target the Hadoop ecosystem and the fingerprint pays for itself the first time an Airflow sensor needs to distinguish “same successful run” from “new successful run at the same prefix”.

M3 — Part Identity Triple (PIT)

Every part referenced by the manifest is uniquely identified by (path, size_bytes, content_fingerprint).

The triple is recorded for each part. On resume, a part is considered the same as the manifested part if and only if all three components match. A part whose path matches but whose size or fingerprint differs is treated as corrupt or stale and quarantined (M9).

content_fingerprint is xxh3_64 over the part body, formatted "xxh3:<16-hex>". xxh3 was chosen because (a) the codebase already depends on xxhash-rust, (b) it streams at ~2 GB/s so the per-part cost is negligible against destination upload latency, and (c) the manifest is a trust contract for integrity, not for security — cryptographic hashes (sha256, blake3) are explicitly out of scope. The encryption / tamper-evidence track is deferred to a separate ADR if and when needed; until then, the xxh3: prefix in the on-wire format reserves the syntactic slot so a future cryptographic hasher can coexist without a schema break.

For 0.7.0, fingerprint is mandatory for new manifests. Pre-0.7.0 runs have no fingerprint and fall under M6.

M4 — Manifest Is Append-Only Per Run

A given run_id produces exactly one manifest. The manifest is never amended in place — a resumed run that completes additional parts writes a fresh manifest atomically (write-then-rename on local; atomic PUT on S3/GCS).

Rationale: partially-written manifests must be impossible to observe. Object stores give per-object atomicity for PUT/upload; local FS gets the same via write-temp-then-rename.

Resume across multiple interruptions does not produce multiple manifests for the same run — the latest write supersedes.

[Update: each manifest write now also leaves an immutable run-unique copy manifest-<run_id>.json (see Artifacts), so one run_id yields the canonical manifest.json plus one sidecar copy at the prefix. The canonical manifest is still never amended in place — the copy exists so repeated runs into one prefix keep every run’s record.]

M5 — SUCCESS Implies Verifiability

If _SUCCESS exists, then for every part listed in the manifest, the part is present at the destination at the recorded byte length.

This is the contract --validate checks on the metadata-only path: it lists the prefix, reads the manifest, and verifies M5 part-by-part. The listing also carries each object’s content MD5 (GCS md5Hash, S3/Azure single-PUT ETag), so --validate confirms content, not just size, with no download — the original “re-download to re-fingerprint” idea (--validate --deep) was rejected as wasteful. A part whose store gives no checksum (streamed multipart, local FS) verifies size-only; exports[].verify: content makes that a failure.

--reconcile adds: row counts in the manifest sum to the source COUNT(*) for the export’s row range.

M6 — Legacy Output Is Labeled, Not Migrated

Runs that completed before 0.7.0 (no manifest at the destination prefix) are not migrated. Operations on legacy prefixes succeed with reduced guarantees and must emit an explicit legacy_run: true label in operator-facing output.

Per the project decision taken at 0.7.0 planning (pre-0.7.0 runs keep the old behavior; the manifest applies only to new runs; every reduced check is explicitly labeled legacy_run, never silent):

  • --resume on a legacy prefix uses the pre-0.7.0 file-log-based logic; no manifest-aware skip.
  • --validate on a legacy prefix falls back to local-file row-count checks; manifest/M5 checks are skipped and reported as such.
  • --reconcile on a legacy prefix uses source-COUNT vs file-log only; the “manifest part-count match” line is omitted.

Silent fallback is forbidden. Every reduced check must say so in the report.

M7 — Manifest Atomicity

The manifest write is observable atomically: a reader either sees the previous state (manifest absent, or older manifest from a prior superseded run) or the new complete manifest. A partially-written manifest is unreachable to readers.

Local FS: write manifest.json.tmp then rename to manifest.json (POSIX rename(2) is atomic on the same filesystem).

S3 / GCS: write manifest.json as a single PUT / upload. Object stores guarantee write-completes-or-fails-with-no-trace.

Multipart uploads MUST NOT be used for the manifest itself — only the data parts. The manifest stays small enough (KB to single MB) that single-PUT suffices and side-steps the multipart abort/cleanup story.

M8 — Resume Decisions Are Deterministic

Given the same destination prefix and the same source/cursor state, --resume makes identical decisions on every run.

Decision matrix per part name:

Manifest entryObject presentSize matchesFingerprint matchesDecision
yesyesyesyesskip (committed)
yesyesyesnoquarantine (M9)
yesyesno—quarantine (M9)
yesno——rewrite (lost)
noyes——quarantine (M9) — untracked artifact
nono——new — write

_SUCCESS present + no --force → refuse to start (operator must opt in to overwrite a successful run).

The “no manifest entry / object present” row does not apply to the run-unique manifest copies (manifest-*.json, see Artifacts): both the reconcile and validate untracked-object scans exempt them via is_run_unique_manifest_name, so prior runs’ sidecar copies are never quarantined as untracked artifacts.

M9 — Untracked / Corrupt Parts Are Quarantined Best-Effort, Never Deleted

When resume finds an unknown or fingerprint-mismatch part, Rivet attempts to move it to a quarantine prefix and emits a warning. The move is best-effort: if it fails, the run still proceeds, the warning escalates, and the object stays where it was. Rivet never deletes unknown objects.

Quarantine layout: <prefix>/_quarantine/<run_id>/<original-name>.

Rationale — defensive: the unknown part may be the operator’s own intentional artifact, or evidence of a bug. Either way, Rivet preserves it and shifts the cost of cleanup to the operator.

Rationale — best-effort: on S3 / GCS the move decomposes into copy + delete, two non-atomic operations. A partial failure (copy succeeds, delete fails; or copy fails outright on a permissions issue) must not abort an otherwise-recoverable run. The reported warning carries enough detail (source path, destination quarantine path, failure reason) for the operator to finish the move manually. If the move never happens, the untracked part remains in place and re-trips M9 on the next resume — that is acceptable; an unmovable artifact is not a correctness problem, just a clutter problem.

Local FS gets the same best-effort behaviour: rename(2) is atomic but can still fail (different mount point, permissions, file-in-use on Windows). The semantics are uniform across backends — never bail on a quarantine failure.


Manifest schema (v1)

Field additions are backwards-compatible (consumers ignore unknowns). Field removals or type changes require a manifest_version bump.

{
  "manifest_version": 1,
  "run_id": "orders_20260521T120000.000",
  "export_name": "public.orders",
  "started_at": "2026-05-21T12:00:00.000Z",
  "finished_at": "2026-05-21T12:14:33.412Z",
  "status": "success",
  "source": {
    "engine": "postgres",
    "schema": "public",
    "table": "orders"
  },
  "destination": {
    "kind": "gcs",
    "uri": "gs://rivet-exports/public.orders/run_20260521T120000/"
  },
  "format": "parquet",
  "compression": "zstd",
  "schema_fingerprint": "xxh3:7f3a91be...",
  "row_count": 2001291,
  "part_count": 41,
  "parts": [
    {
      "part_id": 1,
      "path": "part-000001.parquet",
      "rows": 50000,
      "size_bytes": 123456789,
      "content_fingerprint": "xxh3:8a44e2c1...",
      "status": "committed"
    }
  ]
}

path is relative to the destination prefix so the manifest is portable across copies of the same dataset.

source.schema / source.table capture the logical name; the resolved SQL is not embedded — the manifest is about the output, not the extraction strategy. The run report (.rivet/runs/<run_id>/summary.json) carries the plan-side details.

schema_fingerprint is xxh3_64 over a canonical serialization of [{name, type}] from the existing state::SchemaColumn array. The fingerprint format prefix (xxh3:) is reserved so future fingerprint algorithms can coexist.

status per part: committed (in this manifest) or quarantined (the part listed in a prior superseded manifest that resume found corrupted; retained for audit).


What this does NOT define

  • Per-column encryption metadata.
  • Per-row provenance / lineage fingerprints.
  • Cross-run incremental cursor state (lives in export_state, surfaced by rivet state / rivet metrics).
  • Quotas, retention, or bucket policy.

These are intentionally outside the manifest. A manifest that tries to be a catalog will lose its trust-verdict role.


Decisions locked at ADR review

These items were open in the first draft of this ADR; they are now decided.

  1. _SUCCESS body — decided: carries the manifest fingerprint. See M2. A polling orchestrator can detect manifest changes between two successful runs (a rerun, a resume that completed, a repair) by reading the _SUCCESS body alone, without re-fetching the manifest. The Hadoop empty-marker convention is rejected — Rivet does not target the Hadoop ecosystem.
  2. Run-id segmentation in the destination prefix — decided: no automatic segmentation. Rivet writes parts, manifest, and _SUCCESS directly under the operator-configured destination prefix. Two successive runs against the same prefix produce one observable dataset whose manifest reflects the latest run; the prior run’s parts are reused (M8 skip), rewritten, or quarantined (M9) as the matrix dictates. Operators who want time-segregated historical runs include {run_id} (or {date}) in their destination URI themselves — that policy lives in the destination config, not in the manifest contract. Resume across overwrite is handled by the _SUCCESS gate plus --force (M8).
  3. Cryptographic / encryption-aware fingerprinting — decided: out of scope for 0.7.0. See M3. The xxh3: prefix reserves the slot.

Open questions deferred to implementation

  1. Quarantine TTL: Rivet does not delete quarantined objects. Operators may want a cleanup helper (rivet state remote --gc) — out of scope for 0.7.0.

Test coverage plan

InvariantStatus (2026-05-21)Test
M1✅ writer side coveredmanifest writer commits parts before manifest (pipeline::manifest_writer); kill-mid-write integration test deferred to Phase C-γ
M2✅ writer side covered_SUCCESS written iff status==Success; body = xxh3(manifest.json bytes); covered by success_marker_* tests + tests/offline/trust_artifacts_integration.rs §4 (compiled into the offline suite via tests/offline_suite.rs)
M3✅ write side + no-download content verifyper-part content_fingerprint (xxh3) and content_md5 recorded at write in one pass; --validate confirms content by comparing content_md5 to the store’s listing checksum (no download); resume still trusts size for skip decisions (quarantine on size drift) — covered by pipeline::resume_decisions::tests and pipeline::manifest_reconcile::tests
M4✅tests/offline/trust_artifacts_integration.rs §6 — writing_manifest_twice_replaces_the_previous_artifact
M5✅pipeline::validate_manifest + tests/offline/trust_artifacts_integration.rs §22 (manifest read, part presence, size match)
M6✅legacy_run: true label surfaced by verify_at_destination when no manifest present; covered in validate_manifest unit + integration tests
M7✅ writer relies on Destination::write atomicitylocal: fs::copy; S3/GCS: single PUT (opendal); covered by destination capability tests
M8✅ gate + matrix + chunked-resume executor wired--resume against _SUCCESS refuses without --force (covered §26); pure matrix tested per row in pipeline::resume_decisions::tests and end-to-end against real Destination listing in §27; executor apply_m8_resume_decisions runs as the resume preamble in both chunked runners (pipeline/chunked/resume_m8.rs, called from sequential_checkpoint + parallel_checkpoint)
M9✅ best-effort quarantine move wiredquarantine_move → Destination::move for divergent manifest parts (size/fingerprint) and untracked surplus objects; never fatal, never deletes on partial failure; counted on M8ResumeStats.{quarantined_moved, quarantine_move_failures} (pipeline/chunked/resume_m8.rs)

Each invariant lands with at least one unit test (local FS, fast) and one integration test (S3-compat via MinIO or GCS-compat; nightly).


Amendment 2026-08-27: the CDC cursor belongs in this contract, fenced (M10, proposed)

M1–M9 govern what a run records ABOUT its data in the destination. The CDC cursor — the position a next run resumes from — is not in that record. It is a JSON file on the local filesystem (src/source/cdc/mod.rs:98-139: temp file, fsync, rename; a corrupt or truncated file is refused rather than silently re-anchored). There is no fence and no owner on it.

That split has two silent failures, and neither is a bug in the code above:

  1. The two live in different durability domains. The data lands in object storage; the cursor lands on a container’s disk. A run in a fresh container finds no cursor and re-anchors — the “enable CDC during a quiet period” shape the process rules already records for MySQL, one layer up from the engine.
  2. Nothing arbitrates two writers. PostgreSQL refuses a second consumer of a replication slot, so that engine is protected by the server. MySQL, MongoDB and SQL Server are not: two processes on one checkpoint path both advance it and both report success. has_active_run_on_prefix answers a different question (orphan GC) and is not an owner check.

Proposed M10 — Cursor With The Data, Fenced. The durable CDC cursor is recorded in the destination, in the same write path as the manifest, and carries a monotonic generation plus the owning run id. A run re-reads that fence immediately before every advance and refuses (or invalidates its own generation) when it no longer owns it. The local file is demoted to a cache: it may make a resume faster, it may never be the sole source of truth. The durability ORDER is unchanged (flush → record → ack, per M1/M2 and ADR-0017); what changes is where the record lives and that it is fenced.

Primary prior art. The fencing token — a monotonically increasing generation the resource itself checks on every write, so a stalled or superseded owner cannot resume — is Kleppmann, Designing Data-Intensive Applications, ch. 8 (“The Truth Is Defined by the Majority” → fencing tokens). rivet already applies the shape once, in gc_orphans, where a superseded running row loses to a newer run by started_at rather than to a clock; M10 is the same discipline applied to the cursor.

RED-proof before this leaves Proposed. Two runs of one export against one destination, overlapping in time: the second is refused or invalidates the first’s generation, and the union of delivered rows equals the source. The mutant is the fence read removed from the advance path — that test must go RED. Plus the ephemeral half: delete the local cursor file between two runs and assert the second resumes rather than re-anchors, which is RED today.

Sequencing note: M10 touches the manifest write path, the state store, and every engine’s ack path. It is the largest of the four CDC amendments dated 2026-08-27 (the others are in ADR-0023 and ADR-0025) and the one most likely to need its own ADR before code.

Amendment 2026-09-26: what M1, M2 and M8 do today

M1. Recovery without a destination manifest rebuilds the committed parts from the state DB (completed chunk tasks, file_log) and uses the destination listing only to confirm they are present: a missing part’s chunk is reset and re-exported in the same run (keyset refuses instead). Listed parts the state DB does not name are not adopted. When the listing fails, the parts are declared from the state DB with a warning.

M2. Resume-time re-verification is not implemented for the chunked runners. They mark the chunk run completed before the dispatcher writes the manifest, so a crash between the manifest and _SUCCESS makes --resume plan afresh and re-export. --pool --split repairs a missing marker only when the completed units’ windows tile. The writer-side order (manifest before _SUCCESS) holds.

M8. The per-part decision matrix (apply_m8_resume_decisions) runs only in the two chunked-checkpoint runners, and only against a manifest carrying this run’s run_id. Single, plain chunked, keyset and mongo_parallel have no per-part matrix.