Backup and restore

Backup and restore #

Pizza does not have a first-class backup/restore HTTP API yet — this page describes what exists today and the supported manual procedure.

What exists today #

  • Raft metadata snapshots. The catalog’s Raft log is periodically snapshotted (node.snapshot_per_events), so cluster metadata (namespaces, collections, schemas, shard routing) has its own durability story independent of the data plane.
  • The trash. Deleting a namespace or a collection (or replacing shards) only touches metadata: the data directories are moved into <data_path>/trash/ and physically deleted by a background sweeper only after trash.retention_secs (default 1 day) elapses. A deleted namespace is captured as ONE item with a restore snapshot per collection, so the whole namespace — all collections, original shard placements included — can be brought back via the Trash APIs. Every restore is preceded by a read-only shard-completeness verification, and a restore whose data is provably incomplete is refused rather than silently coming back empty. This gives a grace window against accidental deletes, not a backup.

Manual backup #

A shard’s durable state is its built FIRE segments plus its WAL tail — both live under the node’s data directory. To take a consistent copy of a single-node deployment:

  1. Stop the node (or accept a crash-consistent copy of the data directory — the WAL replay on next start makes that safe for wal.local.durability: strict/batched).
  2. Copy the cluster’s data directory (path.data/<cluster.name>/) while writes are quiesced.
  3. Restore by copying the directory back and starting the node; WAL replay covers the tail.

For multi-node deployments, copying per-node data directories of every node at a quiesced point restores the cluster; per-shard copies of the same shard group are interchangeable up to their WAL positions.

A snapshot/restore API with point-in-time semantics is on the roadmap.

Calendar September 24, 2026
Edit Edit this page