Troubleshooting #
Symptom-first fixes for the issues that come up most often. For operations-level inspection (shards, epochs, slow log, trash) see node observability.
Requests fail with 401 #
Pizza starts with security enabled and prints a bootstrap API key on first boot. Pass it on every request until you mint proper keys:
curl -H "Authorization: Bearer pz_bootstrap_key" http://localhost:28000/
See security to mint scoped keys and administration security for static key configuration. Lost the bootstrap key? Delete the security state file under the data directory and restart — a new key is minted.
Everything returns 503 with a “phase” in the body #
The node is still starting or recovering — the startup gate returns 503 with the current phase until the catalog and shards are ready. Track progress:
curl -H "Authorization: Bearer $KEY" \
http://localhost:28000/_node/_local/recovery
GET /_node/_local/ready flips to 200 when the gate opens. Nothing to fix —
wait or check the recovery endpoint for a stuck phase.
Startup logs warn “ShardGroup … not found in global map” #
At boot, batches of lines like
[WAR] ShardGroup 593c493c2bc54d03b30a not found in global map during UpdateState apply, skipping
[WAR] ShardGroup 593c493c2bc54d03b30a not found in global map during SetRuntime apply, skipping
mean the node is replaying raft log entries that reference a rolling whose reshard lifecycle already finished: its shard groups were cleaned up and tombstoned, and replay skips re-creating them on purpose (resurrecting retired copies once produced phantom shards that double-counted documents). Nothing to fix — the metadata ends in the correct state, and the same shard-group ids repeat because one rolling owns all of them.
A graceful shutdown stores a catalog snapshot before the raft core stops
(shutdown snapshot stored at applied index N in the logs), and the next
boot restores it (restored stored snapshot at applied log ...) and replays
only entries committed after it — on a quiet node, nothing. After one
stop/start cycle these batches disappear. A crash (non-graceful exit) still
replays from the previous snapshot, so a bounded window of such warnings can
appear after crashes; it shrinks to the writes since the last graceful stop.
Connection refused on 28000 #
Check the bind address in pizza.yml (api.network.binding): the default
binds 127.0.0.1:28000 — remote clients are refused until you bind a
reachable address. On multi-instance machines every instance needs its
own port and data directory — see
configuration.
The playground says the wasm bundle or dataset is unavailable #
From a source checkout the browser bundle and dataset segments are build artifacts:
make dataset-seed && make dataset-fire # export teaching-dataset segments
make wasm # build the browser engine bundle
The widget explains the same thing inline when artifacts are missing.
Documents are visible immediately but disk usage grows fast #
That is realtime indexing doing its job: writes land in the mutable layer
and a WAL, then background epochs freeze and compact them. Watch the
pipeline via GET /_node/_local/tasks and trigger compaction early with
POST /{collection}/_compact or the
node-level endpoints in
node observability.
Search returns fewer hits than expected #
track_total_hitscaps counting at 10000 by default (gateway form) — set it totruefor exact counts.- A
termquery matches exact values — analysis is not applied. Usematchfor analyzed text. - Date fields compare in epoch millis; check the field’s
formatif you pass strings (see field parameters). - Cross-shard
join/traverseendpoints must be co-located; the teaching dataset keepsarticlessingle-shard for this reason (see traverse).
A query returns 400 “unknown field” #
The search body has a strict top-level whitelist. Misspellings
("form" instead of "from") surface as this error rather than being
silently ignored — check the key against the
search reference.
Where to look next #
- Cluster health for shard-level status
- Node observability for the full endpoint map
- Limitations for known sharp edges