Cluster management APIs #
Operator-facing /_cluster/* endpoints beyond
health and
state:
the allocation dashboard, node removal, and join-token management.
Regions #
GET /_cluster/_regions
A lightweight summary of every region this node knows about — identity
(name, uuid, zone), leader, member nodes with their API/RPC addresses,
and per-region counts. The region serving the request is marked local,
telling you which /_cluster/_region/_local/... endpoints apply to it.
Allocation dashboard #
GET /_cluster/allocation
The leader’s shard-allocation view: the effective cluster.routing.*
settings, per-node shard load, the recent allocation/relocation task
history (newest first) and how many tasks are in flight. Answers “what
is the allocator doing right now”.
GET /_cluster/allocation/explain
A diagnostic snapshot of why allocation is paced the way it is: whether
cluster.routing.allocation.enable is on, the concurrency limits in
effect (concurrent_tasks, node_concurrent_recoveries), per-node load
and incoming recovery counts, and which gate is currently holding work
back.
The routing policy behind both views is tunable at runtime — see Region Settings.
Node removal #
DELETE /_cluster/nodes/<node_id>
Removes a node from Raft membership and the cluster metadata. This is the operator escape hatch next to the failed-follower detector: a node that is half-dead (TCP accepts, RPC never answers) or retired on purpose needs a manual removal.
Notes:
- Shard failover is not triggered by the removal itself — the allocator re-places the node’s shards on its upcoming ticks.
- Must be called on the leader (the console connects to the catalog manager, which is).
- The node serving the request refuses to remove itself (
409).
Join tokens #
Dynamic cluster-admission tokens. Minting the first token switches
the cluster into enforced mode — from then on, joining nodes must
present a token (catalog.join_token in their pizza.yml). Mint and
revoke replicate through the catalog Raft and ride the catalog snapshot,
so tokens (and the enforced flag) survive restarts. The plaintext secret
is echoed exactly once, at mint time; the list endpoint only shows a
masked prefix.
POST /_cluster/join_token
{
"ttl_secs": 3600,
"region": "pizza_region",
"note": "rack-2 expansion"
}
ttl_secs— (Optional, integer) Seconds until the token expires. Default 1h, capped at 24h.region— (Optional, string) Informational: carried into the generated join script.note— (Optional, string) Free-form note shown in the token list.
GET /_cluster/join_token # list (masked secrets)
DELETE /_cluster/join_token/<id> # revoke
Shard reassignment #
POST /_cluster/shards/reassign
Asks the allocation service to re-run shard placement — used by the console after topology changes.