Configuration

Configuration #

Pizza supports several methods to overwrite the default configuration.

Command lines #

➜  ./bin/pizza --help
A Distributed Real-Time Search & AI-Native Innovation Engine.

Usage: pizza [OPTIONS] [COMMAND]

Commands:
  service  Builtin service management (install, uninstall, start, stop)
  help     Print this message or the help of the given subcommand(s)

Options:
  -l, --log <LEVEL>           Set the logging level, options: trace,debug,info,warn,error
      --debug                 Run in debug mode, panic immediately with full stack trace
  -c, --config <FILE>         
  -p, --pid <FILE>            Place pid to this file
  -E, --override <KEY=VALUE>  
  -h, --help                  Print help
  -V, --version               Print version

Configuration file #

You can fully customize Pizza by utilizing the pizza.yml configuration file (the default lookup path; override it with -c/--config):

# ======================== INFINI Pizza Configuration ==========================

# -------------------------------- Log -----------------------------------------
log:
  level: info

# -------------------------------- API -----------------------------------------
api:
  network:
    binding: 127.0.0.1:28000
    # Optional: the address peers and clients should use when this node
    # is reachable on a different address than it binds (NAT, container
    # port mapping, or a proxy fronting the node). Defaults to `binding`.
    # advertise: 10.0.0.1:28000
    skip_occupied_port: true

# -------------------------------- Cluster -------------------------------------
cluster:
  name: pizza

node:
  name: my_node_1
  network:
    binding: 127.0.0.1:38000
    # Optional: the RPC address other nodes dial (registered in the
    # cluster metadata, used by raft and replication). Defaults to
    # `binding`. Set it when the node binds a private address but is
    # reachable through a public/mapped one.
    # advertise: 10.0.0.1:38000
    skip_occupied_port: true
  # Optional: external embedding services (OpenAI-compatible API) for
  # server-side text→vector inference — write-time derivation, text
  # vector queries, the `semantic` query, and /_embedding.
  # embedding:
  #   endpoints:
  #     - name: openai
  #       url: "https://api.openai.com/v1/embeddings"
  #       api_key: "sk-..."
  #       models: ["text-embedding-3-small"]
  #     - name: local
  #       url: "http://127.0.0.1:11434/v1/embeddings"   # Ollama
  #       models: ["nomic-embed-text"]
# -------------------------------- Storage -------------------------------------
storage:
  compression: ZSTD

# -------------------------------- MemTable ------------------------------------
memtable:
  threshold: 1k

# --------------------------------- WAL ----------------------------------------
# wal:
#   local:
#     durability: batched          # strict | batched | async
#     retention_epoch: 8           # keep WAL of 8 already-built epochs
#     retention_bytes: 536870912   # ... and at most 512 MiB per shard

max_num_of_instances: 2
allow_multi_instance: true

See the configuration reference for the full key catalog.

Multiple instances on one host #

max_num_of_instances and allow_multi_instance control how the node resolves its identity when its cluster’s data directory already contains node directories:

  • With allow_multi_instance: false (the conservative default), the node adopts the identity of the locked node directory even when that directory is locked by a running process, and fails fast if the on-disk identity would drift.
  • With allow_multi_instance: true, every locked directory counts toward max_num_of_instances; when all existing directories are locked, the node starts a fresh node id with empty local metadata (a warning is printed) — useful for test setups running several nodes on one machine, but on an accidental second start it silently forks the node’s identity.
  • skip_instance_detect: true skips the working-directory scan entirely; the node always uses node.id from the config as-is.

api.network.advertise and node.network.advertise decouple the address a node binds from the address it registers with the cluster. When unset (the default) both fall back to binding and nothing changes. Set them when the node is not dialable on its binding address — NAT or a cloud security group, a container port mapping, or a proxy fronting the node:

  • node.network.advertise is the RPC address other nodes dial: it is what lands in the cluster metadata and what raft voting, replication, and the failure detector use.
  • api.network.advertise is the HTTP address peers use to forward requests to this node (shown as the node’s reported address).

The node still listens on binding — advertise only changes what peers see. Local behavior such as skip_occupied_port is unaffected (it never touches advertise).

Override configuration #

You can tweak the configuration by passing the command line option -E with KEY=VALUE style during Pizza start:

./bin/pizza -E log.level=trace -E api.network.binding=127.0.0.1:12200 \
            -E node.network.advertise=10.0.0.1:8100

Dynamic settings #

pizza.yml is read once at startup; the file itself is never reloaded — keys like path.*, network bindings and raft timing are baked into the process. Runtime-tunable behavior instead lives in the layered settings model, edited through the HTTP APIs or visually in the built-in web console (/_ui/ → Settings):

built-in defaults
  └─ cluster settings   persistent (Raft-replicated, survives restart)
       └─               transient (until the next restart)
            └─ node settings    pizza.yml value (restart fallback)
                 └─             runtime override (until the next restart)
                      └─ collection settings   (metadata-persisted)
                           └─ rolling settings  (leaf-most override)

Where several levels can define the same concern, the narrower scope wins — a rolling’s replica count overrides the collection default, a node’s runtime override beats its pizza.yml value, and a cluster transient beats its persistent one.

ScopeWhat lives thereAPI
ClusterShard allocation & rebalance policy (cluster.routing.*)GET/PUT /_cluster/_region/<region_id>/settings
NodeSearch lane budgets, slowlog threshold, default timeout, trash retention, log level, epoch freeze/build, task lanesGET/PUT /_node/_local/settings and siblings
CollectionReplica default, shard-routing allocation policyGET/PUT /<target>/_settings
RollingPer-rolling replica overridePUT /<target>/_rolling/<rolling_id>/_settings

Two rules of thumb:

  • Cluster settings marked dynamic apply without restart via the update API; keys marked static (e.g. region.fault_detection.*) are rejected at runtime and must be set in the yml’s region: section before start.
  • Node settings are transient overrides over the yml value: they take effect immediately on that node only and disappear at restart — set the same key in pizza.yml when you want the change to persist. Each node is addressed individually (/_node/<node_id>/settings), so a cluster-wide change is a fan-out; the console does this for you.

See the API references for Region Settings and Node Settings.

Calendar September 29, 2026
Edit Edit this page