Credits

Credits #

Pizza stands on the shoulders of the Rust and information-retrieval communities. This page credits the software we build on — and the projects whose ideas we learn from.

Reference points #

  • Lucene — decades of query semantics and scoring design that the DSL and scoring model learn from.
  • Elasticsearch — the JSON query DSL and REST API surface (_search, _doc, _bulk) the gateway is modeled on.
  • tantivy — reference point for segment-organized inverted indexes and query execution.

Engine core #

The search engine (lib/engine) is a no_std-capable crate that compiles from server to browser; these crates make that possible:

  • roaring — Roaring bitmaps carry every doc-id set: postings, filters, deletions.
  • fst-no-std and fst-regex — our no_std forks of BurntSushi’s fst: finite-state term dictionaries, with regex and Levenshtein automata running directly over them.
  • cedarwood — double-array tries for dictionary-driven analysis and suggestions (forked to drop its std dependency).
  • hashbrown — the hash map that works everywhere, including the wasm builds.
  • memmap2 — immutable segments are memory-mapped, never copied.
  • bitpacking — SIMD-packed posting lists.
  • zstd (with lz4 and flate2) — segment, column, and stored-source compression; a single .fire segment stacks all three where they win.
  • rayon and crossbeam — parallel query execution across segments and workers.
  • rhai — scripting for scripted aggregations and pipelines.
  • turbovec — TurboQuant 4-bit vector quantization for approximate kNN.
  • Apache Arrow (parquet, arrow-array, arrow-schema) — mounting Parquet files directly as searchable segments.
  • moka — segment cache, ureq — remote segment fetching, and fs4 — advisory file locks for the KV store.
  • nom — parser combinators behind the query_string mini-language.
  • wildmatch — wildcard query and route matching.
  • crc32fast and crc32c — CRC checksums over segments, WAL frames, and manifests; corruption is detected, never guessed around.

Distributed layer #

  • openraft — the Raft consensus that keeps cluster state and the catalog in agreement; we vendor it with openraft-rt-monoio, a runtime binding that lets Raft’s timers and heartbeats run on our monoio reactor instead of tokio.
  • capnp (with capnp-rpc) — the inter-node transport wire format; zero-copy both ways.
  • monoio — the thread-per-core, io_uring-based runtime the gateway and replication build on, with signalfut handling OS signals inside that runtime.
  • tokio — the async runtime for the pieces whose ecosystems assume it: the capnp-rpc transport endpoints and the actix-based modules. The split is deliberate — monoio owns the data path, tokio sits at the edges.
  • rskafka — the client behind the Kafka WAL backend, via our monoio-compatible fork.
  • sled — the embedded KV store holding Raft log and catalog state on disk.

Formal methods #

  • TLA+ / TLC — the protocols we built on top of Raft are specified and model-checked in contrib/formal-models before being trusted in code: promotion and WAL handover, shard allocation, online rolling reshard, doc-id reservation, inplace-column seeding, layered deletion visibility, and the membership fence. The TLC counterexamples have caught 18+ real bugs along the way.

Gateway and operations #

  • actix-web — HTTP types and client plumbing in the gateway.
  • serde and serde_json — every byte of the JSON surface, schema included.
  • clap — the CLI.
  • tracing — structured logs and the per-request trace tree.
  • jemallocator — the global allocator, with profiling and stats enabled.
  • sysinfo — host and disk facts behind node stats.
  • flatten-json-object — flattening nested JSON documents into dot-pathed fields.
  • arc-swap — lock-free swapping of routing tables and security snapshots: readers never pause while a new cluster state publishes.

Text analysis #

The per-language analyzer family (contrib/analysis-*, 40+ languages) is a Lucene-lineage port with embedded dictionaries, leaning on the best Unicode infrastructure available:

  • ICU4X — Unicode segmentation, normalization, casing, and collation.
  • jieba-rs — Chinese segmentation, alongside the SmartCN, IK, simplified↔traditional (stconvert), and pinyin analyzers.
  • lindera — Japanese morphological analysis with the ipadic dictionary (the kuromoji analyzer), and Korean via nori (ko-dic).
  • Stempel and Morfologik — Polish stemming and morphology, straight out of the Lucene contrib family.
  • The European stemmer set (analysis-english, -french, -german, …) and the Indic, Thai, Turkish, and Vietnamese analyzers round out the family.

Bindings and this site #

  • pyo3 — native Python bindings for the engine (lib/engine/bindings/python): the same segments and query engine, importable from Python.
  • meilisearch-docsearch — the search box in these docs is pizza-searchbox, our fork of the docsearch component that the Meilisearch and Tauri ecosystems use.

Browser builds #

  • wasm-bindgen (with js-sys and web-sys) — the engine compiled to WebAssembly, the same segment format mounted read-only in the docs playground.
  • serde-wasm-bindgen — zero-copy search results from wasm into JS.

Maintainers and contributors are listed in the repository and on the community page. Thank you all.

Calendar September 27, 2026
Edit Edit this page