Credits #
Pizza stands on the shoulders of the Rust and information-retrieval communities. This page credits the software we build on — and the projects whose ideas we learn from.
Reference points #
- Lucene — decades of query semantics and scoring design that the DSL and scoring model learn from.
- Elasticsearch — the JSON query
DSL and REST API surface (
_search,_doc,_bulk) the gateway is modeled on. - tantivy — reference point for segment-organized inverted indexes and query execution.
Engine core #
The search engine (lib/engine) is a no_std-capable crate that compiles
from server to browser; these crates make that possible:
- roaring — Roaring bitmaps carry every doc-id set: postings, filters, deletions.
- fst-no-std and
fst-regex — our
no_stdforks of BurntSushi’s fst: finite-state term dictionaries, with regex and Levenshtein automata running directly over them. - cedarwood — double-array tries for dictionary-driven analysis and suggestions (forked to drop its std dependency).
- hashbrown — the hash map that works everywhere, including the wasm builds.
- memmap2 — immutable segments are memory-mapped, never copied.
- bitpacking — SIMD-packed posting lists.
- zstd (with lz4 and flate2) —
segment, column, and stored-source compression; a single
.firesegment stacks all three where they win. - rayon and crossbeam — parallel query execution across segments and workers.
- rhai — scripting for scripted aggregations and pipelines.
- turbovec — TurboQuant 4-bit vector quantization for approximate kNN.
- Apache Arrow (parquet, arrow-array, arrow-schema) — mounting Parquet files directly as searchable segments.
- moka — segment cache, ureq — remote segment fetching, and fs4 — advisory file locks for the KV store.
- nom — parser combinators behind
the
query_stringmini-language. - wildmatch — wildcard query and route matching.
- crc32fast and crc32c — CRC checksums over segments, WAL frames, and manifests; corruption is detected, never guessed around.
Distributed layer #
- openraft — the Raft
consensus that keeps cluster state and the catalog in agreement; we vendor
it with
openraft-rt-monoio, a runtime binding that lets Raft’s timers and heartbeats run on our monoio reactor instead of tokio. - capnp (with capnp-rpc) — the inter-node transport wire format; zero-copy both ways.
- monoio — the thread-per-core, io_uring-based runtime the gateway and replication build on, with signalfut handling OS signals inside that runtime.
- tokio — the async runtime for the pieces whose ecosystems assume it: the capnp-rpc transport endpoints and the actix-based modules. The split is deliberate — monoio owns the data path, tokio sits at the edges.
- rskafka — the client behind the Kafka WAL backend, via our monoio-compatible fork.
- sled — the embedded KV store holding Raft log and catalog state on disk.
Formal methods #
- TLA+ / TLC — the
protocols we built on top of Raft are specified and model-checked in
contrib/formal-modelsbefore being trusted in code: promotion and WAL handover, shard allocation, online rolling reshard, doc-id reservation, inplace-column seeding, layered deletion visibility, and the membership fence. The TLC counterexamples have caught 18+ real bugs along the way.
Gateway and operations #
- actix-web — HTTP types and client plumbing in the gateway.
- serde and serde_json — every byte of the JSON surface, schema included.
- clap — the CLI.
- tracing — structured logs and the per-request trace tree.
- jemallocator — the global allocator, with profiling and stats enabled.
- sysinfo — host and disk facts behind node stats.
- flatten-json-object — flattening nested JSON documents into dot-pathed fields.
- arc-swap — lock-free swapping of routing tables and security snapshots: readers never pause while a new cluster state publishes.
Text analysis #
The per-language analyzer family (contrib/analysis-*, 40+ languages) is
a Lucene-lineage port with embedded dictionaries, leaning on the best
Unicode infrastructure available:
- ICU4X — Unicode segmentation, normalization, casing, and collation.
- jieba-rs — Chinese segmentation, alongside the SmartCN, IK, simplified↔traditional (stconvert), and pinyin analyzers.
- lindera — Japanese morphological analysis with the ipadic dictionary (the kuromoji analyzer), and Korean via nori (ko-dic).
- Stempel and Morfologik — Polish stemming and morphology, straight out of the Lucene contrib family.
- The European stemmer set (analysis-english, -french, -german, …) and the Indic, Thai, Turkish, and Vietnamese analyzers round out the family.
Bindings and this site #
- pyo3 — native Python bindings for
the engine (
lib/engine/bindings/python): the same segments and query engine, importable from Python. - meilisearch-docsearch — the search box in these docs is pizza-searchbox, our fork of the docsearch component that the Meilisearch and Tauri ecosystems use.
Browser builds #
- wasm-bindgen (with js-sys and web-sys) — the engine compiled to WebAssembly, the same segment format mounted read-only in the docs playground.
- serde-wasm-bindgen — zero-copy search results from wasm into JS.
Maintainers and contributors are listed in the repository and on the community page. Thank you all.