Lesson 5 — Relations and hybrid search

Lesson 5 — Relations and hybrid search #

The two features that most distinguish Pizza from a plain text index: relations as first-class query targets, and retrieval pipelines you compose explicitly.

Relations #

The articles collection (44 documents) carries a related_to field of type relation — each document declares the articles it points at. A traverse query walks those edges:

Two documents are reachable at depth 0 (the seeds themselves), 17 at depth 1 — the known answer for this dataset. Change max_depth to 2 and watch the circle grow.

Seeds do not have to be ids or keys — any query can select them:

join is the sibling feature — it pulls matching documents from another collection through a join field ( join reference). It needs both collections co-located, so it is a server-side feature rather than a single-collection playground widget.

Hybrid search, explicitly #

Most engines hide their fusion. Pizza exposes the pipeline as request stages — retrieve runs retrievers in parallel, fuse merges them:

{
  "retrieve": [
    { "bm25": { "query": { "match": { "field": "title", "query": "vector database" } } } },
    { "vector": { "field": "embedding", "query_vector": [0.12, -0.4, 0.7], "k": 50 } }
  ],
  "fuse": { "method": "rrf", "k": 60 }
}

The dataset’s embeddings collection carries 24 documents in three well-separated 8-dim clusters (search, storage, vector) — small enough to stay deterministic, real enough to feel like embeddings. A kNN query with the first cluster’s direction returns exactly that cluster:

The retrieve/fuse pipeline itself is a server-side stage — create a collection with a dense_vector field (see vector search), index embeddings, and run the body above against it.

From there the pipeline continues: score_functions re-shapes scores (boost by popularity, decay by age), rescore runs an expensive second pass on only the top window. Each stage has its own section in the hybrid reference.

What you learned #

  • relation fields + traverse = graph queries without a graph database.
  • Seeds can be ids, keys, or an arbitrary query.
  • Hybrid search in Pizza is an explicit retrieve → fuse → score pipeline, not a magic flag.

Next: Lesson 6 — operations.

Calendar September 27, 2026
Edit Edit this page