Hybrid search and scoring pipeline

Hybrid search and scoring pipeline #

Pizza runs lexical (BM25) and vector retrieval side by side and gives you explicit, composable stages to combine and re-shape scores:

retrieve  →  fuse  →  score_functions  →  rescore
(parallel    (merge    (function           (second pass
 retrievers)  the lists) scoring)            on the top window)

Every stage is optional and combines with a plain query: use score_functions alone on a keyword search, or the full chain for hybrid retrieval.

POST /<collection>/_search

Stage 1: retrieve — parallel retrievers #

Replace the single query with a list of retrievers whose results are fused. BM25 and vector legs side by side over the embeddings dataset:

{
  "retrieve": [
    { "bm25": { "query": { "match": { "field": "text", "query": "vector search" } } } },
    { "vector": { "field": "embedding", "query_vector": [0.98, -0.05, 0.1, 0, -0.1, 0.05, -0.05, 0.1], "k": 5 } }
  ],
  "fuse": { "method": "rrf", "k": 60 },
  "size": 5
}

The fused answer includes every document found by EITHER leg (the match alone finds 24; the vector leg re-orders the top by similarity).

Retriever types #

  • bm25
    (object) Carries a query in the normal query DSL — any leaf or boolean query.
  • vector
    (object) Carries field, query_vector (or query text / query_sparse), k, num_candidates — same options as the vector query.

Stage 2: fuse — combine the result lists #

  • method
    (Required, rrf | linear | disjmax) How lists are merged. rrf (reciprocal rank fusion) needs no comparable scores — it ranks by positions; linear takes a weighted sum of normalized scores; disjmax keeps the max score per document.
  • k
    (Optional, integer, default: 60) RRF rank constant — larger dampens the influence of top ranks.
  • weights
    (Optional, array of floats) Per-retriever weights for linear.

For per-document multi-vector fusion inside one query see multi_vector; the gateway also translates the Elasticsearch top-level knn and retriever forms into these stages (see API conventions).

Stage 3: score_functions — function scoring #

Re-shape the score per document (demote stale docs, boost by popularity). Composes with a plain query — no retrieve needed: products matching “apparel”, scored by price with a log1p modifier and a constant weight:

Compare _score against the plain match (≈1.4 per hit): the price factor lifts expensive products; weight multiplies through.

Function types #

  • functions
    (Required, array) One or more of:
    • field_value_factor — {field, factor, modifier (none|log1p|log2p|sqrt|reciprocal|ln|square), missing}
    • decay — {origin, scale, decay, offset, function (linear|exp|gauss)} — score falls off with distance from origin (date or numeric distance)
    • script — {source, params}
    • weight — constant multiplier
  • score_mode
    (Optional, multiply|sum|avg|first|max|min, default: multiply) How the functions’ values combine with each other.
  • boost_mode
    (Optional, multiply|replace|sum|avg|max|min, default: multiply) How the combined function value combines with the query score.

Stage 4: rescore — second pass on the top window #

Re-score only the top candidates with an expensive query (typically a vector or phrase pass over the top-K). Text match first, vector re-score over the top 10:

{
  "query": { "match": { "field": "text", "query": "vector search" } },
  "rescore": {
    "window": 10,
    "query": { "vector": { "field": "embedding", "query_vector": [0.98, -0.05, 0.1, 0, -0.1, 0.05, -0.05, 0.1], "k": 10 } },
    "query_weight": 0.4,
    "rescore_query_weight": 0.6
  },
  "size": 5
}
  • window
    (Optional, integer, default: 10) Number of top documents re-scored.
  • query
    (Required, object) The rescoring query.
  • query_weight / rescore_query_weight
    (Optional, float) Blend weights of original and rescored scores.

What composes with what #

  • retrieve cannot currently combine with aggs, post_filter, collapse or rescore in the same body — the gateway rejects the mix with a clear 400 (verified live).
  • score_functions and rescore compose with a plain query — and with each other — without restriction.
  • rerank (cross-encoder style re-ranking) is recognized but not implemented: requests fail fast with an explicit error naming the missing reranker-model serving layer.

Choosing a fusion method #

  • Prefer rrf when BM25 and cosine scores must not be compared directly — it works from ranks alone.
  • Prefer linear/weights when you have calibrated, comparable scores.
  • disjmax when a document should score by its BEST leg (dedup-like behavior for multi-source retrieval).
Calendar September 27, 2026
Edit Edit this page