Hybrid search and scoring pipeline #
Pizza runs lexical (BM25) and vector retrieval side by side and gives you explicit, composable stages to combine and re-shape scores:
retrieve → fuse → score_functions → rescore
(parallel (merge (function (second pass
retrievers) the lists) scoring) on the top window)
Every stage is optional and combines with a plain query: use
score_functions alone on a keyword search, or the full chain for
hybrid retrieval.
POST /<collection>/_search
Stage 1: retrieve — parallel retrievers
#
Replace the single query with a list of retrievers whose results are
fused. BM25 and vector legs side by side over the embeddings dataset:
{
"retrieve": [
{ "bm25": { "query": { "match": { "field": "text", "query": "vector search" } } } },
{ "vector": { "field": "embedding", "query_vector": [0.98, -0.05, 0.1, 0, -0.1, 0.05, -0.05, 0.1], "k": 5 } }
],
"fuse": { "method": "rrf", "k": 60 },
"size": 5
}
The fused answer includes every document found by EITHER leg (the match alone finds 24; the vector leg re-orders the top by similarity).
Retriever types #
bm25
(object) Carries aqueryin the normal query DSL — any leaf or boolean query.vector
(object) Carriesfield,query_vector(orquerytext /query_sparse),k,num_candidates— same options as thevectorquery.
Stage 2: fuse — combine the result lists
#
method
(Required,rrf|linear|disjmax) How lists are merged.rrf(reciprocal rank fusion) needs no comparable scores — it ranks by positions;lineartakes a weighted sum of normalized scores;disjmaxkeeps the max score per document.k
(Optional, integer, default:60) RRF rank constant — larger dampens the influence of top ranks.weights
(Optional, array of floats) Per-retriever weights forlinear.
For per-document multi-vector fusion inside one query see
multi_vector; the gateway also translates the
Elasticsearch top-level knn and retriever forms into these stages (see
API conventions).
Stage 3: score_functions — function scoring
#
Re-shape the score per document (demote stale docs, boost by
popularity). Composes with a plain query — no retrieve needed:
products matching “apparel”, scored by price with a log1p
modifier and a constant weight:
Compare _score against the plain match (≈1.4 per hit): the price
factor lifts expensive products; weight multiplies through.
Function types #
functions
(Required, array) One or more of:field_value_factor—{field, factor, modifier (none|log1p|log2p|sqrt|reciprocal|ln|square), missing}decay—{origin, scale, decay, offset, function (linear|exp|gauss)}— score falls off with distance fromorigin(date or numeric distance)script—{source, params}weight— constant multiplier
score_mode
(Optional,multiply|sum|avg|first|max|min, default:multiply) How the functions’ values combine with each other.boost_mode
(Optional,multiply|replace|sum|avg|max|min, default:multiply) How the combined function value combines with the query score.
Stage 4: rescore — second pass on the top window
#
Re-score only the top candidates with an expensive query (typically a vector or phrase pass over the top-K). Text match first, vector re-score over the top 10:
{
"query": { "match": { "field": "text", "query": "vector search" } },
"rescore": {
"window": 10,
"query": { "vector": { "field": "embedding", "query_vector": [0.98, -0.05, 0.1, 0, -0.1, 0.05, -0.05, 0.1], "k": 10 } },
"query_weight": 0.4,
"rescore_query_weight": 0.6
},
"size": 5
}
window
(Optional, integer, default:10) Number of top documents re-scored.query
(Required, object) The rescoring query.query_weight/rescore_query_weight
(Optional, float) Blend weights of original and rescored scores.
What composes with what #
retrievecannot currently combine withaggs,post_filter,collapseorrescorein the same body — the gateway rejects the mix with a clear 400 (verified live).score_functionsandrescorecompose with a plainquery— and with each other — without restriction.rerank(cross-encoder style re-ranking) is recognized but not implemented: requests fail fast with an explicit error naming the missing reranker-model serving layer.
Choosing a fusion method #
- Prefer
rrfwhen BM25 and cosine scores must not be compared directly — it works from ranks alone. - Prefer
linear/weightswhen you have calibrated, comparable scores. disjmaxwhen a document should score by its BEST leg (dedup-like behavior for multi-source retrieval).