Search #
A search query, or query, is a request for information about documents in Pizza collections.
A search consists of one or more queries that are combined and sent to Pizza. Documents that match a search’s queries are returned in the hits, or search results, of the response.
A search may also contain additional information used to better process its queries. For example, a search may be limited to a specific collection or only return a specific number of results.
The canonical DSL at a glance #
Every search body is one JSON object whose keys fall into three layers — what matches, how it ranks, and what comes back:
{
"query": { ... },
"retrieve": [ { "bm25": { ... } }, { "vector": { ... } } ],
"fuse": { "method": "rrf", "k": 60 },
"score_functions": { ... },
"rescore": { ... },
"post_filter": { ... },
"sort": [ { "field": "price", "order": "desc" } ],
"from": 0,
"size": 10,
"highlight": { ... },
"collapse": { ... },
"aggs": { ... }
}
- Matching —
query. A recursively composable tree: leaf queries (match,term,range,vector, …) combined bybool(must/should/filter/must_not, plusminimum_should_match). A body withoutqueryis amatch_all. - Ranking —
retrieve+fuse, thenscore_functions/rescore. Either a plainquery, or a hybridretrievelist — BM25 and vector legs side by side, mutually exclusive with a top-levelquery— merged byfuse(RRF / linear / disjunctive max). See hybrid search. - Shaping — everything else.
sortwithfrom/sizeorsearch_aftercursors,highlight,collapse,aggs, and_source/fieldsfiltering decide what the response contains — the full list is in request-level options and API conventions.
Leaf form #
A leaf query names one field and one condition with explicit keys — the canonical strict form used by every example in this documentation:
{ "term": { "field": "status", "value": "published" } }
{ "match": { "field": "message", "query": "this is a test" } }
{ "range": { "field": "age", "gte": 10, "lte": 20 } }
The same three shapes with ES shorthand ({"term": {"status": "published"}})
are accepted by the gateway as a migration courtesy only; the canonical
form is the documented standard, and unknown top-level body keys are
rejected with HTTP 400 instead of being silently ignored.
Execution model #
The body maps onto one execution pipeline. Matching produces candidates (scored doc ids); shaping reduces them to the returned page:
query ────────┐
│ aggs (computed over ALL matches,
retrieve ──┐ │ independent of pagination)
+ fuse ────┘ │
▼
pre_filter → score_functions → post_filter → rescore
▼
collapse → sort → from/size (or search_after)
▼
fetch _source → highlight
pre_filter cuts candidates before scoring; post_filter narrows the
hit list after scoring while aggregations still see the unfiltered
matches (faceted navigation). The hybrid path runs
retrieve → fuse → pre_filter → score_functions → truncate and rejects
the stages it does not run (aggs, collapse, rescore,
post_filter) by name — a stage is either executed or refused, never
silently skipped.
Examples #
Search all the collections under the default namespace whose names are ended
with -logs, fetch the documents whose field year has value 2024:
POST /default.*-logs/_search
{
"query": {
"term": {
"field": "year",
"value": "2024"
}
}
}
Requests #
POST /<targets>/_search
GET /<targets>/_search
POST /_search
GET /_search
A request without a targets path segment searches every collection the
caller may access.
Path parameters #
targets
(Optional, String) Comma-separated, names of the collection to search (wildcard supported)
Query parameters #
URL parameters override the corresponding body fields (Elasticsearch semantics) and apply to every request shape:
q
(Optional, string) Elasticsearch “lite” search — only consulted for a request without a body; becomes aquery_stringquery.from
(Optional, integer) How many documents to skip, should be non-negative and defaults to 0.size
(Optional, integer) The maximum number of documents to be returned inhits, defaults to 10.track_total_hits
(Optional, boolean) Whether to track the total hit count.key_as_id
(Optional, boolean, default:false) ES-compat hit rendering: keyed documents answer with the user key AS_idin every hit — the default keeps the system_idand reports the key as_key. Also accepted as a body field.
Response #
{
"took": 5,
"timed_out": false,
"hits": {
"total": { ... },
"hits": [ { "_id": "0,0", "_key": "...", "_score": 1.0, "source": { ... } } ]
},
"aggregations": { ... }
}
aggregations is present only when the request carried aggs. A search
that exceeded its deadline returns partial results with
"timed_out": true.
Request bodies are accepted in Pizza’s canonical DSL (the documented
standard — see
API conventions),
and the gateway additionally tolerates common Elasticsearch forms;
a bodyless GET is a match_all query.
Full-text queries #
matchquery
Returns documents that match a provided text, number, date or boolean value. The provided text is analyzed before matching.match_phrasequery
Returns documents that contain the exact sequence of terms of the provided text, in order.multi_matchquery
Returns documents that match a provided text across several fields with per-field boosts.query_stringquery
Returns documents that match a compact query-syntax string, e.g.status:active AND title:search.boolquery
Matches documents matching boolean combinations of other queries.match_allquery
Matches every document — the default when noqueryis sent.match_nonematches none.spanqueries
Low-level positional queries (span_term,span_near,span_or,span_multi) for phrase-adjacent logic.
Term-level queries #
existsquery
Returns documents that contain an indexed value for a field.fuzzyquery
Returns documents that contain terms similar to the search term within an edit distance.prefixquery
Returns documents that contain a specific prefix in a provided field.rangequery
Returns documents that contain terms within a provided range.regexpquery
Returns documents that contain terms matching a regular expression.suffixquery
Returns documents that contain terms ending in the provided value.termquery
Returns documents that contain an exact term in a provided field.termsquery
Returns documents that contain one or more exact terms in a provided field.wildcardquery
Returns documents that contain terms matching a wildcard pattern.
Geo queries #
geo_bounding_boxquery
Returns documents with a geo point inside a rectangle.geo_distancequery
Returns documents with a geo point within a radius.
Vector queries #
vectorquery
kNN similarity search over vector fields; accepts an embedding or text (embedded server-side).multi_vectorquery
Executes several vector queries and fuses the results (weighted sum or RRF).semanticquery
Semantic search with zero field knowledge: the vector field and embedding model resolve from the collection schema, the request is plain text; composes insideboollike any leaf.- Hybrid search — the native
retrieve+fuseDSL (BM25 + vector legs, RRF / linear fusion, strategy optimizer) plus the ESretrieverand top-levelknnsyntaxes translated at the gateway. See the vector reference and API conventions.
Relational queries #
nestedquery
Matches parent documents via queries on nested sub-documents.joinquery
Semi-join against a query on another index.traversequery
Bounded graph traversal over relation-typed edge fields.
Other supported query types #
Two variants of the catalog above round out the query surface:
match_none— matches no documents; useful as a placeholder branch in generated queries. See match all.- Every leaf also accepts
rewrite(constant-score and top-N forms) andcase_insensitivewhere exact matching applies.
Request-level options #
Beyond query, the search body accepts: sort and cursor pagination
(
sort and pagination), highlight
(
highlighting), hybrid retrieval and scoring stages
(
hybrid search), result deduplication
(
field collapsing), streaming delivery
(
streaming search), plus from, size,
track_total_hits, collect_size, timeout, explain, _source,
fields, epoch_range and point-in-time snapshot_id (see
API conventions).
Analysis #
- Analyze text
The analysis workbench: run analyzers and pipelines over sample text. - Built-in analysis components
The shipped catalog: 83 analyzers, 50 tokenizers, 15 normalizers and 327 token filters, from language analyzers to CJK segmentation.