Search

Search #

A search query, or query, is a request for information about documents in Pizza collections.

A search consists of one or more queries that are combined and sent to Pizza. Documents that match a search’s queries are returned in the hits, or search results, of the response.

A search may also contain additional information used to better process its queries. For example, a search may be limited to a specific collection or only return a specific number of results.

The canonical DSL at a glance #

Every search body is one JSON object whose keys fall into three layers — what matches, how it ranks, and what comes back:

{
  "query":           { ... },
  "retrieve":        [ { "bm25": { ... } }, { "vector": { ... } } ],
  "fuse":            { "method": "rrf", "k": 60 },
  "score_functions": { ... },
  "rescore":         { ... },
  "post_filter":     { ... },
  "sort":            [ { "field": "price", "order": "desc" } ],
  "from": 0,
  "size": 10,
  "highlight":       { ... },
  "collapse":        { ... },
  "aggs":            { ... }
}
  • Matching — query. A recursively composable tree: leaf queries (match, term, range, vector, …) combined by bool (must / should / filter / must_not, plus minimum_should_match). A body without query is a match_all.
  • Ranking — retrieve + fuse, then score_functions / rescore. Either a plain query, or a hybrid retrieve list — BM25 and vector legs side by side, mutually exclusive with a top-level query — merged by fuse (RRF / linear / disjunctive max). See hybrid search.
  • Shaping — everything else. sort with from/size or search_after cursors, highlight, collapse, aggs, and _source/fields filtering decide what the response contains — the full list is in request-level options and API conventions.

Leaf form #

A leaf query names one field and one condition with explicit keys — the canonical strict form used by every example in this documentation:

{ "term":  { "field": "status", "value": "published" } }
{ "match": { "field": "message", "query": "this is a test" } }
{ "range": { "field": "age", "gte": 10, "lte": 20 } }

The same three shapes with ES shorthand ({"term": {"status": "published"}}) are accepted by the gateway as a migration courtesy only; the canonical form is the documented standard, and unknown top-level body keys are rejected with HTTP 400 instead of being silently ignored.

Execution model #

The body maps onto one execution pipeline. Matching produces candidates (scored doc ids); shaping reduces them to the returned page:

   query ────────┐
                 │   aggs (computed over ALL matches,
   retrieve ──┐  │         independent of pagination)
   + fuse ────┘  │
                 ▼
        pre_filter → score_functions → post_filter → rescore
                 ▼
        collapse → sort → from/size (or search_after)
                 ▼
        fetch _source → highlight

pre_filter cuts candidates before scoring; post_filter narrows the hit list after scoring while aggregations still see the unfiltered matches (faceted navigation). The hybrid path runs retrieve → fuse → pre_filter → score_functions → truncate and rejects the stages it does not run (aggs, collapse, rescore, post_filter) by name — a stage is either executed or refused, never silently skipped.

Examples #

Search all the collections under the default namespace whose names are ended with -logs, fetch the documents whose field year has value 2024:

POST /default.*-logs/_search
{
  "query": {
    "term": {
      "field": "year",
      "value": "2024"
    }
  }
}

Requests #

POST /<targets>/_search
GET /<targets>/_search
POST /_search
GET /_search

A request without a targets path segment searches every collection the caller may access.

Path parameters #

  • targets
    (Optional, String) Comma-separated, names of the collection to search (wildcard supported)

Query parameters #

URL parameters override the corresponding body fields (Elasticsearch semantics) and apply to every request shape:

  • q
    (Optional, string) Elasticsearch “lite” search — only consulted for a request without a body; becomes a query_string query.

  • from
    (Optional, integer) How many documents to skip, should be non-negative and defaults to 0.

  • size
    (Optional, integer) The maximum number of documents to be returned in hits, defaults to 10.

  • track_total_hits
    (Optional, boolean) Whether to track the total hit count.

  • key_as_id
    (Optional, boolean, default: false) ES-compat hit rendering: keyed documents answer with the user key AS _id in every hit — the default keeps the system _id and reports the key as _key. Also accepted as a body field.

Response #

{
  "took": 5,
  "timed_out": false,
  "hits": {
    "total": { ... },
    "hits": [ { "_id": "0,0", "_key": "...", "_score": 1.0, "source": { ... } } ]
  },
  "aggregations": { ... }
}

aggregations is present only when the request carried aggs. A search that exceeded its deadline returns partial results with "timed_out": true.

Request bodies are accepted in Pizza’s canonical DSL (the documented standard — see API conventions), and the gateway additionally tolerates common Elasticsearch forms; a bodyless GET is a match_all query.

Full-text queries #

  • match query
    Returns documents that match a provided text, number, date or boolean value. The provided text is analyzed before matching.
  • match_phrase query
    Returns documents that contain the exact sequence of terms of the provided text, in order.
  • multi_match query
    Returns documents that match a provided text across several fields with per-field boosts.
  • query_string query
    Returns documents that match a compact query-syntax string, e.g. status:active AND title:search.
  • bool query
    Matches documents matching boolean combinations of other queries.
  • match_all query
    Matches every document — the default when no query is sent. match_none matches none.
  • span queries
    Low-level positional queries (span_term, span_near, span_or, span_multi) for phrase-adjacent logic.

Term-level queries #

  • exists query
    Returns documents that contain an indexed value for a field.
  • fuzzy query
    Returns documents that contain terms similar to the search term within an edit distance.
  • prefix query
    Returns documents that contain a specific prefix in a provided field.
  • range query
    Returns documents that contain terms within a provided range.
  • regexp query
    Returns documents that contain terms matching a regular expression.
  • suffix query
    Returns documents that contain terms ending in the provided value.
  • term query
    Returns documents that contain an exact term in a provided field.
  • terms query
    Returns documents that contain one or more exact terms in a provided field.
  • wildcard query
    Returns documents that contain terms matching a wildcard pattern.

Geo queries #

Vector queries #

  • vector query
    kNN similarity search over vector fields; accepts an embedding or text (embedded server-side).
  • multi_vector query
    Executes several vector queries and fuses the results (weighted sum or RRF).
  • semantic query
    Semantic search with zero field knowledge: the vector field and embedding model resolve from the collection schema, the request is plain text; composes inside bool like any leaf.
  • Hybrid search — the native retrieve + fuse DSL (BM25 + vector legs, RRF / linear fusion, strategy optimizer) plus the ES retriever and top-level knn syntaxes translated at the gateway. See the vector reference and API conventions.

Relational queries #

  • nested query
    Matches parent documents via queries on nested sub-documents.
  • join query
    Semi-join against a query on another index.
  • traverse query
    Bounded graph traversal over relation-typed edge fields.

Other supported query types #

Two variants of the catalog above round out the query surface:

  • match_none — matches no documents; useful as a placeholder branch in generated queries. See match all.
  • Every leaf also accepts rewrite (constant-score and top-N forms) and case_insensitive where exact matching applies.

Request-level options #

Beyond query, the search body accepts: sort and cursor pagination ( sort and pagination), highlight ( highlighting), hybrid retrieval and scoring stages ( hybrid search), result deduplication ( field collapsing), streaming delivery ( streaming search), plus from, size, track_total_hits, collect_size, timeout, explain, _source, fields, epoch_range and point-in-time snapshot_id (see API conventions).

Analysis #

  • Analyze text
    The analysis workbench: run analyzers and pipelines over sample text.
  • Built-in analysis components
    The shipped catalog: 83 analyzers, 50 tokenizers, 15 normalizers and 327 token filters, from language analyzers to CJK segmentation.
Calendar September 30, 2026
Edit Edit this page