How query works

How query works? #

A search request fans out from the coordinating node to every shard copy of the target collections, executes against a snapshot of the index, and merges ranked results back — this page walks that pipeline and the mechanisms that keep it fast.

Snapshot: mutable layer + immutable segments #

A query sees one consistent snapshot: the mutable in-memory layer (the documents behind realtime visibility) plus the immutable FIRE segments built so far. That is why a freshly indexed document is searchable immediately, and why a long-running query’s document set stays stable even as writes land.

One nuance: fields declared inplace: true are authoritative columns, and a column write’s value becomes visible the moment it lands — a snapshot pins the set of documents, not the values of hot counters (see In-place columns).

Dispatch and admission #

The coordinator picks a copy of each shard group to query, skipping copies on nodes whose search lanes are saturated (see below); a group with no eligible copy reports a failed shard and the query returns partial results rather than queueing indefinitely.

Each node runs two admission lanes:

  • FAST lane — regular searches, capped by node.search.max_concurrent (default 64).
  • HEAVY lane — aggregations, deep pagination, very large collect windows, capped by node.search.heavy_concurrency (default 8).

On the serving side, a per-shard cap (node.search.max_per_shard, default 16) bounds searches in flight against one shard copy — queue plus execution — protecting shards remote coordinators cannot see.

All three caps are runtime-tunable via Node Settings.

Timeouts and partial results #

Requests may carry a timeout; without one, the node’s cooperative deadline applies (node.search.default_timeout_ms, default 30s). The engine’s collector loops check the deadline as they work and return what they have with "timed_out": true — searches degrade to partial results instead of hanging.

Top-K pruning #

Ranked queries collect only the top size hits per shard, and the inverted index prunes candidate work with block-max WAND: whole blocks of postings are skipped when their best possible contribution cannot enter the current top-K. Term, terms and range queries over inplace columns are answered by the column itself, with block-level summaries that skip whole 4 096-slot blocks that cannot contain the queried range.

Aggregations #

Aggregations run per shard and merge on the coordinator. Percentile estimation uses TDigest by default (HDR histogram quantization optional). Bucket aggregations nest sub-aggregations under aggs; see the aggregation reference.

dense_vector / sparse_vector fields support kNN search — HNSW graph traversal with a num_candidates pool (default k × 2), or brute force where appropriate; min_score switches to threshold collection (all vectors above a similarity floor). Multi-field or mixed dense/sparse queries fuse rankings with weighted sum or Reciprocal Rank Fusion (see multi_vector).

Quantized indexes are selected per field in the schema (vector_options.index_type): BBQ binary quantization (~32x compression) and TurboQuant scalar quantization (2–4 bit, ~8x at 4-bit, zero training). Quantized scans produce the candidate pool and full-precision vectors rescore the top entries, keeping recall close to exact search at a fraction of the scan cost. Both the quantized codes and the flat vectors persist in the segment file, and segments carrying vector data always merge in the FireV11 format so the quantized sections survive compaction (see vector query). The turbo_quant_hnsw variant additionally persists an HNSW graph over the flat vectors in the same segment; once the segment is memory mapped, queries traverse the graph instead of scanning the codes — distances stay exact (flat vectors), so scores match the brute-force scan while only ef_search-bounded neighborhoods are evaluated. hnsw_flat carries the same graph without quantized codes. Graph shape and query beam width are schema-level parameters (m, ef_construction, ef_search), overridable per query.

Syntax tolerance #

Request bodies are accepted in both Pizza’s canonical DSL and common Elasticsearch forms — the normalization is idempotent and happens before parsing (see API conventions).

Calendar September 27, 2026
Edit Edit this page