How query works? #
A search request fans out from the coordinating node to every shard copy of the target collections, executes against a snapshot of the index, and merges ranked results back — this page walks that pipeline and the mechanisms that keep it fast.
Snapshot: mutable layer + immutable segments #
A query sees one consistent snapshot: the mutable in-memory layer (the documents behind realtime visibility) plus the immutable FIRE segments built so far. That is why a freshly indexed document is searchable immediately, and why a long-running query’s document set stays stable even as writes land.
One nuance: fields declared inplace: true are authoritative columns,
and a column write’s value becomes visible the moment it lands — a
snapshot pins the set of documents, not the values of hot counters
(see
In-place columns).
Dispatch and admission #
The coordinator picks a copy of each shard group to query, skipping copies on nodes whose search lanes are saturated (see below); a group with no eligible copy reports a failed shard and the query returns partial results rather than queueing indefinitely.
Each node runs two admission lanes:
- FAST lane — regular searches, capped by
node.search.max_concurrent(default 64). - HEAVY lane — aggregations, deep pagination, very large collect
windows, capped by
node.search.heavy_concurrency(default 8).
On the serving side, a per-shard cap (node.search.max_per_shard,
default 16) bounds searches in flight against one shard copy — queue
plus execution — protecting shards remote coordinators cannot see.
All three caps are runtime-tunable via Node Settings.
Timeouts and partial results #
Requests may carry a timeout; without one, the node’s cooperative
deadline applies (node.search.default_timeout_ms, default 30s). The
engine’s collector loops check the deadline as they work and return
what they have with "timed_out": true — searches degrade to partial
results instead of hanging.
Top-K pruning #
Ranked queries collect only the top size hits per shard, and the
inverted index prunes candidate work with block-max WAND: whole blocks
of postings are skipped when their best possible contribution cannot
enter the current top-K. Term, terms and range queries over inplace
columns are answered by the column itself, with block-level summaries
that skip whole 4 096-slot blocks that cannot contain the queried
range.
Aggregations #
Aggregations run per shard and merge on the coordinator. Percentile
estimation uses TDigest by default (HDR histogram quantization
optional). Bucket aggregations nest sub-aggregations under aggs; see
the
aggregation reference.
Vector search #
dense_vector / sparse_vector fields support kNN search — HNSW
graph traversal with a num_candidates pool (default k × 2), or
brute force where appropriate; min_score switches to threshold
collection (all vectors above a similarity floor). Multi-field or
mixed dense/sparse queries fuse rankings with weighted sum or Reciprocal
Rank Fusion (see
multi_vector).
Quantized indexes are selected per field in the schema
(vector_options.index_type): BBQ binary quantization (~32x
compression) and TurboQuant scalar quantization (2–4 bit, ~8x at
4-bit, zero training). Quantized scans produce the candidate pool and
full-precision vectors rescore the top entries, keeping recall close
to exact search at a fraction of the scan cost. Both the quantized
codes and the flat vectors persist in the segment file, and segments
carrying vector data always merge in the FireV11 format so the
quantized sections survive compaction (see
vector query). The
turbo_quant_hnsw variant additionally persists an HNSW graph over
the flat vectors in the same segment; once the segment is memory
mapped, queries traverse the graph instead of scanning the codes —
distances stay exact (flat vectors), so scores match the brute-force
scan while only ef_search-bounded neighborhoods are evaluated.
hnsw_flat carries the same graph without quantized codes. Graph
shape and query beam width are schema-level parameters (m,
ef_construction, ef_search), overridable per query.
Syntax tolerance #
Request bodies are accepted in both Pizza’s canonical DSL and common Elasticsearch forms — the normalization is idempotent and happens before parsing (see API conventions).