Semantic query #
Semantic search with zero field knowledge: the vector field and the embedding model resolve from the collection schema, and the request is plain text.
POST /my-collection/_search
{"query": {"semantic": "find documents about search engines"}}
The gateway embeds the text server-side through the field’s configured
model (the same inference path as text
vector queries)
and rewrites the clause into a native vector query — so semantic
composes inside bool like any other leaf:
POST /my-collection/_search
{
"query": {
"bool": {
"should": [
{ "semantic": "search engines" },
{ "match": { "field": "title", "query": "engines" } }
]
}
}
}
Enable semantic search #
Three steps, server config first:
1. Configure an embedding endpoint — the server embeds through
external OpenAI-compatible services (OpenAI, vLLM, Ollama, TEI, …),
declared in pizza.yml under
node.embedding:
node:
embedding:
endpoints:
- name: openai
url: https://api.openai.com/v1/embeddings
api_key: sk-…
models: ["text-embedding-3-small"]
timeout_ms: 30000
max_batch_texts: 64
A model id routes to the endpoint that lists it as an embedding model;
with exactly one endpoint configured, every model routes there
(default_endpoint overrides) — that lenient routing serves ad-hoc
/_embedding calls. Schema declarations are stricter: the model
must be listed as an embedding model by exactly one service, or the
mapping names service explicitly (see
AI Services).
Endpoints behind self-signed TLS set
insecure_skip_verify: true (per endpoint). The registry can also be
re-managed at runtime — see
AI Services
— and every node.embedding key is in the
configuration reference.
2. Map a semantic field — the simplest form is the ES-style
semantic_text
type, which collapses field + wiring into one declaration (pure
semantic — the field is not inverted-indexed):
PUT /my-collection
{
"schema": {
"properties": {
"body": {
"type": "semantic_text",
"service": "openai",
"model": "text-embedding-3-small",
"dims": 1536
}
}
}
}
Two explicit forms count as semantic fields too — a dense_vector
sub-field on a top-level text/keyword field (the parent is the
source, and stays BM25-searchable — hybrid out of the box), or a
standalone dense_vector whose vector_options declare source
(the document’s text field to embed) and model:
PUT /my-collection
{
"schema": {
"properties": {
"title": { "type": "text" },
"embedding": {
"type": "dense_vector",
"vector_options": {
"dims": 768,
"similarity": "cosine",
"source": "title",
"model": "text-embedding-3-small"
}
}
}
}
}
3. Write and search — documents indexed into the collection get
their embedding derived automatically at write time for the field’s
model (you don’t supply vectors); the semantic query embeds its
text through the same path at search time. The
/_embedding API exposes the inference
directly.
Wire forms #
String shorthand — just the text:
{ "semantic": "find documents about search engines" }
Object form — any subset of the parameters:
{ "semantic": { "query": "search engines", "k": 10, "similarity": "cosine" } }
Parameters for semantic
#
query
(Required, string) The natural-language text to embed and search with.field
(Optional, string) The semantic field to search. Resolved from the schema when omitted: with exactly one semantic field it is the default; with several, the request must name one (a400lists the candidates otherwise).k
(Optional, integer, default:10) Nearest neighbors to return.num_candidates
(Optional, integer) Graph beam width for the query (HNSWef); see vector.similarity
(Optional, string)cosine(default),l2ordot_product— must match the field schema.min_score
(Optional, float) Return everything above this similarity instead of a fixed top-k — useful when the semantic clause feeds a hybrid intersection.
Notes #
- The search must target exactly one collection — the schema is
where the field and model live; multi-collection targets answer
400. - A collection whose schema declares no semantic field answers
400with instructions to map asemantic_textfield (or adense_vectorwithvector_options.sourceandmodel). - The query never names a service or model — the binding lives in the mapping, and query-time inference uses the same one (writes and searches stay consistent).
- The full vector parameter surface (
vector_optionsoverrides, quantization, sparse vectors) is on the vector query page.