Vector types #
semantic_text
#
The ES-style field type for semantic search (Elasticsearch 8.15+
alignment): declare one field, writes embed it server-side, and the
semantic query searches
it with plain text — no field names, no manual vectors.
PUT /my-collection
{
"schema": {
"properties": {
"body": {
"type": "semantic_text",
"service": "zai",
"model": "GLM-4.6V",
"dims": 3072,
"similarity": "cosine"
}
}
}
}
Documents carry plain text ({"body": "some text"}); the vector derives
at write time into the internal body.embedding sub-field. The field is
pure semantic (ES semantic_text alignment): the text is kept in
_source but is not inverted-indexed — match/term queries do not
hit it; search goes through the derived vectors. A sibling regular text
field plus bool.should composition gives hybrid retrieval.
Parameters #
model
(Required, string) The embedding model. Must be listed by exactly one AI service — see binding below.service
(Optional, string) The AI service name servingmodel, for when several services list the same model id. Resolution: withservice, the model must be listed by that service; without it, exactly one service must list the model (0 → 400 “listed by no service”, 2+ → 400 naming the candidates and asking forservice). Schema bindings are strict — the ad-hoc default-endpoint fallback ofPOST /_embeddingdoes not apply.dims
(Optional, integer) The embedding dimensionality. Required unless the service’smodelsentry declares it ({"id": "…", "dims": 3072}); a value conflicting with the declared dims is a400at mapping time, and a backend answering with a different width fails the write loudly.similarity
(Optional, string, defaultcosine)cosine|l2|dot_product.
Top-level fields only; the type is rejected in multi-field or sub-object
positions, and accepts no layout parameters (the expansion fixes them).
GET /<target>/_schema presents the field back as semantic_text with
its binding — the declaration round-trips.
Honest gap: no chunking yet — the whole field text is embedded in one call, so very long texts are bounded by the service’s input limit.
Support indexing multiple values #
No — one text, one derived vector per field.
dense_vector
#
An array of floats with a fixed dimension, for kNN vector search. Requires
vector_options
(dims, similarity, index_type — flat, HNSW, BBQ or TurboQuant
variants). See the
vector search page for
the full parameter catalog.
Parameters for dense_vector fields
#
- vector_options (dims, similarity, index_type, m, ef_construction, ef_search, source, model, …)
- realtime
- fields (semantic multi-field — the declarative embedding form, see the vector page)
- profile
- source / column / store
Support indexing multiple values #
No — one vector per field.
sparse_vector
#
A sparse vector as a map of index to weight, for sparse kNN search over
models such as SPLADE. Shares the query surface of dense_vector; see
sparse vectors.
Parameters for sparse_vector fields
#
- vector_options (similarity, …)
- realtime
- profile
- source / column / store
Support indexing multiple values #
No