Vector types

Vector types #

semantic_text #

The ES-style field type for semantic search (Elasticsearch 8.15+ alignment): declare one field, writes embed it server-side, and the semantic query searches it with plain text — no field names, no manual vectors.

PUT /my-collection
{
  "schema": {
    "properties": {
      "body": {
        "type": "semantic_text",
        "service": "zai",
        "model": "GLM-4.6V",
        "dims": 3072,
        "similarity": "cosine"
      }
    }
  }
}

Documents carry plain text ({"body": "some text"}); the vector derives at write time into the internal body.embedding sub-field. The field is pure semantic (ES semantic_text alignment): the text is kept in _source but is not inverted-indexed — match/term queries do not hit it; search goes through the derived vectors. A sibling regular text field plus bool.should composition gives hybrid retrieval.

Parameters #

  • model
    (Required, string) The embedding model. Must be listed by exactly one AI service — see binding below.
  • service
    (Optional, string) The AI service name serving model, for when several services list the same model id. Resolution: with service, the model must be listed by that service; without it, exactly one service must list the model (0 → 400 “listed by no service”, 2+ → 400 naming the candidates and asking for service). Schema bindings are strict — the ad-hoc default-endpoint fallback of POST /_embedding does not apply.
  • dims
    (Optional, integer) The embedding dimensionality. Required unless the service’s models entry declares it ({"id": "…", "dims": 3072}); a value conflicting with the declared dims is a 400 at mapping time, and a backend answering with a different width fails the write loudly.
  • similarity
    (Optional, string, default cosine) cosine | l2 | dot_product.

Top-level fields only; the type is rejected in multi-field or sub-object positions, and accepts no layout parameters (the expansion fixes them). GET /<target>/_schema presents the field back as semantic_text with its binding — the declaration round-trips.

Honest gap: no chunking yet — the whole field text is embedded in one call, so very long texts are bounded by the service’s input limit.

Support indexing multiple values #

No — one text, one derived vector per field.

dense_vector #

An array of floats with a fixed dimension, for kNN vector search. Requires vector_options (dims, similarity, index_type — flat, HNSW, BBQ or TurboQuant variants). See the vector search page for the full parameter catalog.

Parameters for dense_vector fields #

Support indexing multiple values #

No — one vector per field.

sparse_vector #

A sparse vector as a map of index to weight, for sparse kNN search over models such as SPLADE. Shares the query surface of dense_vector; see sparse vectors.

Parameters for sparse_vector fields #

Support indexing multiple values #

No

Calendar September 29, 2026
Edit Edit this page