semantic_text — one-field semantic search

semantic_text — one-field semantic search #

Semantic search used to mean wiring two fields: a text field for the source, a dense_vector field with vector_options pointing back at it. The ES-style semantic_text type (Elasticsearch 8.15+ alignment) collapses all of that into one declaration:

PUT /my-collection
{
  "schema": {
    "properties": {
      "body": {
        "type": "semantic_text",
        "service": "zai",
        "model": "GLM-4.6V",
        "dims": 3072
      }
    }
  }
}

From there everything derives. Writes carry plain text — the vector is embedded server-side at write time, in bulk and after partial updates alike. Search needs no field names and no vectors:

POST /my-collection/_search
{"query": {"semantic": "documents about search engines"}}

# hybrid without the retriever DSL
{"query": {"bool": {"should": [
  {"semantic": "search engines"},
  {"match": {"field": "title", "query": "engines"}}
]}}}

The field is pure semantic, like its ES namesake: the text stays in _source but is never inverted-indexed — match does not hit it, the vectors do. Keep a sibling text field for lexical legs.

The binding to the embedding backend is strict, so a mapping can never silently point at the wrong service:

  • model must be listed by an AI service — exactly one, or the mapping names service explicitly. An unlisted model, or one listed by several services, is a 400 at mapping time, with the candidates named.
  • dims may be omitted when the service declares it on its models entry ({"id": "GLM-4.6V", "dims": 3072}); a conflicting hand-written value is caught at mapping time, and a backend answering with a different width fails the write loudly — no silently dropped vectors.

GET /_schema presents the field back as semantic_text with its binding, so the declaration round-trips.

Honest gap: no chunking yet — the whole field text embeds in one call (fine for titles, names and log lines; long documents are bounded by the model’s input limit). The engine keeps one vector per field, so ES-style chunking is on the roadmap; until then, chunk into child documents and collapse by parent id.

See Vector types, the semantic query and AI Services for the full contract.

Calendar September 29, 2026
Edit Edit this page