semantic_text — one-field semantic search #
Semantic search used to mean wiring two fields: a text field for the
source, a dense_vector field with vector_options pointing back at
it. The ES-style semantic_text type (Elasticsearch 8.15+ alignment)
collapses all of that into one declaration:
PUT /my-collection
{
"schema": {
"properties": {
"body": {
"type": "semantic_text",
"service": "zai",
"model": "GLM-4.6V",
"dims": 3072
}
}
}
}
From there everything derives. Writes carry plain text — the vector is embedded server-side at write time, in bulk and after partial updates alike. Search needs no field names and no vectors:
POST /my-collection/_search
{"query": {"semantic": "documents about search engines"}}
# hybrid without the retriever DSL
{"query": {"bool": {"should": [
{"semantic": "search engines"},
{"match": {"field": "title", "query": "engines"}}
]}}}
The field is pure semantic, like its ES namesake: the text stays in
_source but is never inverted-indexed — match does not hit it, the
vectors do. Keep a sibling text field for lexical legs.
The binding to the embedding backend is strict, so a mapping can never silently point at the wrong service:
modelmust be listed by an AI service — exactly one, or the mapping namesserviceexplicitly. An unlisted model, or one listed by several services, is a400at mapping time, with the candidates named.dimsmay be omitted when the service declares it on itsmodelsentry ({"id": "GLM-4.6V", "dims": 3072}); a conflicting hand-written value is caught at mapping time, and a backend answering with a different width fails the write loudly — no silently dropped vectors.
GET /_schema presents the field back as semantic_text with its
binding, so the declaration round-trips.
Honest gap: no chunking yet — the whole field text embeds in one call (fine for titles, names and log lines; long documents are bounded by the model’s input limit). The engine keeps one vector per field, so ES-style chunking is on the roadmap; until then, chunk into child documents and collapse by parent id.
See Vector types, the semantic query and AI Services for the full contract.