Model features & the /_chat facade #
Every model in the AI-services registry can now declare what it is
for and what it accepts — features (embedding or
generation) and inputs (text or image). The two axes are
orthogonal: “multimodal” is not a category of its own but an inputs
list beyond text.
node:
embedding:
endpoints:
- name: zai
url: https://api.z.ai/api/coding/paas/v4/embeddings
chat_url: https://api.z.ai/api/coding/paas/v4/chat/completions
api_key: sk-…
models:
- embedding-3 # embedding + text (the defaults)
- {id: GLM-4.6V, dims: 3072} # embedding, dims declared
- {id: glm-4-flash, features: [generation]} # text generation
- {id: glm-4v, features: [generation], inputs: [text, image]} # vision LLM
- {id: clip-v2, features: [embedding], inputs: [text, image]} # multimodal embedder
Omitted fields keep the historical defaults, so every existing plain-string listing still means embedding + text — nothing to migrate.
The features are enforced where they matter, not just displayed:
- Semantic bindings (
semantic_text, semantic multi-fields, source-mappeddense_vector) reject a generation-only listing at mapping time (“not an embedding model”) instead of failing at write time, and the console’s model picker lists embedding models only. - Generation gets its own facade, the sibling of
/_embedding:
POST /_chat
{
"model": "glm-4-flash",
"messages": [{"role": "user", "content": "hello"}],
"temperature": 0.2
}
The body is the standard chat-completions shape (plus an optional
service disambiguator that pizza strips); it routes to the service
listing the model with the generation feature, and the provider’s
answer — choices, usage, provider extras — passes through verbatim.
The management API follows suit: services carry a chat_url, and the
connectivity probe tests a model the way it serves (embedding models
answer their dimensionality, generation models answer a reply excerpt).