AI services are the external, OpenAI-compatible model endpoints behind
server-side inference — schema-driven text→vector derivation, text
vector queries, the semantic query and the /_embedding API on the
embedding side, and the POST /_chat generation facade on the other.
Every service speaks the OpenAI wire format: the embeddings API
(POST {url} with {"model", "input": [...]} and a bearer key), and —
once it lists generation models — the chat-completions API at its
chat_url.
The registry has two layers:
- the config baseline —
node.embeddinginpizza.yml, read at boot ( keys); - the runtime layer — this API (or the console’s AI Services page,
/_ui/), which takes over on its first modification and persists to<path.data>/ai_services.json(0600, atomically replaced).
The layer is node-local by design: API keys never enter cluster
metadata or raft logs, and never come back from the API — GET answers
with a masked hint only. Inference resolves an effective view
(source: "dynamic" once the runtime layer exists, "config" while
the yml is authoritative), with three resolution modes:
- Ad-hoc (
POST /_embedding): a model id routes to the service listing it as an embedding model, then the default endpoint, then the single configured service. - Schema-bound (
semantic_textfields, semantic multi-fields, source-mappeddense_vector): strict — the model must be listed as an embedding model by exactly one service, or the mapping must name the service explicitly ("service": "…", accepted onsemantic_textand insidevector_options). Unlisted → 400 “listed by no configured service”; a generation-only listing → 400 “not an embedding model”; listed by several → 400 naming the candidates and asking forservice. - Generation (
POST /_chat, see Chat): strict — the model must be listed with thegenerationfeature by exactly one service (or the request namesservice), and that service must declarechat_url.
Each models entry may declare its output dimensionality, its features
and its input modalities — "text-embedding-3-small" or
{"id": "text-embedding-3-small", "dims": 1536} or
{"id": "llava", "features": ["generation"], "inputs": ["text", "image"]}. A declared dims lets schema mappings omit dims and
catches hand-written mismatches at mapping time instead of dropping
vectors at write time.
features says what the model is for: embedding (the default —
everything before the field existed embeds) or generation (routable
through /_chat). inputs says what it accepts: text (default)
or image. The two are orthogonal — “multimodal” is not a category of
its own but an inputs list beyond text: a CLIP-style embedder is
{"features": ["embedding"], "inputs": ["text", "image"]}, a vision
LLM is {"features": ["generation"], "inputs": ["text", "image"]}.
Omitted fields keep the historical defaults, so existing plain-string
listings stay embedding + text.
Get AI Services #
Returns the effective registry on this node.
Requests #
GET /_node/_local/ai_services
A peer node is read through the unified forward:
GET /_node/<node_id>/ai_services
Response #
{
"services": [
{
"name": "openai",
"url": "https://api.openai.com/v1/embeddings",
"chat_url": "https://api.openai.com/v1/chat/completions",
"models": ["text-embedding-3-small", {"id": "gpt-4o-mini", "features": ["generation"]}],
"api_key_set": true,
"api_key_hint": "sk-…4f2a",
"is_default": false,
"insecure_skip_verify": false
},
{
"name": "local",
"url": "http://127.0.0.1:11434/v1/embeddings",
"chat_url": null,
"models": ["nomic-embed-text"],
"api_key_set": false,
"api_key_hint": "",
"is_default": true,
"insecure_skip_verify": false
}
],
"source": "dynamic",
"settings": {
"default_endpoint": "local",
"timeout_ms": 30000,
"max_batch_texts": 64
}
}
api_key_set/api_key_hint
Whether a key is stored, and a masked fragment (sk-…4f2a) — enough to recognize, impossible to reuse. Empty hint = no key.chat_url
The chat-completions endpoint declared for generation models;nullwhen the service is embeddings-only.models
Entries stay minimal on the wire: a plain string while the model keeps the defaults (embedding, text-only), an object as soon asdims,featuresorinputsare declared.is_default
The service serving model ids nobody lists: the explicit default endpoint, or the single configured service.sourceconfigwhilepizza.ymlis authoritative;dynamicafter the first runtime modification.settings
The effective registry-wide knobs (see Update settings below).
Upsert a service #
Creates or updates one service — the switch to the runtime layer happens here, seeded from the config baseline.
Requests #
PUT /_node/_local/ai_services/services/<name>
{
"url": "https://api.z.ai/api/coding/paas/v4/embeddings",
"chat_url": "https://api.z.ai/api/coding/paas/v4/chat/completions",
"models": ["GLM-4.6V", {"id": "embedding-3", "dims": 1536}, {"id": "glm-4-flash", "features": ["generation"]}],
"api_key": "…",
"insecure_skip_verify": false
}
Request body #
url
(Required, string) The full embeddings endpoint URL; must start withhttp://orhttps://.chat_url
(Optional, string) The full chat-completions endpoint URL — required once the service lists models with"features": ["generation"](that listing is rejected with400otherwise); omitted ornullclears it. Must be http(s).models
(Optional, array) Model ids routed here, each either a plain string or an object{"id": …, "dims": …, "features": […], "inputs": […]}declaring the model’s output dimensionality, features and input modalities (see the intro; omitted fields default toembedding+text); the request’smodelpasses through verbatim. Omitted = empty.dimsis only valid on models with theembeddingfeature.api_key
(Optional, string) Write-only. Omitted ornullkeeps the stored key — edit a service without re-entering the secret;""clears it; any other string replaces it.insecure_skip_verify
(Optional, boolean, defaultfalse) Skip TLS chain validation for this endpoint (self-signed certificates).
Errors: 400 for a missing/non-string url, a non-array or malformed
models, a non-http(s) URL, dims on a model without the embedding
feature, or generation models without a chat_url. The response is
the updated registry view.
Delete a service #
Requests #
DELETE /_node/_local/ai_services/services/<name>
404 when the name is unknown. Deleting the default service clears
the default. Deleting from a pure-config setup seeds the runtime layer
first, so the removal persists instead of being masked by the yml
fallback. The response is the updated registry view.
Update settings #
Registry-wide knobs — any subset of the body, at least one key.
Requests #
PUT /_node/_local/ai_services/settings
{
"default_endpoint": "local",
"timeout_ms": 30000,
"max_batch_texts": 64
}
Request body #
default_endpoint
(Optional, string | null) Service name serving model ids no service lists.nullclears it (single-service setups need none). Must name an existing service —400otherwise.timeout_ms
(Optional, integer) Per-request timeout for embedding calls; values below100are raised to100.max_batch_texts
(Optional, integer) Texts per embedding request — larger write batches split; clamped to1..=1024.
Unset keys fall through to the config-file values. The response is the updated registry view.
Test a service #
Runs one live round-trip through the service, with its stored key and
TLS settings — routed by the probed model’s feature: an embedding
model embeds "ping" and answers the vector dimensionality; a
generation model completes "ping" and answers a reply excerpt.
Requests #
POST /_node/_local/ai_services/services/<name>/_test
{ "model": "GLM-4.6V" }
model
(Optional, string) The model id to test with, probed per its declared feature. Defaults to the service’s first embedding model, falling back to its first generation model.400when the named model is not listed by the service, or the service lists none.
Response #
{
"service": "GLM",
"model": "GLM-4.6V",
"feature": "embedding",
"dims": 3072,
"latency_ms": 812,
"ok": true
}
feature
Which path was probed —embeddingorgeneration.dims
(embedding probe) The returned vector’s dimension — the number to use indense_vectorfield declarations, or to declare once on the service’smodelsentry ({"id": …, "dims": …}) so schema mappings can omit it.reply
(generation probe) A short excerpt of the model’s first choice — proof it actually generated.
Failures surface the upstream error (connection refused, 401/403, 429, …).
Chat (generation facade) #
POST /_chat is the generation sibling of POST /_embedding: the
standard chat-completions body, routed through the AI-services registry
— one stable URL while pizza fans out to whichever service lists the
model with the generation feature.
Requests #
POST /_chat
{
"model": "gpt-4o-mini",
"service": "openai",
"messages": [{"role": "user", "content": "hello"}],
"temperature": 0.2
}
model
(Required, string) The generation model — must be listed withfeatures: ["generation"]by exactly one service, or the request must nameservice;400otherwise (“not a generation model” for an embedding-only listing, the candidate names for an ambiguity).service
(Optional, string) The AI service to route through — the registry disambiguator, stripped before the call; never forwarded upstream.- everything else
messages(required, non-empty) and any provider options —temperature,max_tokens,response_format, … — pass through verbatim, and the provider’s answer returns verbatim:choices,usageand provider extras all survive.
"stream": true is rejected with 400 — the facade is non-streaming;
call the provider directly for streaming. The registry’s timeout_ms
bounds the call. 502 when the endpoint is unreachable or answers
non-2xx.
Reset to the config file #
Drops the runtime layer and its persisted file — pizza.yml becomes
authoritative again, without a restart.
Requests #
DELETE /_node/_local/ai_services
The response is the post-reset registry view (source: "config").
Remote nodes and the console #
Reads and writes on a peer node go through the unified forward
(GET/PUT/POST /_node/<node_id>/ai_services…); the two
DELETEs are local-only — the console reaches each node’s API address
directly over CORS for them. Every node manages its own services; a
cluster-wide change is a fan-out (the console does this for you).