AI Services

AI services are the external, OpenAI-compatible model endpoints behind server-side inference — schema-driven text→vector derivation, text vector queries, the semantic query and the /_embedding API on the embedding side, and the POST /_chat generation facade on the other. Every service speaks the OpenAI wire format: the embeddings API (POST {url} with {"model", "input": [...]} and a bearer key), and — once it lists generation models — the chat-completions API at its chat_url.

The registry has two layers:

  • the config baseline — node.embedding in pizza.yml, read at boot ( keys);
  • the runtime layer — this API (or the console’s AI Services page, /_ui/), which takes over on its first modification and persists to <path.data>/ai_services.json (0600, atomically replaced).

The layer is node-local by design: API keys never enter cluster metadata or raft logs, and never come back from the API — GET answers with a masked hint only. Inference resolves an effective view (source: "dynamic" once the runtime layer exists, "config" while the yml is authoritative), with three resolution modes:

  • Ad-hoc (POST /_embedding): a model id routes to the service listing it as an embedding model, then the default endpoint, then the single configured service.
  • Schema-bound (semantic_text fields, semantic multi-fields, source-mapped dense_vector): strict — the model must be listed as an embedding model by exactly one service, or the mapping must name the service explicitly ("service": "…", accepted on semantic_text and inside vector_options). Unlisted → 400 “listed by no configured service”; a generation-only listing → 400 “not an embedding model”; listed by several → 400 naming the candidates and asking for service.
  • Generation (POST /_chat, see Chat): strict — the model must be listed with the generation feature by exactly one service (or the request names service), and that service must declare chat_url.

Each models entry may declare its output dimensionality, its features and its input modalities — "text-embedding-3-small" or {"id": "text-embedding-3-small", "dims": 1536} or {"id": "llava", "features": ["generation"], "inputs": ["text", "image"]}. A declared dims lets schema mappings omit dims and catches hand-written mismatches at mapping time instead of dropping vectors at write time.

features says what the model is for: embedding (the default — everything before the field existed embeds) or generation (routable through /_chat). inputs says what it accepts: text (default) or image. The two are orthogonal — “multimodal” is not a category of its own but an inputs list beyond text: a CLIP-style embedder is {"features": ["embedding"], "inputs": ["text", "image"]}, a vision LLM is {"features": ["generation"], "inputs": ["text", "image"]}. Omitted fields keep the historical defaults, so existing plain-string listings stay embedding + text.

Get AI Services #

Returns the effective registry on this node.

Requests #

GET /_node/_local/ai_services

A peer node is read through the unified forward:

GET /_node/<node_id>/ai_services

Response #

{
  "services": [
    {
      "name": "openai",
      "url": "https://api.openai.com/v1/embeddings",
      "chat_url": "https://api.openai.com/v1/chat/completions",
      "models": ["text-embedding-3-small", {"id": "gpt-4o-mini", "features": ["generation"]}],
      "api_key_set": true,
      "api_key_hint": "sk-…4f2a",
      "is_default": false,
      "insecure_skip_verify": false
    },
    {
      "name": "local",
      "url": "http://127.0.0.1:11434/v1/embeddings",
      "chat_url": null,
      "models": ["nomic-embed-text"],
      "api_key_set": false,
      "api_key_hint": "",
      "is_default": true,
      "insecure_skip_verify": false
    }
  ],
  "source": "dynamic",
  "settings": {
    "default_endpoint": "local",
    "timeout_ms": 30000,
    "max_batch_texts": 64
  }
}
  • api_key_set / api_key_hint
    Whether a key is stored, and a masked fragment (sk-…4f2a) — enough to recognize, impossible to reuse. Empty hint = no key.
  • chat_url
    The chat-completions endpoint declared for generation models; null when the service is embeddings-only.
  • models
    Entries stay minimal on the wire: a plain string while the model keeps the defaults (embedding, text-only), an object as soon as dims, features or inputs are declared.
  • is_default
    The service serving model ids nobody lists: the explicit default endpoint, or the single configured service.
  • source
    config while pizza.yml is authoritative; dynamic after the first runtime modification.
  • settings
    The effective registry-wide knobs (see Update settings below).

Upsert a service #

Creates or updates one service — the switch to the runtime layer happens here, seeded from the config baseline.

Requests #

PUT /_node/_local/ai_services/services/<name>
{
  "url": "https://api.z.ai/api/coding/paas/v4/embeddings",
  "chat_url": "https://api.z.ai/api/coding/paas/v4/chat/completions",
  "models": ["GLM-4.6V", {"id": "embedding-3", "dims": 1536}, {"id": "glm-4-flash", "features": ["generation"]}],
  "api_key": "…",
  "insecure_skip_verify": false
}

Request body #

  • url
    (Required, string) The full embeddings endpoint URL; must start with http:// or https://.
  • chat_url
    (Optional, string) The full chat-completions endpoint URL — required once the service lists models with "features": ["generation"] (that listing is rejected with 400 otherwise); omitted or null clears it. Must be http(s).
  • models
    (Optional, array) Model ids routed here, each either a plain string or an object {"id": …, "dims": …, "features": […], "inputs": […]} declaring the model’s output dimensionality, features and input modalities (see the intro; omitted fields default to embedding + text); the request’s model passes through verbatim. Omitted = empty. dims is only valid on models with the embedding feature.
  • api_key
    (Optional, string) Write-only. Omitted or null keeps the stored key — edit a service without re-entering the secret; "" clears it; any other string replaces it.
  • insecure_skip_verify
    (Optional, boolean, default false) Skip TLS chain validation for this endpoint (self-signed certificates).

Errors: 400 for a missing/non-string url, a non-array or malformed models, a non-http(s) URL, dims on a model without the embedding feature, or generation models without a chat_url. The response is the updated registry view.

Delete a service #

Requests #

DELETE /_node/_local/ai_services/services/<name>

404 when the name is unknown. Deleting the default service clears the default. Deleting from a pure-config setup seeds the runtime layer first, so the removal persists instead of being masked by the yml fallback. The response is the updated registry view.

Update settings #

Registry-wide knobs — any subset of the body, at least one key.

Requests #

PUT /_node/_local/ai_services/settings
{
  "default_endpoint": "local",
  "timeout_ms": 30000,
  "max_batch_texts": 64
}

Request body #

  • default_endpoint
    (Optional, string | null) Service name serving model ids no service lists. null clears it (single-service setups need none). Must name an existing service — 400 otherwise.
  • timeout_ms
    (Optional, integer) Per-request timeout for embedding calls; values below 100 are raised to 100.
  • max_batch_texts
    (Optional, integer) Texts per embedding request — larger write batches split; clamped to 1..=1024.

Unset keys fall through to the config-file values. The response is the updated registry view.

Test a service #

Runs one live round-trip through the service, with its stored key and TLS settings — routed by the probed model’s feature: an embedding model embeds "ping" and answers the vector dimensionality; a generation model completes "ping" and answers a reply excerpt.

Requests #

POST /_node/_local/ai_services/services/<name>/_test
{ "model": "GLM-4.6V" }
  • model
    (Optional, string) The model id to test with, probed per its declared feature. Defaults to the service’s first embedding model, falling back to its first generation model. 400 when the named model is not listed by the service, or the service lists none.

Response #

{
  "service": "GLM",
  "model": "GLM-4.6V",
  "feature": "embedding",
  "dims": 3072,
  "latency_ms": 812,
  "ok": true
}
  • feature
    Which path was probed — embedding or generation.
  • dims
    (embedding probe) The returned vector’s dimension — the number to use in dense_vector field declarations, or to declare once on the service’s models entry ({"id": …, "dims": …}) so schema mappings can omit it.
  • reply
    (generation probe) A short excerpt of the model’s first choice — proof it actually generated.

Failures surface the upstream error (connection refused, 401/403, 429, …).

Chat (generation facade) #

POST /_chat is the generation sibling of POST /_embedding: the standard chat-completions body, routed through the AI-services registry — one stable URL while pizza fans out to whichever service lists the model with the generation feature.

Requests #

POST /_chat
{
  "model": "gpt-4o-mini",
  "service": "openai",
  "messages": [{"role": "user", "content": "hello"}],
  "temperature": 0.2
}
  • model
    (Required, string) The generation model — must be listed with features: ["generation"] by exactly one service, or the request must name service; 400 otherwise (“not a generation model” for an embedding-only listing, the candidate names for an ambiguity).
  • service
    (Optional, string) The AI service to route through — the registry disambiguator, stripped before the call; never forwarded upstream.
  • everything else
    messages (required, non-empty) and any provider options — temperature, max_tokens, response_format, … — pass through verbatim, and the provider’s answer returns verbatim: choices, usage and provider extras all survive.

"stream": true is rejected with 400 — the facade is non-streaming; call the provider directly for streaming. The registry’s timeout_ms bounds the call. 502 when the endpoint is unreachable or answers non-2xx.

Reset to the config file #

Drops the runtime layer and its persisted file — pizza.yml becomes authoritative again, without a restart.

Requests #

DELETE /_node/_local/ai_services

The response is the post-reset registry view (source: "config").

Remote nodes and the console #

Reads and writes on a peer node go through the unified forward (GET/PUT/POST /_node/<node_id>/ai_services…); the two DELETEs are local-only — the console reaches each node’s API address directly over CORS for them. Every node manages its own services; a cluster-wide change is a fan-out (the console does this for you).

Calendar September 29, 2026
Edit Edit this page