Graph projection block

Graph projection block #

Add a graph block to any search request and the response carries a graph — nodes and edges — built from the hit set, alongside the normal hits. Where the traverse query searches by walking edges, the graph block renders: it takes the documents a query already found and projects them into a node/edge structure a UI (or a GraphRAG pipeline) can consume directly.

Live over the articles teaching dataset (44 documents with a related_to relation field) — a term query picks the seeds, the graph block renders everything they reach:

And the ad-hoc form over a plain keyword field — no relation schema, every tags value on the hit documents becomes an entity node:

POST /articles/_search
{
  "size": 20,
  "query": { "match_all": {} },
  "graph": { "field": "related_to" }
}

Two projection tiers #

The engine picks a tier per relation slice:

  • Relation fields (adjacency). When the field is declared type: relation in the schema (or you pass the relation hint), the projection walks the field’s adjacency index — the same structure traverse uses — and emits document-to-document edges for every recorded edge whose endpoints are in the hit set.
  • Ad-hoc column projection. For an ordinary keyword field, every distinct value on the hit documents becomes a synthetic entity node and each document links to its values (doc → ent:<field>:<value>). This turns any faceting-style column into a bipartite graph — no relation schema needed.

Shapes #

Single relation (the shorthand):

{
  "query": { "term": { "field": "author", "value": "medcl" } },
  "graph": { "field": "genres", "max_values_per_doc": 8 }
}

Multiple relations in one round trip — each slice merges into one graph; doc nodes are deduplicated across slices, entity nodes stay per-field. Live: one projection over both the related_to edges and the tags column:

The equivalent request as JSON:

{
  "query": { "match_all": {} },
  "graph": {
    "relations": [
      { "field": "related_to" },
      { "field": "tags", "max_values_per_doc": 4 }
    ],
    "max_seeds": 100
  }
}

Response #

The response gains a graph object as a peer of hits and aggregations:

{
  "hits": { … },
  "graph": {
    "nodes": [
      { "id": "doc:A01", "kind": "doc", "key": "A01", "degree": 3 },
      { "id": "ent:tags:search", "kind": "ent", "field": "tags", "value": "search" }
    ],
    "edges": [
      { "from": "doc:A01", "to": "doc:A09", "via": "related_to" },
      { "from": "doc:A01", "to": "ent:tags:search", "via": "tags" }
    ],
    "plan": "slices=2 seeds=3 [tier=edge field=related_to reached=10, tier=column field=tags]"
  }
}
  • Node ids are stable identifiers for consumers to dedupe on: "doc:<key>" for document nodes, "ent:<field>:<value>" for entity nodes. Internal numeric document ids are never exposed — documents without a _key render as doc:<id>.
  • kind is "doc" or "ent"; doc nodes carry their key and a degree (edges touching the node), entity nodes carry the field they were derived from and its stringified value.
  • With a relation field, edge documents appear as doc nodes too — the endpoints they connect link through them, mirroring how the edges are stored.
  • Edges carry from, to (node ids) and via (the field the edge was derived from).
  • plan is a short human-readable summary of what actually executed (tier chosen per slice, field, seed count, reach).

Parameters for graph #

  • field
    (string) Single-relation shorthand — the field to project. Set this or relations, not both.
  • relation
    (Optional, string) Explicit hint naming a schema relation field, routing the slice through the adjacency tier even when field alone would be ambiguous.
  • relations
    (array) Multi-relation form — one entry per slice, each with field, optional relation hint and max_values_per_doc.
  • max_values_per_doc
    (Optional, integer, default: 16) Cap on distinct values projected per source document (ad-hoc tier) — keeps a long array field from blowing up the node set. In the relations form, set it per entry.
  • max_seeds
    (Optional, integer, default: the request’s size) Cap on how many hit documents feed the projection — prevents accidental projections over million-hit baselines.
  • seed_kind
    (Optional, string, default: "doc") Kind label for document nodes — set it to the collection name (e.g. "movie") for more semantic rendering.
  • key_field
    (Optional, string) _source field used to key document nodes that are NOT part of the returned hit page (reached only through edges) — the engine looks the value up from the document store so their ids carry a business key instead of the internal ordinal. Nodes in the hit page always use their _key.
  • label_field
    (Optional, string) _source field populating each doc node’s human-readable value (e.g. the title) — verified live: works for in-page and edge-reached nodes alike.
  • extra_fields
    (array of strings) Additional _source fields attached to each doc node under extras (e.g. ["author"]) so a UI can render metadata without a second round trip.

Notes #

  • The block composes with any query, sort, and size — the graph is derived from whatever the query selected.
  • Like traverse, relation-field projection resolves endpoints within the shard that stores the edge document — keep graph collections on a single shard (see traverse — shard placement).
Calendar September 30, 2026
Edit Edit this page