Graph projection block #
Add a graph block to any search request and the response carries a
graph — nodes and edges — built from the hit set, alongside the
normal hits. Where the
traverse query searches
by walking edges, the graph block renders: it takes the documents
a query already found and projects them into a node/edge structure a
UI (or a GraphRAG pipeline) can consume directly.
Live over the articles teaching dataset (44 documents with a
related_to relation field) — a term query picks the seeds, the graph
block renders everything they reach:
And the ad-hoc form over a plain keyword field — no relation schema,
every tags value on the hit documents becomes an entity node:
POST /articles/_search
{
"size": 20,
"query": { "match_all": {} },
"graph": { "field": "related_to" }
}
Two projection tiers #
The engine picks a tier per relation slice:
- Relation fields (adjacency). When the field is declared
type: relationin the schema (or you pass therelationhint), the projection walks the field’s adjacency index — the same structuretraverseuses — and emits document-to-document edges for every recorded edge whose endpoints are in the hit set. - Ad-hoc column projection. For an ordinary keyword field, every
distinct value on the hit documents becomes a synthetic entity node
and each document links to its values (
doc → ent:<field>:<value>). This turns any faceting-style column into a bipartite graph — no relation schema needed.
Shapes #
Single relation (the shorthand):
{
"query": { "term": { "field": "author", "value": "medcl" } },
"graph": { "field": "genres", "max_values_per_doc": 8 }
}
Multiple relations in one round trip — each slice merges into one
graph; doc nodes are deduplicated across slices, entity nodes stay
per-field. Live: one projection over both the related_to edges and
the tags column:
The equivalent request as JSON:
{
"query": { "match_all": {} },
"graph": {
"relations": [
{ "field": "related_to" },
{ "field": "tags", "max_values_per_doc": 4 }
],
"max_seeds": 100
}
}
Response #
The response gains a graph object as a peer of hits and
aggregations:
{
"hits": { … },
"graph": {
"nodes": [
{ "id": "doc:A01", "kind": "doc", "key": "A01", "degree": 3 },
{ "id": "ent:tags:search", "kind": "ent", "field": "tags", "value": "search" }
],
"edges": [
{ "from": "doc:A01", "to": "doc:A09", "via": "related_to" },
{ "from": "doc:A01", "to": "ent:tags:search", "via": "tags" }
],
"plan": "slices=2 seeds=3 [tier=edge field=related_to reached=10, tier=column field=tags]"
}
}
- Node ids are stable identifiers for consumers to dedupe on:
"doc:<key>"for document nodes,"ent:<field>:<value>"for entity nodes. Internal numeric document ids are never exposed — documents without a_keyrender asdoc:<id>. kindis"doc"or"ent"; doc nodes carry theirkeyand adegree(edges touching the node), entity nodes carry thefieldthey were derived from and its stringifiedvalue.- With a relation field, edge documents appear as doc nodes too — the endpoints they connect link through them, mirroring how the edges are stored.
- Edges carry
from,to(node ids) andvia(the field the edge was derived from). planis a short human-readable summary of what actually executed (tier chosen per slice, field, seed count, reach).
Parameters for graph
#
field
(string) Single-relation shorthand — the field to project. Set this orrelations, not both.relation
(Optional, string) Explicit hint naming a schemarelationfield, routing the slice through the adjacency tier even whenfieldalone would be ambiguous.relations
(array) Multi-relation form — one entry per slice, each withfield, optionalrelationhint andmax_values_per_doc.max_values_per_doc
(Optional, integer, default:16) Cap on distinct values projected per source document (ad-hoc tier) — keeps a long array field from blowing up the node set. In therelationsform, set it per entry.max_seeds
(Optional, integer, default: the request’ssize) Cap on how many hit documents feed the projection — prevents accidental projections over million-hit baselines.seed_kind
(Optional, string, default:"doc") Kind label for document nodes — set it to the collection name (e.g."movie") for more semantic rendering.key_field
(Optional, string)_sourcefield used to key document nodes that are NOT part of the returned hit page (reached only through edges) — the engine looks the value up from the document store so their ids carry a business key instead of the internal ordinal. Nodes in the hit page always use their_key.label_field
(Optional, string)_sourcefield populating each doc node’s human-readablevalue(e.g. the title) — verified live: works for in-page and edge-reached nodes alike.extra_fields
(array of strings) Additional_sourcefields attached to each doc node underextras(e.g.["author"]) so a UI can render metadata without a second round trip.
Notes #
- The block composes with any
query,sort, andsize— the graph is derived from whatever the query selected. - Like
traverse, relation-field projection resolves endpoints within the shard that stores the edge document — keep graph collections on a single shard (see traverse — shard placement).