Aggregation #
An aggregation summarizes your data as metrics, statistics, or other
analytics. Aggregations nest: any bucket aggregation accepts aggs
sub-aggregations computed per bucket.
Pizza organizes aggregations into three categories:
- Metric aggregations compute values (a sum, a percentile, a centroid) from field values.
- Bucket aggregations group documents into buckets keyed by field values, ranges, filters, tiles, …
- Pipeline aggregations consume the output of other aggregations — 17 dedicated pages (running totals, derivatives, window statistics, per-bucket arithmetic).
Metric aggregations #
Single-value metrics:
avg— average of numeric values.cardinality— count of distinct values.max/min— extremes of numeric values.sum— sum of numeric values.value_count— count of extracted values.median_absolute_deviation— robust spread of numeric values.weighted_avg— average weighted by another field.
Multi-value statistics:
stats— count, min, max, avg, sum in one.extended_stats— stats plus variance, standard deviation and bounds.percentiles— values at given percentiles.percentile_ranks— percentage of values at or below given values.boxplot— five-number summary (min/q1/median/q3/max).string_stats— length and entropy stats for keyword fields.matrix_stats— correlation matrix across numeric fields.t_test— paired t-test between two numeric fields.rate— per-interval totals for counters.
Geospatial metrics:
geo_bounds— bounding box of geo points.geo_centroid— centroid of geo points.
Document pickers:
top_hits— the most relevant documents per bucket.top_metrics— field values of the top documents.
Bucket aggregations #
By field value:
terms— one bucket per unique value.multi_terms— buckets keyed by value tuples.rare_terms— only values below a count threshold.significant_terms— values unusually frequent in scope vs. background.
By numeric/date range:
range— user-defined numeric ranges.date_range— user-defined date ranges.histogram— fixed-interval numeric buckets.variable_width_histogram— equal-count adaptive buckets.date_histogram— calendar buckets (day, week, month …).auto_date_histogram— you pick the bucket count, the engine picks the interval.ip_range— CIDR / IP ranges on ip fields.
By query:
filter— one bucket narrowed by a query.filters— named filters, one bucket each.adjacency_matrix— filters plus their pairwise intersections.global— all documents, ignoring the query scope.missing— documents lacking a field value.
Sampling:
sampler— sub-aggs over a document sample.diversified_sampler— sample diversified by a field.
Geospatial buckets:
geo_distance— distance rings around an origin.geohash_grid— geohash cells for map clustering.geotile_grid— web-mapzoom/x/ytiles.
Relational:
nested— sub-aggs inside a nested object path (plusreverse_nested).children— sub-aggs over relation children.
Pagination:
composite— multi-source buckets streamed with anaftercursor.
Pipeline aggregations #
Dedicated pages under
pipeline aggregations:
sibling statistics (avg_bucket, sum_bucket, min_bucket,
max_bucket, stats_bucket, extended_stats_bucket,
percentiles_bucket) and parent transforms (derivative,
cumulative_sum, cumulative_cardinality, serial_diff,
moving_avg, moving_fn, bucket_script, bucket_selector,
bucket_sort, normalize).
Compatibility notes #
Two additional Elasticsearch aggregation names parse for wire compatibility but do not compute yet:
scripted_metric— accepted, returns an empty object.significant_text— requires a stored column on a text field; text fields are columnless by design, so it returns no buckets (usesignificant_termson a keyword field).
A few implementations are simplified relative to Elasticsearch — noted
on their pages:
top_metrics (returns the first
matching docs),
children (scope-keyed, not
per-parent),
t_test (paired evaluation for all types)
and
rate (sum semantics).