Built-in analysis components

Built-in analysis components #

The server ships 83 analyzers, 50 tokenizers, 15 normalizers and 327 token filters, registered at startup. The authoritative, always-current list on YOUR node is:

GET /_analysis/components

Components are referenced by name — in a field’s schema ("analyzer": "english"), in a custom analyzer definition, or in an ad-hoc _analysis/analyze pipeline.

Analyzers #

General purpose:

fingerprint, keyword, pattern, simple, standard, stop, whitespace

CJK and transliteration:

cjk, ik_max_word, ik_smart, jieba, kuromoji, nori, pinyin, smartcn stconvert_s2t, stconvert_t2s

Language analyzers (stop words + stemming per language):

afrikaans, amharic, arabic, armenian, aymara, azerbaijani, basque, bengali brazilian, bulgarian, catalan, cherokee, croatian, czech, danish, dutch english, estonian, filipino, finnish, french, galician, georgian, german greek, guarani, hebrew, hindi, hungarian, indonesian, irish, italian khmer, latvian, lithuanian, malay, marathi, mongolian, nahuatl, navajo nepali, norwegian, persian, polish, portuguese, quechua, romanian, russian scottish_gaelic, serbian, slovak, slovenian, sorani, spanish, swahili, swedish tagalog, tamil, thai, tibetan, turkish, ukrainian, urdu, vietnamese welsh, yiddish

The CJK tokenizers (ik_max_word/ik_smart, jieba, smartcn, kuromoji, nori) load their dictionaries at runtime from the external <config>/analysis/ directory (make init-analysis-dicts); without the dictionary files the analyzers answer but degrade to simpler segmentation.

Tokenizers #

burmese, camel_case, char_group, chinese_char, classic, code, compound_word, edge_ngram elision, email, emoji, fingerprint, hyphenated, icu_tokenizer, ik_max_word, ik_smart jieba, json_field, keyword, kuromoji_tokenizer, letter, log, lowercase, markdown microblog, ngram, nori_tokenizer, path_hierarchy, pattern, phone_number, pinyin, punctuation reverse, script_boundary, sentence, simple_pattern, simple_pattern_split, sliding_window, smartcn_tokenizer, standard stconvert_s2t, stconvert_t2s, structured_id, tab_separated, thai, truncate, uax_url_email, url whitespace, wildcard

Normalizers #

Applied to raw text BEFORE tokenization (keyword fields use them at index time via the field’s normalizer schema parameter):

collapse_whitespace, html_strip, lowercase, mapping, pattern_replace, pinyin, pinyin_first_letter, stconvert_s2t stconvert_t2s, trim, unicode_nfc, unicode_nfd, unicode_nfkc, unicode_nfkd, uppercase

Token filters #

327 filters covering stemming (per language), stop words, folding (asciifolding, cjk_width, decimal_digit), n-grams (edge_ngram, ngram), truncation, delimiters, security masking (credit_card_mask, email_mask), encoding (base64_encode) and more. See GET /_analysis/components for the live list; the analyze API accepts any of them by name in the filter array.

Calendar September 27, 2026
Edit Edit this page