Built-in analysis components #
The server ships 83 analyzers, 50 tokenizers, 15 normalizers and 327 token filters, registered at startup. The authoritative, always-current list on YOUR node is:
GET /_analysis/components
Components are referenced by name — in a field’s schema
("analyzer": "english"), in a
custom analyzer
definition, or in an ad-hoc
_analysis/analyze pipeline.
Analyzers #
General purpose:
fingerprint, keyword, pattern, simple, standard, stop, whitespace
CJK and transliteration:
cjk, ik_max_word, ik_smart, jieba, kuromoji, nori, pinyin, smartcn
stconvert_s2t, stconvert_t2s
Language analyzers (stop words + stemming per language):
afrikaans, amharic, arabic, armenian, aymara, azerbaijani, basque, bengali
brazilian, bulgarian, catalan, cherokee, croatian, czech, danish, dutch
english, estonian, filipino, finnish, french, galician, georgian, german
greek, guarani, hebrew, hindi, hungarian, indonesian, irish, italian
khmer, latvian, lithuanian, malay, marathi, mongolian, nahuatl, navajo
nepali, norwegian, persian, polish, portuguese, quechua, romanian, russian
scottish_gaelic, serbian, slovak, slovenian, sorani, spanish, swahili, swedish
tagalog, tamil, thai, tibetan, turkish, ukrainian, urdu, vietnamese
welsh, yiddish
The CJK tokenizers (ik_max_word/ik_smart, jieba, smartcn,
kuromoji, nori) load their dictionaries at runtime from the external
<config>/analysis/ directory (make init-analysis-dicts); without the
dictionary files the analyzers answer but degrade to simpler
segmentation.
Tokenizers #
burmese, camel_case, char_group, chinese_char, classic, code, compound_word, edge_ngram
elision, email, emoji, fingerprint, hyphenated, icu_tokenizer, ik_max_word, ik_smart
jieba, json_field, keyword, kuromoji_tokenizer, letter, log, lowercase, markdown
microblog, ngram, nori_tokenizer, path_hierarchy, pattern, phone_number, pinyin, punctuation
reverse, script_boundary, sentence, simple_pattern, simple_pattern_split, sliding_window, smartcn_tokenizer, standard
stconvert_s2t, stconvert_t2s, structured_id, tab_separated, thai, truncate, uax_url_email, url
whitespace, wildcard
Normalizers #
Applied to raw text BEFORE tokenization (keyword fields use them at
index time via the field’s normalizer schema parameter):
collapse_whitespace, html_strip, lowercase, mapping, pattern_replace, pinyin, pinyin_first_letter, stconvert_s2t
stconvert_t2s, trim, unicode_nfc, unicode_nfd, unicode_nfkc, unicode_nfkd, uppercase
Token filters #
327 filters covering stemming (per language), stop words,
folding (asciifolding, cjk_width, decimal_digit), n-grams
(edge_ngram, ngram), truncation, delimiters, security masking
(credit_card_mask, email_mask), encoding (base64_encode) and more.
See GET /_analysis/components for the live list; the
analyze API accepts any of them by name in the
filter array.