Lesson 2 — Text search #
How text actually matches: analysis, exact terms, phrases, and the
error-tolerant family. Everything below runs live on the logs collection
(150 documents with message, level, host, url.path).
Exact terms vs analyzed text #
A
term query matches the stored value
exactly — no tokenization, no casing help. A
match query runs the field’s analyzer
first, so casing and punctuation stop mattering:
Compare with a term on the keyword field level — exact, and therefore
fast and predictable:
25 of the 150 log lines are errors. If you had written "ERROR" the count
would be 0 — that is the analyzer difference in one number.
Phrases #
phrase (alias match_phrase)
requires the tokens in order, with slop budget for gaps. The dataset’s
known answer — HTTP/1.1 200 appears in 75 documents — shows why analysis
matters: the tokenizer folds HTTP/1.1 into one token, and the phrase
matches it exactly:
The error-tolerant family #
Typos (
fuzzy), prefixes
(
prefix), patterns
(
wildcard and
regexp), and endings
(
suffix). The fuzzy query below
misspells search as serch — one dropped letter — and still finds the
same 25 GET /search requests the exact match at the top of this page
found:
One string, full mini-language #
query_string packs booleans,
fields, ranges and wildcards into one expression:
Combining with bool #
bool is the workhorse — must for
scoring clauses, filter for cheap yes/no:
What you learned #
term= exact value,match= analyzed text.- Phrases order tokens;
slopbuys distance. - Fuzzy/prefix/wildcard/regexp/suffix cover the typo and pattern cases.
boolcomposes everything, withfilterfor scoring-free clauses.
Next: Lesson 3 — structure and numbers, or read the analysis API to inspect what any analyzer does to a string.