Completion¶
complete returns ranked suggestions for what a user has typed so far. Every match kind is collected and
ranked in one request; you do not choose between prefix, infix or fuzzy search.
The examples on this page use this index:
from completr import Index
index = Index.from_documents([
{"id": "ml", "text": "Machine Learning", "popularity": 0.9, "abbreviations": ["ML"], "contexts": ["courses"]},
{"id": "mv", "text": "Machine Vision", "popularity": 0.4, "contexts": ["books"]},
{"id": "ds", "text": "Data Science", "popularity": 0.7, "synonyms": ["data analytics"], "contexts": ["courses", "books"]},
{"id": "cs", "text": "Computer Science", "popularity": 0.6},
{"id": "nlp", "text": "Natural Language Processing", "popularity": 0.5, "abbreviations": ["NLP"]},
])
Match kinds¶
| Query | Top suggestion | kind |
Highlights |
|---|---|---|---|
machine learning |
Machine Learning | exact |
Machine Learning |
mach |
Machine Learning, Machine Vision | prefix |
Machine Learning |
Data Sc |
Data Science | prefix |
Data Science |
nlp |
Natural Language Processing | abbreviation |
none |
science |
Data Science, Computer Science | infix |
Data Science |
vison |
Machine Vision | fuzzy |
Machine Vision |
machne lerning |
Machine Learning | fuzzy |
Machine Learning |
datascience |
Data Science | exact |
Data Science |
for query in ["machine learning", "mach", "nlp", "science", "machne lerning", "datascience"]:
print(f"{query!r:18}", [(s.text, s.kind) for s in index.complete(query, limit=2)])
'machine learning' [('Machine Learning', 'exact')]
'mach' [('Machine Learning', 'prefix'), ('Machine Vision', 'prefix')]
'nlp' [('Natural Language Processing', 'abbreviation')]
'science' [('Data Science', 'infix'), ('Computer Science', 'infix')]
'machne lerning' [('Machine Learning', 'fuzzy')]
'datascience' [('Data Science', 'exact')]
Notes on each kind:
- Exact and prefix matching works on the whole text and word by word, so
Data Sccompletes Data Science. Every query word must start a word of the text. - Abbreviations match exactly and case-insensitively:
k8sfinds a document with the abbreviationK8S, butk8does not, so short codes do not flood the results. - Infix matches a word inside the text. Words shorter than
min_word_chars(3) are not indexed for infix matching. - Fuzzy matching corrects up to
max_edit_distance(2) edits per word, SymSpell-style, and verifies candidates with Levenshtein distance. Fuzzy hits always rank below the weakest direct match. - Word decomposition splits run-together input, such as
datascience, into dictionary words.
Highlights¶
highlights lists the (start, end) character ranges of text that matched, sorted, ready to render in
bold. Abbreviation matches have no highlights, since the query does not appear in the text.
Contexts¶
Tag documents with contexts and pass contexts= to any request (complete, complete_aliases,
vector_search and hybrid_search). A document qualifies when it has any of the requested contexts.
Filtering happens while candidates are collected, so a filtered request still returns up to limit
results.
Contexts are the only filter completr has. Use them for categories, tenants or languages, or use separate indexes and layers when the partitions are large.
Synonyms¶
Synonyms are searched with complete_aliases, which prefix-matches the synonyms and returns the documents
they belong to. Keeping them separate from complete means an alternative name never displaces a direct
match; call both if your box shows both.
print(index.complete_aliases("data an")) # [AliasSuggestion(id='ds', text="Data Science", score=...)]
An AliasSuggestion has id, text (the document's text, not the synonym), score and layer.
Limits¶
limit (10 by default) caps the number of suggestions. Queries of up to short_query_chars characters
(3 by default) are answered from a per-index cache of short_query_limit (100) results, so one- to
three-character prefixes, which match large parts of an index, stay fast. Requests with contexts bypass
this cache.
Ranking¶
Suggestions are ordered by score, then shorter text, then id. The score combines the match kind, the text
length, whole-word bonuses and popularity; Concepts lists the formula.
popularity_weight (0.4 by default) sets how much popularity counts: