Indexes and segments¶
A Segment is an immutable set of documents, plus the ids it deletes from older segments. An Index is a
list of segments, oldest to newest, searched as one. An Engine holds indexes by name and searches them
as layers. None of them needs a database.
In-memory indexes¶
Index.from_documents builds one segment in memory and returns an index over it. It takes the same
document input as Collection.add.
from completr.lowlevel import Index
docs = [
{"id": "ps5", "text": "PlayStation 5 Console", "popularity": 0.9, "abbreviations": ["PS5"]},
{"id": "psvr", "text": "PlayStation VR2", "popularity": 0.4},
{"id": "anc", "text": "Wireless Noise Cancelling Headphones", "popularity": 0.7, "synonyms": ["bluetooth headphones"]},
]
index = Index.from_documents(docs)
for query in ["play", "PS5", "wirless", "cancelling"]:
print(query, [(s.text, s.kind, round(s.score, 3)) for s in index.complete(query, limit=3)])
use completr::{Document, Index};
let index = Index::from_documents([
Document::keyed("ps5", "PlayStation 5 Console", 0.9).with_abbreviation("PS5"),
Document::keyed("psvr", "PlayStation VR2", 0.4),
Document::keyed("anc", "Wireless Noise Cancelling Headphones", 0.7)
.with_synonym("bluetooth headphones"),
])?;
for query in ["play", "PS5", "wirless", "cancelling"] {
for s in index.complete(query, 3) {
println!("{query}: {} {} {:.3}", s.text, s.kind.as_str(), s.score);
}
}
play [('PlayStation 5 Console', 'prefix', 0.594), ('PlayStation VR2', 'prefix', 0.338)]
PS5 [('PlayStation 5 Console', 'abbreviation', 0.594)]
wirless [('Wireless Noise Cancelling Headphones', 'fuzzy', 0.275)]
cancelling [('Wireless Noise Cancelling Headphones', 'infix', 0.287)]
An index's suggestions have no layer. Synonyms are searched separately with complete_aliases, which
returns AliasSuggestions with id, text (the document's, not the synonym), score and layer:
[AliasSuggestion(id='anc', text="Wireless Noise Cancelling Headphones", score=0.5474)]
Document(id='ps5', text="PlayStation 5 Console", popularity=0.9) 3
get(id) returns a stored document, by int id, string key or key_id; in Rust, document(id) and
document_by_key(key). complete, complete_aliases, vector_search and hybrid_search all take
contexts=; in Rust, the _with variants take SearchOptions:
use completr::SearchOptions;
let games = SearchOptions::new(10).contexts(["games"]);
let hits = index.complete_with("play", &games);
Segments¶
Segment.build builds one segment from documents, and deletes lists ids it hides in older segments.
With path=, it streams the build into a file and memory-maps it. A newer copy of an id supersedes older
ones, and an index of many segments ranks exactly like one compacted segment.
from completr.lowlevel import Segment
base = Segment.build(docs, path="products.seg")
delta = Segment.build([{"id": "fryer", "text": "Air Fryer", "popularity": 0.8}], deletes=["psvr"])
index = Index([Segment.open("products.seg"), delta])
print([s.text for s in index.complete("play")], [s.text for s in index.complete("air")])
merged = index.compact() # one segment with the same documents
print(len(merged), merged.deletes())
| Method | Does |
|---|---|
Segment.open(path) |
Memory-maps a segment file and checks its structure; reads only its headers. |
segment.verify() |
Checks the checksum and every section. |
Segment.from_bytes(data), segment.to_bytes(), segment.save(path) |
Converts and stores segments; from_bytes verifies. |
segment.documents(), ids(), deletes(), len(segment) |
The segment's contents. |
index.compact(path=None) |
Merges an index's segments into one, byte for byte as a rebuild would. |
index.segments() |
The index's segments, oldest first. |
Offline builds¶
A SegmentWriter writes segment files into a directory, starting a new file whenever building the
current one would pass its memory_budget (256 MB by default). So any corpus indexes in bounded memory,
and the segments rank exactly like one. Of several documents with one id, the last added wins.
from completr.lowlevel import SegmentWriter
writer = SegmentWriter("segments", memory_budget=1_000_000)
writer.add([{"id": f"p{i}", "text": f"product {i}"} for i in range(20_000)])
segments = writer.finish() # segments/000000.seg, 000001.seg, ...
index = Index(segments)
print(len(segments), len(index))
use std::sync::Arc;
use completr::{BuildOptions, Document, Index, IndexOptions, SegmentWriter};
let mut writer = SegmentWriter::new(BuildOptions::default(), "segments")?.memory_budget(1 << 20);
for i in 0..20_000 {
writer.add(Document::keyed(format!("p{i}"), format!("product {i}"), 0.0))?;
}
let segments = writer.finish()?.into_iter().map(Arc::new).collect();
let index = Index::new(segments, IndexOptions::default())?;
Indexing the 124,440 HN titles takes 0.29 s and peaks at 59 MB; all of English Wikipedia, 7.2 million
titles, is written and compacted into one segment in 25.6 s. Segments can be shipped as files and opened
with Segment.open wherever they are served.
Ranking¶
Suggestions are ordered by score, then shorter text, then id. popularity_weight (0.4 by default) sets
how much popularity counts, and max_score the raw score that normalises to 1.0:
flat = Index.from_documents(
[{"id": "a", "text": "PlayStation VR2", "popularity": 0.0}, {"id": "b", "text": "PlayStation 5 Console", "popularity": 1.0}],
popularity_weight=0.0,
)
print([s.text for s in flat.complete("play")]) # shorter text first
An index estimates max_score from its documents unless you pass one. Pin it, with max_score= or
Transaction.set_max_score, when scores must stay comparable across rebuilds.
Configuration lists every index option.
Semantic and hybrid search¶
Pass vectors= to Index.from_documents or Segment.build. vector_search returns the nearest
documents as Suggestions with kind set to semantic. hybrid_search runs lexical completion and
vector search and fuses both lists; it returns HybridSuggestions, which also report lexical_score and
semantic_score, either of which is None when the document came from one side only.
import zlib
import numpy as np
def embed(texts):
"""Stands in for your embedding model: one normalised float32 row per text."""
rng = np.random.default_rng(zlib.crc32("\n".join(texts).encode()))
rows = rng.standard_normal((len(texts), 64), dtype=np.float32)
return rows / np.linalg.norm(rows, axis=1, keepdims=True)
vectors = embed([d["text"] for d in docs])
semantic = Index.from_documents(docs, vectors=vectors, vector_bits=4)
for s in semantic.vector_search(vectors[2], limit=2):
print(s.id, s.kind, round(s.score, 3))
for s in semantic.hybrid_search("wireless no", embed(["wireless no"])[0], limit=3, fusion="rrf"):
print(s.id, s.kind, round(s.score, 4), s.lexical_score, s.semantic_score)
use completr::{Fusion, HybridOptions};
let query_vector: Vec<f32> = embed("wireless no"); // your model
let nearest = index.vector_search(&query_vector, 10)?; // Vec<Suggestion>, kind Semantic
let options = HybridOptions::default().fusion(Fusion::ReciprocalRank { k: 60.0 });
let hits = index.hybrid_search("wireless no", &query_vector, 10, &options)?;
anc semantic 0.998
psvr semantic -0.091
anc prefix 0.0325 0.4825822422419287 0.037495680153369904
ps5 semantic 0.0164 None 0.20452365279197693
psvr semantic 0.0159 None 0.031032711267471313
fusion |
Score | Parameters |
|---|---|---|
rrf (default) |
1 / (k + lexical_rank) + 1 / (k + semantic_rank), ranks from 1 |
rrf_k, 60 by default; needs no calibration |
weighted |
(1 - w) * lexical + w * max(semantic, 0) |
semantic_weight (w), 0.5 by default |
lexical_first |
lexical hits in their order, then semantic-only hits | none |
semantic.hybrid_search("wireless no", embed(["wireless no"])[0], limit=3, fusion="weighted", semantic_weight=0.3)
semantic.hybrid_search("wireless no", embed(["wireless no"])[0], limit=3, fusion="lexical_first")
candidates sets how many hits are taken from each side before fusing; by default, max(2 * limit, 20).
Engines and layers¶
An Engine holds indexes by name and publishes new versions atomically. A search names a list of
indexes, its layers; later layers override earlier ones per document id, as
collection layers do.
from completr.lowlevel import Engine
shared = Index.from_documents([
{"id": "ps5", "text": "PlayStation 5 Console", "popularity": 0.9},
{"id": "psvr", "text": "PlayStation VR2", "popularity": 0.4},
])
acme = Index.from_documents([
{"id": "psvr", "text": "PlayStation VR2 Bundle", "popularity": 1.0}, # overrides "psvr"
{"id": "cards", "text": "Playing Cards", "popularity": 0.3}, # adds a document
])
engine = Engine()
engine.publish({"shared": shared, "acme": acme})
for s in engine.complete("play", ["shared", "acme"], limit=10):
print(s.text, s.layer)
use std::sync::Arc;
use completr::Engine;
let engine = Engine::new();
engine.publish([
("shared".to_owned(), Some(Arc::new(shared))),
("acme".to_owned(), Some(Arc::new(acme))),
]);
for layered in engine.complete(&["shared", "acme"], "play", 10) {
println!("{} from layer {}", layered.suggestion.text, layered.layer);
}
In Rust, the engine returns LayeredSuggestions whose layer is the position in the layer list.
- A layer name that is not published counts as an empty layer.
publishreplaces the named indexes atomically. A search already running keeps the versions it started with.engine.publish({"acme": None})removes a layer.engine.get(name)returns a publishedIndex, andengine.names()lists them.complete_aliases,vector_searchandhybrid_searchtake a list of layers too.
A layered search takes limit * overfetch candidates from each layer before overrides are applied, so
that overridden documents do not leave the result short. The default, 2, suits layers that override a
small share of the documents. Raise it for layers that shadow many results: