Skip to content

Indexes and segments

A Segment is an immutable set of documents, plus the ids it deletes from older segments. An Index is a list of segments, oldest to newest, searched as one. An Engine holds indexes by name and searches them as layers. None of them needs a database.

In-memory indexes

Index.from_documents builds one segment in memory and returns an index over it. It takes the same document input as Collection.add.

from completr.lowlevel import Index

docs = [
    {"id": "ps5", "text": "PlayStation 5 Console", "popularity": 0.9, "abbreviations": ["PS5"]},
    {"id": "psvr", "text": "PlayStation VR2", "popularity": 0.4},
    {"id": "anc", "text": "Wireless Noise Cancelling Headphones", "popularity": 0.7, "synonyms": ["bluetooth headphones"]},
]
index = Index.from_documents(docs)

for query in ["play", "PS5", "wirless", "cancelling"]:
    print(query, [(s.text, s.kind, round(s.score, 3)) for s in index.complete(query, limit=3)])
use completr::{Document, Index};

let index = Index::from_documents([
    Document::keyed("ps5", "PlayStation 5 Console", 0.9).with_abbreviation("PS5"),
    Document::keyed("psvr", "PlayStation VR2", 0.4),
    Document::keyed("anc", "Wireless Noise Cancelling Headphones", 0.7)
        .with_synonym("bluetooth headphones"),
])?;

for query in ["play", "PS5", "wirless", "cancelling"] {
    for s in index.complete(query, 3) {
        println!("{query}: {} {} {:.3}", s.text, s.kind.as_str(), s.score);
    }
}
play [('PlayStation 5 Console', 'prefix', 0.594), ('PlayStation VR2', 'prefix', 0.338)]
PS5 [('PlayStation 5 Console', 'abbreviation', 0.594)]
wirless [('Wireless Noise Cancelling Headphones', 'fuzzy', 0.275)]
cancelling [('Wireless Noise Cancelling Headphones', 'infix', 0.287)]

An index's suggestions have no layer. Synonyms are searched separately with complete_aliases, which returns AliasSuggestions with id, text (the document's, not the synonym), score and layer:

print(index.complete_aliases("bluetooth head"))
print(index.get("ps5"), len(index))
[AliasSuggestion(id='anc', text="Wireless Noise Cancelling Headphones", score=0.5474)]
Document(id='ps5', text="PlayStation 5 Console", popularity=0.9) 3

get(id) returns a stored document, by int id, string key or key_id; in Rust, document(id) and document_by_key(key). complete, complete_aliases, vector_search and hybrid_search all take contexts=; in Rust, the _with variants take SearchOptions:

use completr::SearchOptions;

let games = SearchOptions::new(10).contexts(["games"]);
let hits = index.complete_with("play", &games);

Segments

Segment.build builds one segment from documents, and deletes lists ids it hides in older segments. With path=, it streams the build into a file and memory-maps it. A newer copy of an id supersedes older ones, and an index of many segments ranks exactly like one compacted segment.

from completr.lowlevel import Segment

base = Segment.build(docs, path="products.seg")
delta = Segment.build([{"id": "fryer", "text": "Air Fryer", "popularity": 0.8}], deletes=["psvr"])

index = Index([Segment.open("products.seg"), delta])
print([s.text for s in index.complete("play")], [s.text for s in index.complete("air")])

merged = index.compact()   # one segment with the same documents
print(len(merged), merged.deletes())
['PlayStation 5 Console'] ['Air Fryer']
3 []
Method Does
Segment.open(path) Memory-maps a segment file and checks its structure; reads only its headers.
segment.verify() Checks the checksum and every section.
Segment.from_bytes(data), segment.to_bytes(), segment.save(path) Converts and stores segments; from_bytes verifies.
segment.documents(), ids(), deletes(), len(segment) The segment's contents.
index.compact(path=None) Merges an index's segments into one, byte for byte as a rebuild would.
index.segments() The index's segments, oldest first.

Offline builds

A SegmentWriter writes segment files into a directory, starting a new file whenever building the current one would pass its memory_budget (256 MB by default). So any corpus indexes in bounded memory, and the segments rank exactly like one. Of several documents with one id, the last added wins.

from completr.lowlevel import SegmentWriter

writer = SegmentWriter("segments", memory_budget=1_000_000)
writer.add([{"id": f"p{i}", "text": f"product {i}"} for i in range(20_000)])
segments = writer.finish()   # segments/000000.seg, 000001.seg, ...

index = Index(segments)
print(len(segments), len(index))
use std::sync::Arc;
use completr::{BuildOptions, Document, Index, IndexOptions, SegmentWriter};

let mut writer = SegmentWriter::new(BuildOptions::default(), "segments")?.memory_budget(1 << 20);
for i in 0..20_000 {
    writer.add(Document::keyed(format!("p{i}"), format!("product {i}"), 0.0))?;
}
let segments = writer.finish()?.into_iter().map(Arc::new).collect();
let index = Index::new(segments, IndexOptions::default())?;
6 20000

Indexing the 124,440 HN titles takes 0.29 s and peaks at 59 MB; all of English Wikipedia, 7.2 million titles, is written and compacted into one segment in 25.6 s. Segments can be shipped as files and opened with Segment.open wherever they are served.

Ranking

Suggestions are ordered by score, then shorter text, then id. popularity_weight (0.4 by default) sets how much popularity counts, and max_score the raw score that normalises to 1.0:

flat = Index.from_documents(
    [{"id": "a", "text": "PlayStation VR2", "popularity": 0.0}, {"id": "b", "text": "PlayStation 5 Console", "popularity": 1.0}],
    popularity_weight=0.0,
)
print([s.text for s in flat.complete("play")])   # shorter text first
['PlayStation VR2', 'PlayStation 5 Console']

An index estimates max_score from its documents unless you pass one. Pin it, with max_score= or Transaction.set_max_score, when scores must stay comparable across rebuilds. Configuration lists every index option.

Pass vectors= to Index.from_documents or Segment.build. vector_search returns the nearest documents as Suggestions with kind set to semantic. hybrid_search runs lexical completion and vector search and fuses both lists; it returns HybridSuggestions, which also report lexical_score and semantic_score, either of which is None when the document came from one side only.

import zlib

import numpy as np

def embed(texts):
    """Stands in for your embedding model: one normalised float32 row per text."""
    rng = np.random.default_rng(zlib.crc32("\n".join(texts).encode()))
    rows = rng.standard_normal((len(texts), 64), dtype=np.float32)
    return rows / np.linalg.norm(rows, axis=1, keepdims=True)

vectors = embed([d["text"] for d in docs])
semantic = Index.from_documents(docs, vectors=vectors, vector_bits=4)

for s in semantic.vector_search(vectors[2], limit=2):
    print(s.id, s.kind, round(s.score, 3))
for s in semantic.hybrid_search("wireless no", embed(["wireless no"])[0], limit=3, fusion="rrf"):
    print(s.id, s.kind, round(s.score, 4), s.lexical_score, s.semantic_score)
use completr::{Fusion, HybridOptions};

let query_vector: Vec<f32> = embed("wireless no");   // your model
let nearest = index.vector_search(&query_vector, 10)?;   // Vec<Suggestion>, kind Semantic
let options = HybridOptions::default().fusion(Fusion::ReciprocalRank { k: 60.0 });
let hits = index.hybrid_search("wireless no", &query_vector, 10, &options)?;
anc semantic 0.998
psvr semantic -0.091
anc prefix 0.0325 0.4825822422419287 0.037495680153369904
ps5 semantic 0.0164 None 0.20452365279197693
psvr semantic 0.0159 None 0.031032711267471313
fusion Score Parameters
rrf (default) 1 / (k + lexical_rank) + 1 / (k + semantic_rank), ranks from 1 rrf_k, 60 by default; needs no calibration
weighted (1 - w) * lexical + w * max(semantic, 0) semantic_weight (w), 0.5 by default
lexical_first lexical hits in their order, then semantic-only hits none
semantic.hybrid_search("wireless no", embed(["wireless no"])[0], limit=3, fusion="weighted", semantic_weight=0.3)
semantic.hybrid_search("wireless no", embed(["wireless no"])[0], limit=3, fusion="lexical_first")

candidates sets how many hits are taken from each side before fusing; by default, max(2 * limit, 20).

Engines and layers

An Engine holds indexes by name and publishes new versions atomically. A search names a list of indexes, its layers; later layers override earlier ones per document id, as collection layers do.

from completr.lowlevel import Engine

shared = Index.from_documents([
    {"id": "ps5", "text": "PlayStation 5 Console", "popularity": 0.9},
    {"id": "psvr", "text": "PlayStation VR2", "popularity": 0.4},
])
acme = Index.from_documents([
    {"id": "psvr", "text": "PlayStation VR2 Bundle", "popularity": 1.0},   # overrides "psvr"
    {"id": "cards", "text": "Playing Cards", "popularity": 0.3},          # adds a document
])

engine = Engine()
engine.publish({"shared": shared, "acme": acme})
for s in engine.complete("play", ["shared", "acme"], limit=10):
    print(s.text, s.layer)
use std::sync::Arc;
use completr::Engine;

let engine = Engine::new();
engine.publish([
    ("shared".to_owned(), Some(Arc::new(shared))),
    ("acme".to_owned(), Some(Arc::new(acme))),
]);
for layered in engine.complete(&["shared", "acme"], "play", 10) {
    println!("{} from layer {}", layered.suggestion.text, layered.layer);
}
PlayStation 5 Console shared
PlayStation VR2 Bundle acme
Playing Cards acme

In Rust, the engine returns LayeredSuggestions whose layer is the position in the layer list.

  • A layer name that is not published counts as an empty layer.
  • publish replaces the named indexes atomically. A search already running keeps the versions it started with.
  • engine.publish({"acme": None}) removes a layer.
  • engine.get(name) returns a published Index, and engine.names() lists them.
  • complete_aliases, vector_search and hybrid_search take a list of layers too.

A layered search takes limit * overfetch candidates from each layer before overrides are applied, so that overridden documents do not leave the result short. The default, 2, suits layers that override a small share of the documents. Raise it for layers that shadow many results:

engine = Engine(overfetch=4)