Skip to content

Building blocks

completr.connect and its collections are built from public parts, in completr.lowlevel in Python and at the crate root in Rust. Use them when the client does not fit:

  • Offline builds: build segment files in a batch job, in bounded memory, and ship them.
  • In-memory indexes: complete over documents without a database, for tests and scripts.
  • Full control: transactions with several indexes, strict conflict checks, leases, snapshots of old versions, other fusion methods, or index options such as popularity_weight.
  • Many writers: processes submit change sets to an inbox, and one lease-elected ingestor commits them.
from completr.lowlevel import ChangeSet, Engine, Index, Ingestor, Segment, SegmentWriter, open_database
Part Role Page
Segment An immutable, self-contained set of documents plus the ids it deletes from older segments. One file, 8-byte aligned, checksummed, read in place through a memory map. Indexes and segments
SegmentWriter Writes segment files in bounded memory, starting a new one when a memory budget is reached. Indexes and segments
Index A list of segments, oldest to newest, searched as one. Indexes and segments
Engine Indexes by name, replaced atomically and searched as layers. Indexes and segments
Database Versioned manifests of named indexes in a directory or bucket. Databases and engines
Transaction An optimistic, atomic commit of one or more indexes. Databases and engines
Replica Follows a database and publishes changed indexes to an Engine. Databases and engines
Lease An expiring, fenced lock in the database. Databases and engines
ChangeSet, Ingestor Changes submitted to an inbox, and the lease-elected process that commits them. Many writers
Store Raw object-store access: get, put, create-only put, list, segments. Python API

A collection is an index of the same name. A client and the building blocks work on the same database, and db.database returns the Database under a client:

import completr

db = completr.connect("./data")
db.get_or_create_collection("products").add([{"id": "fryer", "text": "Air Fryer", "popularity": 0.8}])

database = db.database
print(database.index_names(), database.latest_version())
['products'] 2

With an inbox

Writes through a client commit directly. With the building blocks, processes can instead submit change sets to an inbox in the database, and an ingestor commits them:

flowchart LR
    P["Any process<br/>database.submit(changes)"] -- "create-only put" --> I
    subgraph B["Bucket or directory: the database"]
        I["_inbox/ change sets"]
        M["_versions/ manifests"]
        S["segments/ immutable files"]
        L["_locks/ ingestor lease"]
    end
    I --> W["Ingestor<br/>any process holding the lease"]
    W -- "commit" --> M
    W -- "write, compact" --> S
    W -. "lease, fencing" .-> L
    M -- "new versions" --> F1["Your app + completr<br/>Replica, Engine"]
    S -- "changed segments" --> F1
    M --> F2["Your app + completr"]
    S --> F2
  • Ingestion is a role, not a service. Any process may run an Ingestor; the one holding an expiring, fenced lease commits the inbox, and another takes over when it stops.
  • Readers scale with your application. Each serving process runs a Replica that loads only changed, immutable segments and switches versions atomically.