Skip to content

Python API

The reference for the completr package, generated from its typed stubs (__init__.pyi). Every name on this page is importable from completr.

Connecting

connect

connect(
    url: str,
    options: Mapping[str, str] | None = None,
    cache_dir: str | PathLike[str] | None = None,
    **build_options,
) -> Database

Opens the database at url: a local path, or an s3://, gs://, az:// or memory:// URL.

connect takes the same arguments as Database: url, options, cache_dir, and the keyword-only build options min_word_chars, max_edit_distance, fuzzy_prefix_chars, vector_bits, compact_keys and build_threads.

connect_async async

connect_async(
    url: str,
    options: Mapping[str, str] | None = None,
    cache_dir: str | PathLike[str] | None = None,
    **build_options: Any,
) -> AsyncDatabase

Opens the database at url without blocking the event loop.

Documents

Document

Document(
    id: Id,
    text: str,
    popularity: float = 0.0,
    *,
    synonyms: Sequence[str] = ...,
    abbreviations: Sequence[str] = ...,
    contexts: Sequence[str] = ...,
    vector: Vector | None = None,
)

id property

id: Id

text property

text: str

popularity property

popularity: float

synonyms property

synonyms: list[str]

abbreviations property

abbreviations: list[str]

contexts property

contexts: list[str]

to_dict

to_dict() -> DocumentDict

DocumentDict

Bases: TypedDict

id instance-attribute

id: Id

text instance-attribute

text: str

popularity instance-attribute

popularity: float

synonyms instance-attribute

synonyms: Sequence[str]

abbreviations instance-attribute

abbreviations: Sequence[str]

contexts instance-attribute

contexts: Sequence[str]

vector instance-attribute

vector: Vector

Id module-attribute

Id = int | str

A document id: an int, or a string key.

Documents module-attribute

Documents = Iterable[DocumentDict | Document] | Any

Dicts, Documents, or a pandas, polars or pyarrow table with those columns.

Vector module-attribute

Vector = Buffer | Sequence[float]

An embedding: a 1-D float32 array or a list of floats.

Vectors module-attribute

Vectors = Buffer | Sequence[Sequence[float]]

One embedding row per document: a float32 array of shape (n, dim), or lists.

key_id

key_id(key: str) -> int

Suggestions

Suggestion

id instance-attribute

id: Id

text instance-attribute

text: str

score instance-attribute

score: float

kind instance-attribute

kind: MatchKind

highlights instance-attribute

highlights: list[tuple[int, int]]

(start, end) character offsets of text that matched the query.

layer instance-attribute

layer: str | None

HybridSuggestion

id instance-attribute

id: Id

text instance-attribute

text: str

score instance-attribute

score: float

kind instance-attribute

kind: MatchKind

lexical_score instance-attribute

lexical_score: float | None

semantic_score instance-attribute

semantic_score: float | None

highlights instance-attribute

highlights: list[tuple[int, int]]

layer instance-attribute

layer: str | None

AliasSuggestion

id instance-attribute

id: Id

text instance-attribute

text: str

score instance-attribute

score: float

layer instance-attribute

layer: str | None

MatchKind module-attribute

MatchKind = Literal[
    "exact",
    "prefix",
    "abbreviation",
    "infix",
    "fuzzy",
    "semantic",
]

FusionKind module-attribute

FusionKind = Literal['rrf', 'weighted', 'lexical_first']

Segments, indexes and engines

Segment

vector_dim property

vector_dim: int | None

size_bytes property

size_bytes: int

build staticmethod

build(
    documents: Documents,
    deletes: Sequence[Id] = ...,
    vectors: Vectors | None = None,
    *,
    min_word_chars: int = 3,
    max_edit_distance: int = 2,
    fuzzy_prefix_chars: int = 7,
    vector_bits: int = 4,
    compact_keys: bool = False,
    build_threads: int = 1,
    path: str | PathLike[str] | None = None,
) -> Segment

open staticmethod

open(path: str | PathLike[str]) -> Segment

from_bytes staticmethod

from_bytes(data: bytes) -> Segment

verify

verify() -> None

to_bytes

to_bytes() -> bytes

save

save(path: str | PathLike[str]) -> None

documents

documents() -> list[Document]

ids

ids() -> list[int]

deletes

deletes() -> list[int]

Index

Index(
    segments: Sequence[Segment],
    *,
    max_score: float | None = None,
    popularity_weight: float = 0.4,
    short_query_chars: int = 3,
    short_query_limit: int = 100,
    short_query_cache_entries: int = 10000,
    vector_threads: int = 1,
)

vector_dim property

vector_dim: int | None

max_score property

max_score: float

from_documents staticmethod

from_documents(
    documents: Documents,
    vectors: Vectors | None = None,
    *,
    max_score: float | None = None,
    popularity_weight: float = 0.4,
    min_word_chars: int = 3,
    max_edit_distance: int = 2,
    fuzzy_prefix_chars: int = 7,
    vector_bits: int = 4,
    build_threads: int = 1,
) -> Index

complete

complete(
    query: str,
    limit: int = 10,
    *,
    contexts: Sequence[str] | None = None,
) -> list[Suggestion]

complete_aliases

complete_aliases(
    query: str,
    limit: int = 10,
    *,
    contexts: Sequence[str] | None = None,
) -> list[AliasSuggestion]
vector_search(
    vector: Vector,
    limit: int = 10,
    *,
    contexts: Sequence[str] | None = None,
) -> list[Suggestion]
hybrid_search(
    text: str,
    vector: Vector,
    limit: int = 10,
    *,
    fusion: FusionKind = "rrf",
    rrf_k: float = 60.0,
    semantic_weight: float = 0.5,
    candidates: int | None = None,
    contexts: Sequence[str] | None = None,
) -> list[HybridSuggestion]

get

get(id: Id) -> Document | None

compact

compact(path: str | PathLike[str] | None = None) -> Segment

segments

segments() -> list[Segment]

Engine

Engine(overfetch: int = 2)

version property

version: int | None

publish

publish(updates: Mapping[str, Index | None]) -> None

sync

sync() -> int | None

Loads the database's latest version; only for engines from Database.engine().

get

get(name: str) -> Index | None

names

names() -> list[str]

complete

complete(
    query: str,
    layers: Sequence[str],
    limit: int = 10,
    *,
    contexts: Sequence[str] | None = None,
) -> list[Suggestion]

complete_aliases

complete_aliases(
    query: str,
    layers: Sequence[str],
    limit: int = 10,
    *,
    contexts: Sequence[str] | None = None,
) -> list[AliasSuggestion]
vector_search(
    vector: Vector,
    layers: Sequence[str],
    limit: int = 10,
    *,
    contexts: Sequence[str] | None = None,
) -> list[Suggestion]
hybrid_search(
    text: str,
    vector: Vector,
    layers: Sequence[str],
    limit: int = 10,
    *,
    fusion: FusionKind = "rrf",
    rrf_k: float = 60.0,
    semantic_weight: float = 0.5,
    candidates: int | None = None,
    contexts: Sequence[str] | None = None,
) -> list[HybridSuggestion]

Databases

Database

Database(
    url: str,
    options: Mapping[str, str] | None = None,
    cache_dir: str | PathLike[str] | None = None,
    *,
    min_word_chars: int = 3,
    max_edit_distance: int = 2,
    fuzzy_prefix_chars: int = 7,
    vector_bits: int = 4,
    compact_keys: bool = False,
    build_threads: int = 1,
)

versions

versions() -> list[int]

latest_version

latest_version() -> int

index_names

index_names(version: int | None = None) -> list[str]

manifest

manifest(version: int | None = None) -> dict[str, Any]

begin

begin(version: int | None = None) -> Transaction

open_index

open_index(
    name: str,
    version: int | None = None,
    *,
    max_score: float | None = None,
    popularity_weight: float = 0.4,
    short_query_chars: int = 3,
    short_query_limit: int = 100,
    short_query_cache_entries: int = 10000,
    vector_threads: int = 1,
) -> Index

engine

engine(
    *,
    group_separator: str | None = None,
    overfetch: int = 2,
    popularity_weight: float = 0.4,
    short_query_chars: int = 3,
    short_query_limit: int = 100,
    short_query_cache_entries: int = 10000,
    vector_threads: int = 1,
) -> Engine

compact

compact(
    index: str,
    *,
    fanout: int = 4,
    max_segments: int = 16,
    max_hidden_fraction: float = 0.25,
    until_done: bool = False,
) -> int | None

cleanup

cleanup(
    *,
    keep_versions: int = 10,
    older_than_seconds: float = 3600.0,
) -> dict[str, int]

acquire_lease

acquire_lease(
    name: str, owner: str, ttl_seconds: float
) -> Lease | None

submit

submit(changes: ChangeSet) -> str

pending_change_sets

pending_change_sets() -> int

Transaction

read_version property

read_version: int

append

append(
    index: str,
    documents: Documents,
    deletes: Sequence[Id] = ...,
    vectors: Vectors | None = None,
) -> None

overwrite

overwrite(
    index: str,
    documents: Documents,
    vectors: Vectors | None = None,
) -> None

drop_index

drop_index(index: str) -> None

set_max_score

set_max_score(index: str, max_score: float) -> None

set_metadata

set_metadata(key: str, value: str | None) -> None

strict

strict(strict: bool = True) -> None

max_retries

max_retries(retries: int) -> None

commit

commit() -> dict[str, Any]

ChangeSet

ChangeSet()

upsert

upsert(
    index: str,
    documents: Documents,
    vectors: Vectors | None = None,
) -> None

delete

delete(index: str, ids: Sequence[Id]) -> None

Ingestor

Ingestor(
    database: Database,
    owner: str,
    *,
    lease_ttl_seconds: float = 30.0,
    max_change_sets: int = 1000,
    compact: bool = True,
)

is_active property

is_active: bool

run_once

run_once() -> dict[str, Any]

release

release() -> None

Replica

Replica(
    database: Database,
    engine: Engine,
    *,
    popularity_weight: float = 0.4,
    short_query_chars: int = 3,
    short_query_limit: int = 100,
    short_query_cache_entries: int = 10000,
    vector_threads: int = 1,
    group_separator: str | None = None,
)

version property

version: int

sync

sync() -> int | None

Lease

generation property

generation: int

renew

renew(ttl_seconds: float) -> bool

release

release() -> None

Store

Store(
    url: str,
    options: Mapping[str, str] | None = None,
    cache_dir: str | PathLike[str] | None = None,
)

prune_cache

prune_cache(keep: Sequence[str]) -> int

get

get(key: str) -> bytes

put

put(key: str, data: bytes) -> None

put_if_absent

put_if_absent(key: str, data: bytes) -> bool

delete

delete(key: str) -> None

list

list(prefix: str = '') -> list[str]

put_segment

put_segment(key: str, segment: Segment) -> None

get_segment

get_segment(key: str) -> Segment

asyncio

AsyncDatabase

AsyncDatabase(database: Database)

A Database whose storage operations are awaitable.

database property

database: Database

versions async

versions() -> list[int]

latest_version async

latest_version() -> int

index_names async

index_names(version: int | None = None) -> list[str]

manifest async

manifest(version: int | None = None) -> dict[str, Any]

begin async

begin(version: int | None = None) -> Transaction

commit async

commit(transaction: Transaction) -> dict[str, Any]

open_index async

open_index(
    name: str, version: int | None = None, **options: Any
) -> Index

engine async

engine(**options: Any) -> AsyncEngine

submit async

submit(changes: ChangeSet) -> str

pending_change_sets async

pending_change_sets() -> int

compact async

compact(index: str, **policy: Any) -> int | None

cleanup async

cleanup(**policy: Any) -> dict[str, int]

run_ingestor async

run_ingestor(ingestor: Ingestor) -> dict[str, Any]

One round of ingestor.run_once().

AsyncEngine

AsyncEngine(engine: Engine)

An Engine that follows a database, with sync() awaitable.

sync async

sync() -> int | None

Errors

Every error derives from CompletrError. NotFoundError is also a LookupError, InvalidInputError a ValueError, and StorageError an OSError.

CompletrError

Bases: Exception

Base class of every completr error.

ConflictError

Bases: CompletrError

A concurrent commit changed what this transaction depends on.

CorruptionError

Bases: CompletrError

Stored data failed its checksum or structural validation.

NotFoundError

Bases: CompletrError, LookupError

A requested object, version or index does not exist.

InvalidInputError

Bases: CompletrError, ValueError

An argument or document was invalid.

StorageError

Bases: CompletrError, OSError

The file system or object store failed.