TurboVec trades exact vectors for a smaller in-process index
At 2 or 4 bits per dimension, TurboVec searches compressed embedding codes with Rust SIMD kernels. Python bindings expose the same basic index. You can add vectors, search by inner product, remove them, filter by allowed IDs, and save the index to a local file. The result belongs inside an application that wants its retrieval data nearby. There is no daemon to deploy and no hosted account to open.
Compression changes what a match means. Inputs are float32, but stored codes are quantized, so a returned neighbor can differ from the nearest float32 vector. Scores use inner products. If you want cosine similarity, you normalize the vectors yourself before adding them. The supported bit widths are 2 and 4, while dimensions must be multiples of 8 and cannot exceed 16,384. Those are design boundaries, not settings to discover after ingesting a corpus.
Linear scanning suits moderate collections, not every scale
A query over 100,000 stored vectors examines all 100,000 of them. TurboVec uses compression and CPU-specific code to make that scan cheaper, but its cost still grows with the collection. The README itself recommends a graph or IVF index for hundreds of millions of vectors under strict latency targets. That candor makes the choice easier: use TurboVec when avoiding a training phase and keeping updates immediate matters more than sublinear lookup.
The project's published comparison uses 100,000-vector corpora, two embedding widths, 1,000 queries, and median results from 5 runs on ARM and x86 machines. It includes scripts and raw JSON. We did not rerun that benchmark, so its speed and recall figures remain the author's measurements. Before adoption, replay those scripts with your CPU, batch size, embedding model, filters, and target k. A fast result on a different distribution will not settle your retrieval quality.
What happened when we ran it
Our sandbox installed commit aa4a359 in 9 seconds and added 63 packages. The release build succeeded in 46 seconds. Cargo test then ran for 500 seconds and reported 883 passed, 0 failed out of 883. The unprivileged container had 3 CPUs and 12 GB of RAM. We measured build and correctness checks here, not search throughput, memory use, or recall.
The checkout contained 882 files, about 79,015 lines of source, and occupied 8 MB before installation. We found 12 CI workflow files and no Dockerfile or top-level tests directory. Rust projects commonly keep unit tests beside source, and the 883 passing cases show that the absence of a tests folder did not mean an absent suite here. Five hundred seconds is long enough that a team may split quick checks from the full pre-merge run.
Staged search can change which IDs come back
At 32,768 vectors, eligible 2-bit and 4-bit indexes can switch to staged search. The early stage examines fewer bit planes to make a shortlist, then later stages rescore that shortlist. Returned scores match the full scan for the IDs selected, but the candidate set can differ because a vector dropped from the shortlist never reaches final scoring. Environment switches let operators disable the staged path and keep the full compressed scan.
Open issue 562 gives that tradeoff a concrete shape. On one 101,000-vector PubMed collection using 768-dimensional MedCPT embeddings, the reporter found identical top-10 ID sets between staged and full scans for about 83% of queries. The issue also describes searching with a larger k or disabling the staged path as workarounds. It is one external experiment, not a universal failure rate, but it is enough reason to test the exact embedding distribution you plan to serve.
Stable IDs and incremental saves cover the local-app basics
TurboQuantIndex returns slot positions, and deletions may move the last vector into the removed slot. IdMapIndex adds caller-supplied unsigned 64-bit IDs so references survive those moves. Its allowlist search is useful when another system has already applied an access-control or time-window filter. Masks over raw slots need more care because any mutation can change which vector a slot names.
The file layer has whole-index writes and incremental sync calls. A sync writes changes since the previous sync to that path, while load accepts the resulting file. That is useful for a desktop tool or one application process. It does not add cross-host replication, concurrent service ownership, or a metadata database. One open issue reports that an integration's JSON sidecar reached 24 MB beside a 4.4 MB index at 100,000 documents, so document metadata deserves separate sizing.
Current activity is high, but there is no GitHub release entry
GitHub showed 17,330 stars and 22 open issues and pull requests on October 6, 2026. The API list split those into 21 issues and 1 pull request. The repository was pushed that same day, and the open queue contains recent work on staged-search behavior and persistence accessors. The latest-release endpoint returned no release, so release tags cannot be used as its health signal. Current commits and issue activity are the better evidence.
TurboVec is easiest to recommend when your boundary is equally clear: one Rust or Python process, a collection that tolerates a full compressed scan, and a repeatable recall check using your own queries. The clean build and 883 passing tests make that experiment credible. If the application already needs a shared endpoint, failover, or graph search, start with a vector database or HNSW library instead of building those missing layers around it.

