PageIndex added 543 GitHub stars in the daily snapshot behind this article by moving the expensive part of retrieval-augmented generation into the search path. Instead of embedding chunks once and looking them up by similarity, it asks a language model to walk a document tree for each question. Developers can drop the vector database, but retrieval now depends on model judgment and the time and money each query takes.
That exchange is more interesting than the phrase "vectorless RAG" suggests. The MIT-licensed PageIndex repository has turned a familiar infrastructure complaint into a different systems problem: whether a model can navigate the shape of a long document more reliably than a similarity score can recover the right passage. A star is only a bookmark, so the one-day rise does not establish adoption. It does show that many developers want to test the premise.
A tree takes the place of the vector index
A conventional RAG pipeline usually divides a document into passages, converts them into embedding vectors and retrieves the nearest matches to a query. PageIndex builds a hierarchy instead. Headings and nested sections become nodes tied to pages. The system writes summaries for those nodes, then gives a chat model tools to inspect the tree and open the page content it judges relevant.
The distinction matters when a question and its answer use different language. A similarity search may favor a passage that repeats the query's terms. Tree search can use the document's organization and the conversation around the question to choose a section whose wording is less obvious. The retrieval trail is also inspectable because the model's route ends at named nodes and pages rather than an unlabeled list of nearby vectors.
PageIndex's current local path uses its Flash indexer. According to the v0.2.20 release notes, Flash derives the initial tree from layout statistics and trusted PDF bookmarks. A model still writes node summaries, and the default optimization can ask an LLM to expand the tree. Version 0.2.20, released on September 28, runs summaries while tree expansion is still underway. The project reports that this cut one 222-page indexing run from 98 seconds to 73, and a 758-page run from 174 seconds to 137. Those are the project's measurements, not independent benchmarks.
The design is easiest to picture as a generated table of contents with an active reader attached. The index supplies the map. The chat model chooses which branch to follow, reads the underlying pages and forms the answer. Removing embeddings does not remove retrieval. It changes the retrieval algorithm from a distance calculation into a sequence of model decisions.
Vectorless still involves two model jobs
PageIndex separates indexing from answering. In local mode, one model summarizes and refines the tree while another searches it at query time. The setup documentation recommends a cheaper model for indexing and a stronger one for chat, because answer quality depends on how well the latter reasons over the hierarchy. The same client can route those calls to OpenAI, Anthropic, OpenRouter or an OpenAI-compatible endpoint.
Local mode means that PageIndex stores the tree and source documents on the user's machine. It does not automatically mean that document content stays there. A hosted model endpoint still receives the material needed for summaries and answers. A team that requires fully local processing would need to point the client at a compatible local model and verify that no fallback route sends data elsewhere.
The absence of a vector database removes one service to run and populate. It also makes the chat model part of retrieval rather than only the final answer generator. A weak retrieval model can choose the wrong branch before it ever sees the relevant passage. A stronger model may recover more answers while increasing cost. This is a recurring expense because the tree is reused but the model searches it again for each question.
That trade can be sensible for long reports whose section structure carries useful meaning. It is less attractive for workloads that demand a fixed retrieval result, very low latency or operation without model calls. PageIndex can return citations, yet a citation proves where an answer came from after the search. It does not make the route through the tree deterministic.
The benchmark is useful within a narrow lane
The project's open benchmark repository publishes its PDFs, questions, runner and machine-readable results. It contains 62 lookup questions over 34 PDFs totaling 1,945 pages. Every answer is stated in running text. The set excludes charts, tables, figures, counting and arithmetic so that the test concentrates on finding and reading a passage.
It also excludes documents that Flash could not index. The maintainers say this directly: each included PDF was checked for successful indexing, and the benchmark says nothing about documents outside that accepted set. That choice makes the results reproducible while preventing the headline scores from describing arbitrary PDFs.
Within that lane, the model choice is visible in the numbers. Using the same trees, the published results range from 85.5 percent accuracy at an average $0.0031 per question with the least expensive configuration to 100 percent at about $0.08 per question with stronger models and reasoning settings. Several configurations reached 96.8 percent or better. These are project-run figures with model-based judging, so they are a reason to reproduce the test, not a substitute for an evaluation on a team's own documents.
A separate open issue shows why ingestion belongs in that evaluation. In issue 518, a user testing Flash on FinanceBench reported that some files were missing their opening pages and described coverage of only 50 percent. One user's result cannot establish a system-wide rate. Version 0.2.20 later added page-level fallbacks when no hierarchy is detected. The open issue does not say that its specific failure has been fixed.
Our review of PageIndex covers the setup reality, including a 106-package install, 280 MB environment, 703 reported test passes and one dependency-audit finding. That run checked package health rather than answer quality. The distinction is worth keeping: a clean test suite can show that the software behaves as its authors expect without showing that it retrieves the right evidence from a new document collection.
Local and cloud have different document boundaries
The open local client currently targets text-based PDFs and returns page-level citations. Scans, charts and image-heavy files fall outside that simple path because Flash reads the PDF text layer and layout. PageIndex Cloud adds OCR, image understanding, block-level citations, managed storage, folders, metadata and an MCP server, according to the local-versus-cloud documentation.
That division makes "open source" and "self-hosted" separate questions. Developers can run the Python client, local storage and tree search code without a PageIndex API credential. Some document types and operational features still require the hosted product. Cloud documents can be searched with the customer's chosen model, which creates two data boundaries to inspect: PageIndex for document processing and storage, then the model provider for retrieval and answering.
The package metadata labels the SDK as alpha and requires Python 3.10 or newer. The release history is moving quickly. Version 0.2.19 renamed resolve_citations to get_citations, while assigning the old name a different return shape. Version 0.2.20 carried that change forward. Pinning the dependency and testing citation handling before an upgrade is ordinary care here, especially for an application that displays source evidence to users.
The repository has no documented security reporting route today. Issue 240, opened in April and still open at reporting time, asks how researchers should report a vulnerability without a SECURITY.md file. The missing route raises the cost of responsible disclosure for software that may process sensitive financial or medical documents.
Test the retrieval decision, not the slogan
A fair PageIndex trial should use questions whose answers are already known and documents that resemble production files. Run the same set through the current vector or hybrid retriever, then record whether each system found the supporting passage, whether the citation points to the right page, how long the query took and what it cost. Add malformed, scanned and unusually structured PDFs instead of limiting the test to documents that index cleanly.
The result may favor different systems for different corpora. A legal filing with dependable headings can reward tree navigation. A large pool of short, loosely structured notes may still suit embeddings or a hybrid search. PageIndex itself now has a cloud-only file tree for searching across many documents, which means the single-document hierarchy does not erase the need to select the right file at corpus scale.
The 543-star day has earned PageIndex a close look. Production evidence will come from wider, independent comparisons. The project also has to settle its open coverage reports. Its alpha API needs to stabilize, and the boundary between local and cloud must become clearer. The practical question is specific: on your documents, does a model walking the tree find evidence that similarity search misses often enough to pay for that model on every query?