Static explanations and live prices take separate paths
Crypto RAG's best design decision is simple: concept documents belong in the retrieval index, while changing market values should be fetched when the question arrives. The local Indonesian corpus covers technology, trading, risks, protocols, security, metrics, regulation, and other educational topics. Price, funding, TVL, sentiment, and network routes call public sources instead of retrieving yesterday's number from an embedding.
The router decides whether a question needs retrieval, a market tool, or both. Hybrid search combines BM25 keyword matching with FAISS embeddings, then merges their rankings. An optional cross-encoder can rerank the result. Clear questions use regular-expression routes without an LLM, while complex requests can let an OpenAI-compatible model choose tools. This is a compact example of keeping deterministic paths for common work.
Six exchange quotes are not automatically one arbitrage market
The price aggregator queries Binance, OKX, Bybit, KuCoin, Kraken, and Coinbase in parallel. Four routes request USDT pairs, while Kraken and Coinbase request USD pairs. The comparison function then sorts the raw numeric prices, picks the lowest and highest, and computes a spread. It preserves the quote label in output, but it does not convert USD and USDT onto one common value first.
That is acceptable for a rough screen when the two quotes are near parity. It is insufficient for a trade signal. Fees, transfer time, withdrawal limits, liquidity, account access, and currency conversion can erase a displayed gap. The code warns that fees and transfer costs are excluded, yet the underlying quote mismatch deserves its own check before the word arbitrage enters a decision.
Order-book and slippage calculations use Binance depth rather than all 6 venues. The portfolio tool stores amount and average cost in data/portfolio.json, then values holdings from live prices. It has no exchange connection, custody, tax-lot handling, transaction import, encryption, or order execution. Think of it as a local calculator attached to a research CLI.
What happened when we ran it
Our sandbox installed commit 1b2cba4 in 15 seconds, adding 35 packages and using 37 MB. The build succeeded in 5 seconds. The checkout contained 55 files, about 4,374 source lines, and occupied 0.3 MB. We used Python 3.12 in a fresh unprivileged Debian container with 3 CPUs, 8 GB of RAM, and no secrets.
There was no tests script or target, so tests were skipped. Pip-audit reported 0 known vulnerabilities. The repository included 3 CI workflow files, no Dockerfile, and no tests directory. One workflow invokes pytest, but the measured checkout did not expose a test target for our harness. Workflow presence and an empty audit do not establish correct routing or market arithmetic.
Our run did not fetch CoinGecko data, create the FAISS indexes, download the embedding model, call exchanges, or execute the included retrieval evaluation. It also did not configure an LLM. The passing build confirms that the Python project prepared successfully in our container; it is not a benchmark for answer quality, data freshness, or market coverage.
Optional synthesis sends live numbers through the LLM
The README says market numbers never pass through an LLM. That is true for rule-based output, but the optional synthesis path does pass a context containing live values into /chat/completions. The system prompt tells the model to reproduce market numbers exactly, name the live source and time, and admit when context is insufficient. Those instructions reduce risk but cannot make generated text immune to transcription or omission errors.
Users can avoid that layer with the extractive mode. That is the safer choice when the number itself matters more than conversational wording. If synthesis is enabled, compare the answer with the structured tool result and retain its timestamp. An Indonesian disclaimer at the end of a generated answer does not correct a misplaced decimal or an unavailable source.
The agent route also lets a model select more than one tool for a complex query. Its prompt refuses investment advice and predictions, which is sensible. Tool selection remains another behavior to test. A model can choose the wrong symbol, skip a needed source, or combine observations that were sampled at different moments.
The retrieval evaluation is useful, but it is not a release gate
eval_rag.py contains a golden set of Indonesian questions and checks whether an expected section appears near the top of retrieval results. It reports hit rate and reciprocal rank, using the same knowledge-search path as the application. That is a better starting point than judging retrieval from a few hand-picked chats. The repository does not publish a measured result in the README or run this evaluation as the lab test target.
Three GitHub workflows are present, but two are generic release templates. The SLSA workflow still creates placeholder artifacts, and the Python publishing workflow retains comments asking the maintainer to add real build steps. There is no tagged release or package configuration shown in the top-level tree. These files should not be read as a finished supply-chain process.
A September 13 push leaves a short maintenance record
GitHub showed 465 stars, 0 combined issues and pull requests, and a last push on September 13, 2026. The repository was created on September 5 and had no published release. Zero open issues can mean a quiet tracker; with such a short history, it does not prove that the many data routes are stable or actively supported.
Crypto RAG is worth reading and adapting if Indonesian crypto education is the use case. The static-versus-live split is sound, the source list is visible, and the no-key rule mode keeps the first experiment accessible. Before trusting the numbers, normalize quote currencies, add fixtures for every provider, and turn the retrieval and routing checks into repeatable release gates.

