mrkeyoor.com_
Sat 03 Oct 15:34 UTC
AI Toolsevaluationupdated 03 Oct 2026

production-agentic-rag-course review

Production Agentic RAG Course is a seven-week Python course built around an arXiv research assistant. It teaches the whole path from PostgreSQL and keyword search to hybrid retrieval, a local Ollama model, tracing, caching, LangGraph decisions, and a Telegram bot.

Verdict

Our run installed 236 packages and used 6,297 MB, but 14 tests failed and 23 hit collection or setup errors, so this repo is better as a guided lab than as a clean production starter. Use it if the seven-week progression matches how you learn and you can debug the supplied environment. Do not treat the word production in the title as a deployment approval.

We ran it

Lab card: what happened when we ran production-agentic-rag-courseScreenshot of production-agentic-rag-course (github.com/jamwithai/production-agentic-rag-course)
Install✓ · 155s236 packages · 6297 MB
Build✓ · 6s
Tests✗ · 57s58 passed · 14 failed · 23 errors of 95 (pytest)
Known vulns0(pip-audit)
Repo160 files~10,759 lines of source · 11.5 MB · 0 CI workflows · Dockerfile · tests dir

Answers from our run

Does production-agentic-rag-course build from source?

Dependencies installed in 155 seconds (236 packages), and the build succeeded in 6 seconds. We cloned commit 424a0eb into a clean Debian container with 3 CPUs and no project-specific setup.

Do production-agentic-rag-course's tests pass?

Not all of them: 58 of 95 passed and 14 failed when we ran the project's own test command (pytest), with 23 collection errors. Some failures need services or credentials a bare container does not have.

Does production-agentic-rag-course have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use production-agentic-rag-course?

Developers seeking a small RAG starter: our install used 6,297 MB before the Compose services and model data were running.

What are the alternatives to production-agentic-rag-course?

RAG From Scratch, RAG Techniques, LLM Zoomcamp. Our run installed 236 packages and used 6,297 MB, but 14 tests failed and 23 hit collection or setup errors, so this repo is better as a guided lab than as a clean production starter.

Setup2/56,297 MB installed and the full test command failed
Docs4/5Seven staged guides, with a Python version mismatch
Community4/59,302 stars with issue and pull request activity in 2026
Maturity3/5Broad working stack, but our 95-test run was not green

Who it’s for

Python developers who learn best by extending one working system across seven weekly stages.
AI engineers who want to study BM25 before adding embeddings and generated answers.
Teams evaluating how OpenSearch, Airflow, Ollama, Redis, Langfuse, and LangGraph fit together.
Learners with Docker resources and time to diagnose a large local stack.

Who it’s NOT for

Developers seeking a small RAG starter: our install used 6,297 MB before the Compose services and model data were running.
Teams that require a green test suite before adoption: 14 tests failed and 23 ended in collection or setup errors in our run.
Users who need a broad Python support window: the README says Python 3.12+, while pyproject.toml requires 3.12 and rejects 3.13.
Production operators expecting secure defaults: the supplied local Compose file disables OpenSearch security and includes development credentials for PostgreSQL and Langfuse.
Learners who want a credential-free path through every week: hybrid search requires a Jina API key, and the Telegram stage needs a bot token.

Setup reality

Our sandbox installed commit 424a0eb in 155 seconds, adding 236 packages and using 6,297 MB on disk. The build passed in 6 seconds. Pytest failed after 57 seconds: 58 passed, 14 failed, and 23 collection or setup errors were reported out of 95. Pip-audit found 0 known vulnerabilities.

The application stack needs Docker Compose plus PostgreSQL, OpenSearch, Airflow, Ollama, Redis, and Langfuse. Week 4 requires a Jina API key, Week 7 needs a Telegram bot token, and Langfuse keys are optional. The README asks for at least 8 GB of RAM and 20 GB of free disk.

The package accepts Python 3.12 but excludes 3.13, despite the README's broader 3.12+ label. Our log tail repeatedly said async test functions were not natively supported. It does not establish why the async test support failed, so users should reproduce the suite before changing its test dependencies.

Seven weeks turn one paper curator into an agent

The course starts with a FastAPI service, PostgreSQL, OpenSearch, Airflow, and Ollama. Each week changes the same arXiv paper curator instead of presenting a fresh toy. Week 2 fetches and parses papers. Week 3 builds BM25 search. Week 4 adds section-aware chunks, Jina embeddings, and reciprocal rank fusion. The generated-answer layer arrives in Week 5, followed by Redis and Langfuse in Week 6.

Week 7 is where the title earns the word agentic. A LangGraph workflow validates the query, retrieves documents, grades them, rewrites weak queries, and generates an answer. A Telegram bot exposes that path on a phone. This sequence is the repo's best teaching choice: you can see why each extra service exists because the previous week's limits are visible first.

The local stack needs 20 GB before it needs an agent

The README asks for Python 3.12, Docker Compose, at least 8 GB of RAM, and more than 20 GB of free disk. Its Compose file then brings up the API, two PostgreSQL databases, OpenSearch and its dashboard, Airflow, Ollama, Redis, ClickHouse, Langfuse, and MinIO-related services. That is a useful systems lesson. It is also a lot of moving parts for someone who only wants to understand retrieval.

Our checkout was 11.5 MB across 160 files and about 10,759 lines of source. Installing its Python environment expanded disk use to 6,297 MB before Docker images, database volumes, downloaded papers, or Ollama model data entered the picture. The README's 20 GB warning is therefore believable. A learner on a small laptop should choose a narrower alternative or run one week at a time.

What happened when we ran it

Our sandbox installed commit 424a0eb in 155 seconds. It pulled 236 packages and occupied 6,297 MB on disk. The build succeeded in 6 seconds, and pip-audit found 0 known vulnerabilities. We ran this in a fresh unprivileged Debian container with 3 CPUs, 8 GB of RAM, Python 3.12, and no secrets.

The test command failed after 57 seconds. Pytest reported 58 passed, 14 failed, and 23 collection or setup errors out of 95. The tail showed failures in the arXiv client, metadata fetcher, and PDF parser tests. Each quoted failure said async functions are not natively supported. That is the observed error, not a diagnosis of which dependency or setting should change.

A green build and 58 passing tests still tell us a fair amount of code loaded and ran. They do not cancel 14 failures or 23 errors. The repository has a tests directory but no CI workflow files, so a prospective user cannot point to a visible GitHub Actions run for the commit we tested. Reproduce the suite on your own branch before using these services as a base.

Python support is narrower than the README says

The README labels the requirement as Python 3.12+. The actual package metadata requires >=3.12,<3.13, which means Python 3.13 does not qualify. That small mismatch matters because the setup uses uv and a large dependency set. Let uv create a 3.12 environment rather than asking it to solve against whichever interpreter happens to be installed.

Configuration also grows with the weekly path. Jina credentials are required for hybrid search. Telegram needs a BotFather token. Langfuse keys are optional, while its local deployment needs secrets and an encryption key. Ollama keeps answer generation local, but model weights still have to be pulled. The supplied environment file explains these variables, including warnings to replace development values.

The Compose defaults belong on a learning machine

OpenSearch runs with its security plugin disabled. The Compose file contains fixed local database passwords and initializes Langfuse with a sample admin email and password. Those choices reduce friction in a course environment. They also make the stack unsuitable for an exposed host until you replace credentials, enable the controls you need, and decide which dashboards should be reachable.

The repo calls the system production-grade, but the safer reading is production-shaped. You get queues, caching, tracing, health checks, a search engine, and separate persistence layers. You do not get a finished security or deployment policy. The course teaches the pieces and their connections. Your organization still owns authentication, network boundaries, backups, upgrades, cost limits, and incident response.

Current activity exists outside the last release

GitHub showed 9,302 stars and 29 combined open issues and pull requests on October 3, 2026. The default branch was last pushed on June 5. The latest tagged release, week7.0, was published on November 26, 2025. Those dates alone could look stale, but the open queue included an issue from September 27 and a pull request from October 3.

The queue split was 13 open issues and 16 open pull requests when checked. One July issue reports large CUDA dependencies entering the Airflow image through Docling, which is relevant beside our 6,297 MB install. The activity shows that people still work around the project. It does not give us a passing suite for commit 424a0eb. Treat the repo as course material worth studying, then promote only the parts your own tests can defend.

Alternatives

ProjectWhat it isPick it when
RAG From ScratchA compact set of notebooks and videos that explains core RAG ideas in smaller pieces.pick this instead when you want retrieval concepts without operating the full service stack.
RAG Techniques gh↗A notebook collection that isolates advanced retrieval methods for direct comparison.pick this instead when your goal is to test retrieval techniques rather than build one staged application.
LLM ZoomcampA free course on building and evaluating an LLM application around a knowledge base.pick this instead when you want a broader course with evaluation and operations taught outside this arXiv stack.

What people are saying

  1. [github-trending] jamwithai/production-agentic-rag-course

Sources

  1. Production Agentic RAG Course repository
  2. Course README
  3. Python package configuration
  4. Docker Compose stack
  5. Week 7 release
  6. Airflow CUDA dependency issue

More ai tools reviews

chandra · GPT-as-Policy · NeuralScreen · mural · recurrent-looped-tranformer · gpu-time · the whole board →