Qlib is a 75,114-line research system, not a backtest helper
Our clone contained 619 files and about 75,114 lines of source in an 8.1 MB checkout. That size matches Qlib's scope. It stores and transforms market data, trains models, records experiments, turns predictions into portfolio decisions, simulates execution, and produces analysis. Components can be used separately. The main attraction is a shared workflow from raw data to an evaluated strategy. A team choosing Qlib is adopting research infrastructure, not adding one function to a notebook.
The repository includes 6 CI workflow files, a Dockerfile, and a tests directory. Its model examples span tree methods, neural networks, transformers, market-dynamics work, and reinforcement learning. Alpha158 and Alpha360 are prepared feature sets for equity research, while strategies and executors cover portfolio and order decisions. Qlib supplies machinery for asking financial questions. It does not supply a dependable trading edge, and an example result still needs independent checks for data leakage, costs, benchmark choice, and out-of-sample behavior.
What happened when we ran it
Our sandbox install succeeded in 107 seconds and brought in 214 packages, leaving 964 MB on disk. Building the project then succeeded in 36 seconds. Our test method ran commit 79633dd in an unprivileged Debian container with 3 CPUs, 8 GB of RAM, and the python3.12-bookworm image. That is a workable build path. It is much heavier than the README's single pip install pyqlib command makes the first encounter feel. The checkout itself was only 8.1 MB; dependencies accounted for nearly all of the installed footprint.
The test step failed with exit code 1 after 230 seconds. Pytest reported 51 passed, 10 failed, 1 skipped, and 9 collection or setup errors; the harness recorded 70 tests and 38,363 warnings. Two workflow failures ended with MLflow exceptions saying its filesystem tracking backend was in maintenance mode. Two setup errors named missing test_all_flow_mlruns and test_contrib_mlruns directories. Six reinforcement-learning test modules and tests/test_pit.py also appeared in the error list. The log tail does not establish causes for those seven errors, so we will not invent one. Pip-audit also reported 1 known vulnerability.
Python 3.12 works, but usable data still comes from elsewhere
Python 3.12 installed and built Qlib in our container, which agrees with the README's support table for Python 3.8 through 3.12. A useful experiment still starts with a separate data job. The official Qlib dataset is temporarily disabled under a stricter data-security policy. The quick start points to a community-maintained archive, says its public data came from Yahoo Finance, and warns that the data may be imperfect. Qlib includes collectors, conversion scripts, daily updates, and a health checker. Those tools cannot establish provenance or repair a weak source.
The 214-package environment is only the common starting point. Qlib's README says individual baselines have different dependencies and calls out TFT as requiring TensorFlow 1.15 with Python 3.6 or 3.7. Its multi-model runner creates a separate virtual environment per model, runs on Linux only, and cannot execute repeated instances of the same model in parallel. Conda is recommended because source installs may otherwise miss headers. Apple Silicon users may need Homebrew's OpenMP package before LightGBM will build. Reproducing the model zoo is an environment-management project of its own.
Ten failed tests keep the examples from being contracts
Qlib's suite produced 10 failed tests and 9 collection or setup errors, which should change how a team uses it. The configuration-driven qrun path is appealing because one YAML file can connect a data handler, model, recorder, strategy, and backtest. Yet a workflow that completes is not automatically a sound financial experiment. Pin the commit, dependency set, dataset snapshot, trading calendar, universe, transaction costs, and benchmark construction. Recalculate important metrics outside Qlib before a research result influences capital.
The project has 6 CI workflow files. Their coverage boundary matters. Issue #1278, opened in 2022 and updated on 2026-08-25, asks for the main workflow_by_code.ipynb example to be executed in CI because existing checks can miss errors in its report. PR #2241 proposes that job and was still open. Readers can learn from the notebook, but Qlib did not yet treat its full execution as a merged continuous test.
July code and August issue activity show a live, busy project
GitHub showed 469 open issues and PRs combined, while the last repository push was 2026-07-23 and contributions were still being updated on 2026-08-25. That combined count is evidence of attention and a sizable review queue. The latest tagged release was v0.9.7 on 2025-08-15; it added Parquet support and data-layer changes alongside fixes. The release date alone does not make Qlib abandoned because contribution activity continued into August 2026. Its 47,938 stars show reach, not correctness.
A 964 MB environment buys breadth that smaller tools avoid
The 964 MB environment buys one vocabulary for data handlers, ML experiments, portfolio logic, execution simulation, and reports. vectorbt is easier to justify for vectorized strategy sweeps. FinRL is more focused when reinforcement learning is the whole question. LEAN is the better comparison when live brokerage operation matters. None is a drop-in replacement, because each draws the system boundary differently. Qlib's value is the breadth inside that boundary, and its cost is the amount of code and method you must audit.
Our run spent 107 seconds installing and 230 seconds on a suite that did not pass. That is still enough evidence to recommend Qlib for capable quant ML teams: the build works, the research surface is unusually broad, and the documentation gives concrete workflows. It is not a sensible default for a developer who only needs a quick backtest or expects bundled market data. Start with one model and one independently checked dataset, then expand only after the test failures and 1 audit finding are understood in your environment.

