BettaFish joins collection, analysis, and reporting in one system
BettaFish is built around a Chinese public-opinion workflow. A Query Agent searches news and the web, a Media Agent handles multimodal sources, and an Insight Agent reads a private opinion database. A Forum component shares progress and guidance among them. The Report Agent then combines their output into an intermediate document format and renders HTML, PDF, or Markdown.
The main README is Chinese. A substantial English README and an English contribution guide are available, so an English-speaking developer can understand the architecture and setup. Daily operation still points back to Chinese platform names, Chinese templates, and Chinese issue discussions. That orientation is useful for teams studying Weibo, Xiaohongshu, Douyin, or Kuaishou, and it adds friction for a team that cannot review Chinese logs and output.
Seven model roles make configuration a project of its own
The sample environment separates Insight, Media, Query, Report, MindSpider, Forum Host, and Keyword Optimizer models. Each role gets an API key, base URL, and model name. The clients use an OpenAI-compatible interface, so one vendor is not mandatory. The maintainer FAQ warns that protocol compatibility does not prove the model has enough context, multimodal support, structured output, or instruction following for its assigned job.
Search has more moving parts. Query uses Tavily, while Media chooses Anspire or Bocha. The social crawler needs Playwright Chromium, a checked-out MediaCrawler submodule, and saved login state for target platforms. PostgreSQL is recommended, though MySQL is supported. The bundled Compose file starts BettaFish and PostgreSQL, mounts reports and logs, and publishes ports 5000 plus 8501 through 8503.
What happened when we ran it
Our sandbox installed commit b67928f in 117 seconds. It pulled 207 packages and occupied 6,564 MB on disk. The checkout was already 265.8 MB, with 556 files and about 83,803 lines of source. The unprivileged Debian container had 3 CPUs, 8 GB of RAM, Python 3.12, and no secrets.
The build failed with exit code 1 after 3 seconds. Its log tail showed SyntaxWarning messages for invalid escape sequences in models_bigdata.py, html_renderer.py, and the Weibo sentiment utility. The supplied tail did not include a final exception or state that those warnings caused the failure. The honest finding is limited: this commit did not complete the build step in our fresh container.
Pytest failed after 44 seconds. It reported 25 passed, 3 failed, 18 warnings, and 16 collection/setup errors out of 44 tests. All 3 named failures came from Redis cache tests that received connection refused at 127.0.0.1:6379. Sixteen other modules failed during collection or setup, and the supplied tail does not give their underlying exceptions. We cannot reduce those errors to the missing Redis service without evidence.
Pip-audit reported 91 known vulnerabilities. Our scan also found a Dockerfile, Compose file, tests directory, and 2 CI workflow files. Those assets make the intended deployment visible, but they do not cancel the audit or failed suite. A team considering real data should triage the 91 findings against reachable code, upgrade what it can, and rerun the full checks before connecting credentials.
The default web app is for a trusted network
The maintainer-reviewed FAQ says the default web application should not be exposed directly to the public internet. It says the configuration interface lacks a complete application-level authentication boundary and could reveal database credentials or API keys or allow configuration changes. Remote access needs authentication, TLS, access control, and secret management in front of the application. Changing the bind address or firewall alone is not presented as sufficient.
That warning matters because the Compose file publishes 4 application ports and the sample environment holds credentials for 7 model roles, a database, and search providers. Keep the first evaluation on a local machine or trusted network. If the workflow survives evaluation, design the access boundary before deployment and verify that API responses, logs, report files, and saved crawler sessions do not expose secrets.
A generated report remains research material
BettaFish can check report structure and repair some malformed model output, but its FAQ says those mechanisms cannot guarantee every fact, citation, or calculation. The Insight Agent depends on real private data; an empty database or missing upstream report can prompt plausible-looking filler. Important figures and links need human review against the collected sources.
The workflow can also take time without being deadlocked. Three agents search and reflect, the Forum Host speaks, and the Report Agent generates chapters. Open issue 721 describes a run that produced no final report and stopped logging for more than 30 minutes. The project's troubleshooting advice is to inspect process health, provider retries, logs, and growing output files instead of judging by elapsed time alone.
October code activity is newer than the v3.0.0 release
GitHub showed 42,330 stars and 5 combined issues and pull requests on October 2, 2026. Two were open issues and 3 were pull requests. The repository was pushed on October 1, the same day issue 721 and a path-sanitization pull request were active. That combination shows current work even though the latest tagged release, v3.0.0, dates to December 23, 2025.
A maintainer-reviewed July FAQ explicitly warns that the release tag and main branch differ and tells users to record the exact commit. Follow that advice. BettaFish is worth studying if its Chinese social sources and multi-stage reports match the job, but our failed build, 19 failed or errored tests, and 91 audit findings make commit-level validation mandatory before real credentials or sensitive data enter the system.

