mrkeyoor.com_
Fri 02 Oct 14:58 UTC
AI Toolsevaluationupdated 02 Oct 2026

BettaFish review

BettaFish is a Chinese-first multi-agent system for collecting online discussion, analyzing sentiment, and producing public-opinion research reports. An English README and contribution guide exist, though the main documentation, issue activity, and bundled report templates are primarily Chinese. It combines web and social collection, private database analysis, several model roles, agent discussion, and HTML, PDF, or Markdown report generation.

Verdict

Our BettaFish run installed 207 packages into 6,564 MB, then the build failed and pytest finished with 25 passes, 3 failures, and 16 errors. Treat it as an inspectable research system for Chinese public-opinion work, not a ready public service or an unquestioned source of facts. Run a small local case first, audit the 91 known vulnerabilities, and require a human to verify every decision-bearing claim.

We ran it

Lab card: what happened when we ran BettaFishScreenshot of BettaFish (deepwiki.com/666ghj/BettaFish)
Install✓ · 117s207 packages · 6564 MB
Build✗ · 3s
Tests✗ · 44s25 passed · 3 failed · 16 errors of 44 (pytest)
Known vulns91(pip-audit)
Repo556 files~83,803 lines of source · 265.8 MB · 2 CI workflows · Dockerfile · tests dir

Answers from our run

Does BettaFish build from source?

Dependencies installed in 117 seconds (207 packages), and the build failed. We cloned commit b67928f into a clean Debian container with 3 CPUs and no project-specific setup.

Do BettaFish's tests pass?

Not all of them: 25 of 44 passed and 3 failed when we ran the project's own test command (pytest), with 16 collection errors. Some failures need services or credentials a bare container does not have.

Does BettaFish have known vulnerabilities in its dependencies?

pip-audit flagged 91 known advisories in the dependency tree at the time of our run.

Who should not use BettaFish?

Anyone planning to expose the default web app directly to the internet: the maintainer-reviewed FAQ says it lacks a complete application-level authentication boundary.

What are the alternatives to BettaFish?

GPT Researcher, MediaCrawler, AutoGen. Our BettaFish run installed 207 packages into 6,564 MB, then the build failed and pytest finished with 25 passes, 3 failures, and 16 errors.

Setup1/56,564 MB install, failed build, and 19 failed or errored tests
Docs4/5Detailed Chinese docs, an English README, and a candid FAQ
Community5/542,330 stars with an October push and same-day issue activity
Maturity2/5Broad v3 system, but our build and full test run both failed

Who it’s for

Chinese-speaking researchers testing a self-hosted public-opinion workflow across web, social, and private data.
Python teams willing to inspect crawlers, model prompts, intermediate reports, and final claims by hand.
Analysts who already have PostgreSQL or MySQL, several OpenAI-compatible model endpoints, and search API access.
Developers who want a worked multi-agent system rather than a small agent framework.

Who it’s NOT for

Anyone planning to expose the default web app directly to the internet: the maintainer-reviewed FAQ says it lacks a complete application-level authentication boundary.
Teams with a clean-test or dependency-security release gate: our run found 91 known vulnerabilities and ended with 3 failed tests plus 16 collection/setup errors.
Operators expecting one model key and one database: the sample environment defines 7 separate model roles plus Tavily and an Anspire or Bocha search service.
Buyers who need reports to be accepted without review: the project FAQ says its structure checks cannot guarantee facts, citations, or calculations.
Small disks or minimal containers: our 265.8 MB checkout expanded to 6,564 MB after installation.

Setup reality

Our run installed commit b67928f in 117 seconds, adding 207 packages and using 6,564 MB. The build exited 1 after 3 seconds. Pytest exited 1 after 44 seconds: 25 passed, 3 failed, and 16 collection/setup errors were recorded across 44 tests. Pip-audit found 91 known vulnerabilities.

The 3 test failures were Redis cache cases that could not connect to 127.0.0.1:6379. The build log tail showed invalid escape sequence warnings in Python and embedded JavaScript strings, but no final exception that explains the exit, so blaming those warnings would be guesswork.

A real run needs PostgreSQL or MySQL, Playwright Chromium, model endpoints for 7 roles, and web-search credentials. PDF export adds system libraries. Docker Compose supplies the app and PostgreSQL, while source installs also require the MediaCrawler Git submodule and platform login state for social collection.

BettaFish joins collection, analysis, and reporting in one system

BettaFish is built around a Chinese public-opinion workflow. A Query Agent searches news and the web, a Media Agent handles multimodal sources, and an Insight Agent reads a private opinion database. A Forum component shares progress and guidance among them. The Report Agent then combines their output into an intermediate document format and renders HTML, PDF, or Markdown.

The main README is Chinese. A substantial English README and an English contribution guide are available, so an English-speaking developer can understand the architecture and setup. Daily operation still points back to Chinese platform names, Chinese templates, and Chinese issue discussions. That orientation is useful for teams studying Weibo, Xiaohongshu, Douyin, or Kuaishou, and it adds friction for a team that cannot review Chinese logs and output.

Seven model roles make configuration a project of its own

The sample environment separates Insight, Media, Query, Report, MindSpider, Forum Host, and Keyword Optimizer models. Each role gets an API key, base URL, and model name. The clients use an OpenAI-compatible interface, so one vendor is not mandatory. The maintainer FAQ warns that protocol compatibility does not prove the model has enough context, multimodal support, structured output, or instruction following for its assigned job.

Search has more moving parts. Query uses Tavily, while Media chooses Anspire or Bocha. The social crawler needs Playwright Chromium, a checked-out MediaCrawler submodule, and saved login state for target platforms. PostgreSQL is recommended, though MySQL is supported. The bundled Compose file starts BettaFish and PostgreSQL, mounts reports and logs, and publishes ports 5000 plus 8501 through 8503.

What happened when we ran it

Our sandbox installed commit b67928f in 117 seconds. It pulled 207 packages and occupied 6,564 MB on disk. The checkout was already 265.8 MB, with 556 files and about 83,803 lines of source. The unprivileged Debian container had 3 CPUs, 8 GB of RAM, Python 3.12, and no secrets.

The build failed with exit code 1 after 3 seconds. Its log tail showed SyntaxWarning messages for invalid escape sequences in models_bigdata.py, html_renderer.py, and the Weibo sentiment utility. The supplied tail did not include a final exception or state that those warnings caused the failure. The honest finding is limited: this commit did not complete the build step in our fresh container.

Pytest failed after 44 seconds. It reported 25 passed, 3 failed, 18 warnings, and 16 collection/setup errors out of 44 tests. All 3 named failures came from Redis cache tests that received connection refused at 127.0.0.1:6379. Sixteen other modules failed during collection or setup, and the supplied tail does not give their underlying exceptions. We cannot reduce those errors to the missing Redis service without evidence.

Pip-audit reported 91 known vulnerabilities. Our scan also found a Dockerfile, Compose file, tests directory, and 2 CI workflow files. Those assets make the intended deployment visible, but they do not cancel the audit or failed suite. A team considering real data should triage the 91 findings against reachable code, upgrade what it can, and rerun the full checks before connecting credentials.

The default web app is for a trusted network

The maintainer-reviewed FAQ says the default web application should not be exposed directly to the public internet. It says the configuration interface lacks a complete application-level authentication boundary and could reveal database credentials or API keys or allow configuration changes. Remote access needs authentication, TLS, access control, and secret management in front of the application. Changing the bind address or firewall alone is not presented as sufficient.

That warning matters because the Compose file publishes 4 application ports and the sample environment holds credentials for 7 model roles, a database, and search providers. Keep the first evaluation on a local machine or trusted network. If the workflow survives evaluation, design the access boundary before deployment and verify that API responses, logs, report files, and saved crawler sessions do not expose secrets.

A generated report remains research material

BettaFish can check report structure and repair some malformed model output, but its FAQ says those mechanisms cannot guarantee every fact, citation, or calculation. The Insight Agent depends on real private data; an empty database or missing upstream report can prompt plausible-looking filler. Important figures and links need human review against the collected sources.

The workflow can also take time without being deadlocked. Three agents search and reflect, the Forum Host speaks, and the Report Agent generates chapters. Open issue 721 describes a run that produced no final report and stopped logging for more than 30 minutes. The project's troubleshooting advice is to inspect process health, provider retries, logs, and growing output files instead of judging by elapsed time alone.

October code activity is newer than the v3.0.0 release

GitHub showed 42,330 stars and 5 combined issues and pull requests on October 2, 2026. Two were open issues and 3 were pull requests. The repository was pushed on October 1, the same day issue 721 and a path-sanitization pull request were active. That combination shows current work even though the latest tagged release, v3.0.0, dates to December 23, 2025.

A maintainer-reviewed July FAQ explicitly warns that the release tag and main branch differ and tells users to record the exact commit. Follow that advice. BettaFish is worth studying if its Chinese social sources and multi-stage reports match the job, but our failed build, 19 failed or errored tests, and 91 audit findings make commit-level validation mandatory before real credentials or sensitive data enter the system.

Alternatives

ProjectWhat it isPick it when
GPT Researcher gh↗An autonomous research assistant focused on gathering web sources and producing cited reports.pick this instead when web research and report generation matter more than social-platform crawling and sentiment models.
MediaCrawler gh↗A crawler for collecting posts and comments from major Chinese social platforms.pick this instead when you need the collection layer and will build analysis and reporting separately.
AutoGen gh↗A framework for constructing and coordinating custom multi-agent applications.pick this instead when you want to design your own agent roles and data sources rather than adopt BettaFish's opinion-analysis pipeline.

What people are saying

  1. [velocity-scout] 666ghj/BettaFish

Sources

  1. BettaFish repository and Chinese README
  2. BettaFish English README
  3. Maintainer-reviewed BettaFish FAQ
  4. BettaFish v3.0.0 release
  5. Issue 721: report run stopped producing logs

More ai tools reviews

xialingguo-ip · reelbench-skills · SoL-Pi · flybook · crypto-rag · anything2explainer · the whole board →