mrkeyoor.com_
Tue 29 Sept 16:54 UTC
AI Toolsevaluationupdated 26 Aug 2026

VidBee review

VidBee is a desktop media library that downloads video or audio, imports local files, and turns speech into searchable transcripts on your computer. It can label speakers, jump from text to a timestamp, and send a transcript to a chosen AI provider for summaries, translation, questions, or custom prompts.

+17stars / 7d
Verdict

Our VidBee install pulled 1,400 packages and occupied 1,579 MB, then built in 17 seconds with no test target available, so this is a capable trial with a large trust gap. Use it when searchable local transcripts and download organization are worth one substantial desktop app. Keep a fallback downloader nearby, and do not send sensitive transcripts to a hosted model provider by accident.

We ran it

Lab card: what happened when we ran VidBeeScreenshot of VidBee (vidbee.org)
Install✓ · 35s1433 packages · 1628 MB
Build✓ · 16s
Testsn/ano test script
Repo879 files~136,593 lines of source · 12.9 MB · 8 CI workflows

Answers from our run

Does VidBee build from source?

Dependencies installed in 35 seconds (1433 packages), and the build succeeded in 16 seconds. We cloned commit feeea6b into a clean Debian container with 3 CPUs and no project-specific setup.

Does VidBee have tests you can run?

Not through a standard command: the project exposes no test script or target that our harness could run.

Who should not use VidBee?

Anyone who needs downloads from every advertised site to work consistently: current issues show HTTP 403 or extractor failures for Rumble, YouTube, TikTok, and Instagram.

What are the alternatives to VidBee?

yt-dlp, Whisper, LosslessCut. Our VidBee install pulled 1,400 packages and occupied 1,579 MB, then built in 17 seconds with no test target available, so this is a capable trial with a large trust gap.

Setup3/5Build is quick, but 1,579 MB and model downloads are substantial
Docs4/5Workflows, privacy boundary, formats, desktop, and web are clear
Community4/510,386 stars with active August 2026 releases and reports
Maturity3/5v2.0.2 is active, but no test target and downloader bugs remain

Who it’s for

Researchers and journalists who need searchable transcripts tied to exact playback positions.
Podcast or lecture collectors who want downloads, RSS intake, local speech recognition, and exports in one app.
Users who prefer local transcription but still want a choice of hosted or local AI providers for transcript analysis.
NAS owners who can run the newer web and API services with persistent download storage.

Who it’s NOT for

Anyone who needs downloads from every advertised site to work consistently: current issues show HTTP 403 or extractor failures for Rumble, YouTube, TikTok, and Instagram.
Teams that require an automated test signal before deployment: our measured checkout had no test script or target, so tests were skipped.
Users with little local disk or bandwidth: our dependency install alone used 1,579 MB before media files or speech models.
Privacy-sensitive users who select a hosted AI provider without reviewing its terms: the README says prompts and transcript content go to that provider.
Operators seeking a settled appliance: the README calls VidBee active development, and issue 463 reports the v2.0.2 download engine unavailable on an Intel Mac.

Setup reality

Our sandbox installed 1,400 pnpm packages in 54 seconds and used 1,579 MB on disk. The build succeeded in 17 seconds. There was no tests script or target, so the test step was skipped rather than passed.

The desktop path uses yt-dlp and FFmpeg, while transcription adds a local Whisper, SenseVoice, Parakeet, or Qwen3-ASR model. AI prompts need a local endpoint such as Ollama or LM Studio, or credentials for the chosen hosted provider.

This is a 687-file monorepo with about 89,057 source lines, 8 CI workflow files, a Compose file, and no Dockerfile or tests directory in our scan. Model downloads, media storage, cookies for restricted sites, and platform-specific engine behavior are the practical costs beyond a successful build.

Downloads become a searchable media library

VidBee combines jobs that usually live in separate tools. It downloads video, audio, playlists, and channels through a queue, or accepts a local media file. Finished items can be transcribed, searched by spoken words or speaker names, and played from a clicked timestamp. RSS subscriptions can watch a feed and apply their own download directory, filename template, tags, keyword filter, and automatic-download choice.

Our checkout at commit 5155e90 held 687 files and roughly 89,057 source lines across desktop, browser extension, API, web, and shared packages. Support for more than 1,000 sites comes from extractors that must keep up with changing pages. VidBee gives those commands a queue and library, but a blocked extractor remains dependent on the source site.

Transcription stays local until an AI prompt runs

Speech recognition can use local Whisper, SenseVoice, Parakeet, or Qwen3-ASR model families. VidBee recommends models based on the computer and language, lets the user switch downloaded models, and can start transcription after a download. Speaker detection can infer a count or accept a fixed number. Search, timestamp navigation, relabeling, copying, and text or Markdown export remain useful without an AI provider.

Our dependency install used 1,579 MB before any media file or speech model was added. That figure matters on a laptop more than the 54-second install time. Model size and transcription speed were not present in the supplied lab result, so we will not guess at either. Plan separate storage for the pnpm workspace, model cache, download library, captions, and generated exports.

Hosted prompts send transcript content to the selected provider

VidBee can send transcript prompts to OpenAI, Anthropic, Google, DeepSeek, Groq, Azure, Hugging Face, OpenRouter, or xAI. It also accepts a custom base URL and model ID. The README states the boundary plainly: transcription happens on the computer, while a prompt and transcript content go to the selected provider. Ollama and LM Studio can keep that second step local.

The build completed successfully in 17 seconds, but that result did not exercise any provider connection or API key. Each hosted service has its own credentials, retention terms, limits, and bill. For interviews or unpublished reporting, an administrator should lock down approved providers rather than leaving each user to turn a local transcript into a cloud request.

What happened when we ran it

Our sandbox installed 1,400 pnpm packages in 54 seconds, occupying 1,579 MB, and built the monorepo in 17 seconds. We used an unprivileged Node 22 Debian container with 3 CPUs, 8 GB of RAM, and no secrets. The checkout had 8 CI workflow files, a Compose file, monorepo workspaces, no Dockerfile, and no tests directory. Nothing in the supplied install or build result failed.

There was no test script or target, so our harness skipped tests. That is different from a passing suite. We did not validate a desktop package, run yt-dlp against a live site, download an ASR model, transcribe audio, identify speakers, or call an AI provider. The 17-second build shows the source compiled; it does not establish that the media workflow works from download through export.

Site extractors can fail even when the app builds

Issue 461 records a Rumble request ending in HTTP 403 from the underlying extractor. Issue 449 shows a TikTok page returning an unexpected response even with the reported yt-dlp version current. Other August reports cover YouTube 403 responses, Instagram extraction failure, and Windows browser-cookie decryption through DPAPI. These logs show different failures, so there is no single cause to assign.

A 1,400-package install does not prove that an extractor works. Keep the generated yt-dlp command from a failure and try it independently when diagnosing. Some sources need current cookies or a JavaScript runtime; others may reject the request anyway. If a site is essential, pin a real sample URL in a scheduled check and retain yt-dlp as a direct fallback.

Desktop and server deployment solve different problems

The desktop app is the complete personal workflow shown in the screenshots. The monorepo also contains a Fastify API, a TanStack Start web client, and a shared downloader core. pnpm run start:web starts the API and web app together. Compose examples expose ports 3100 and 3000, persist a SQLite history database, and mount storage that can point at a NAS.

Our scan found a Compose file but no Dockerfile at the repository root, although the README names published GHCR images. Treat the desktop and server routes as different products. A NAS can solve shared storage, while local ASR and playback may still belong on a user's machine. Confirm where models execute before assuming the web stack reproduces every desktop feature.

v2.0.2 is active, while regression proof is thin

GitHub records the last push and release v2.0.2 on August 23, 2026. That release fixed a Windows transcript-player crash. The repository had 32 open issues excluding pull requests and 10,386 stars when fetched. Issue 463 then reported that v2.0.2 on macOS 12.7.1 could not prepare a download engine despite locally installed FFmpeg and yt-dlp.

The successful 17-second build makes VidBee worth testing for work that begins with saved media and ends with notes tied to the recording. The missing test target keeps it from earning unattended trust. Try representative files and sites on the exact operating system, budget for 1,579 MB plus models and media, and decide whether transcript prompts may leave the computer before adding provider keys.

Alternatives

ProjectWhat it isPick it when
yt-dlp gh↗A command-line downloader with wide site support and detailed format controls.pick this instead when downloading is the whole job and scripts are preferable to a desktop library.
Whisper gh↗A local speech-recognition model and Python package without a media-library interface.pick this instead when you want to script transcription yourself and do not need VidBee's queue or transcript UI.
LosslessCutA desktop app for quick lossless trimming and rearranging of audio and video files.pick this instead when editing local media matters more than downloading, transcription, or AI analysis.

What people are saying

  1. [github-trending] nexmoe/VidBee

Sources

  1. VidBee README
  2. VidBee v2.0.2 release
  3. Issue 463: download engine unavailable on macOS
  4. Issue 461: Rumble extractor HTTP 403
  5. Issue 449: TikTok extractor failure

More ai tools reviews

voltagent · InferenceX · Bonsai-demo · qwen-audio-agent · wechat-intelligence-hub · dlss5-visual-enhancer · the whole board →