Downloads become a searchable media library
VidBee combines jobs that usually live in separate tools. It downloads video, audio, playlists, and channels through a queue, or accepts a local media file. Finished items can be transcribed, searched by spoken words or speaker names, and played from a clicked timestamp. RSS subscriptions can watch a feed and apply their own download directory, filename template, tags, keyword filter, and automatic-download choice.
Our checkout at commit 5155e90 held 687 files and roughly 89,057 source lines across desktop, browser extension, API, web, and shared packages. Support for more than 1,000 sites comes from extractors that must keep up with changing pages. VidBee gives those commands a queue and library, but a blocked extractor remains dependent on the source site.
Transcription stays local until an AI prompt runs
Speech recognition can use local Whisper, SenseVoice, Parakeet, or Qwen3-ASR model families. VidBee recommends models based on the computer and language, lets the user switch downloaded models, and can start transcription after a download. Speaker detection can infer a count or accept a fixed number. Search, timestamp navigation, relabeling, copying, and text or Markdown export remain useful without an AI provider.
Our dependency install used 1,579 MB before any media file or speech model was added. That figure matters on a laptop more than the 54-second install time. Model size and transcription speed were not present in the supplied lab result, so we will not guess at either. Plan separate storage for the pnpm workspace, model cache, download library, captions, and generated exports.
Hosted prompts send transcript content to the selected provider
VidBee can send transcript prompts to OpenAI, Anthropic, Google, DeepSeek, Groq, Azure, Hugging Face, OpenRouter, or xAI. It also accepts a custom base URL and model ID. The README states the boundary plainly: transcription happens on the computer, while a prompt and transcript content go to the selected provider. Ollama and LM Studio can keep that second step local.
The build completed successfully in 17 seconds, but that result did not exercise any provider connection or API key. Each hosted service has its own credentials, retention terms, limits, and bill. For interviews or unpublished reporting, an administrator should lock down approved providers rather than leaving each user to turn a local transcript into a cloud request.
What happened when we ran it
Our sandbox installed 1,400 pnpm packages in 54 seconds, occupying 1,579 MB, and built the monorepo in 17 seconds. We used an unprivileged Node 22 Debian container with 3 CPUs, 8 GB of RAM, and no secrets. The checkout had 8 CI workflow files, a Compose file, monorepo workspaces, no Dockerfile, and no tests directory. Nothing in the supplied install or build result failed.
There was no test script or target, so our harness skipped tests. That is different from a passing suite. We did not validate a desktop package, run yt-dlp against a live site, download an ASR model, transcribe audio, identify speakers, or call an AI provider. The 17-second build shows the source compiled; it does not establish that the media workflow works from download through export.
Site extractors can fail even when the app builds
Issue 461 records a Rumble request ending in HTTP 403 from the underlying extractor. Issue 449 shows a TikTok page returning an unexpected response even with the reported yt-dlp version current. Other August reports cover YouTube 403 responses, Instagram extraction failure, and Windows browser-cookie decryption through DPAPI. These logs show different failures, so there is no single cause to assign.
A 1,400-package install does not prove that an extractor works. Keep the generated yt-dlp command from a failure and try it independently when diagnosing. Some sources need current cookies or a JavaScript runtime; others may reject the request anyway. If a site is essential, pin a real sample URL in a scheduled check and retain yt-dlp as a direct fallback.
Desktop and server deployment solve different problems
The desktop app is the complete personal workflow shown in the screenshots. The monorepo also contains a Fastify API, a TanStack Start web client, and a shared downloader core. pnpm run start:web starts the API and web app together. Compose examples expose ports 3100 and 3000, persist a SQLite history database, and mount storage that can point at a NAS.
Our scan found a Compose file but no Dockerfile at the repository root, although the README names published GHCR images. Treat the desktop and server routes as different products. A NAS can solve shared storage, while local ASR and playback may still belong on a user's machine. Confirm where models execute before assuming the web stack reproduces every desktop feature.
v2.0.2 is active, while regression proof is thin
GitHub records the last push and release v2.0.2 on August 23, 2026. That release fixed a Windows transcript-player crash. The repository had 32 open issues excluding pull requests and 10,386 stars when fetched. Issue 463 then reported that v2.0.2 on macOS 12.7.1 could not prepare a download engine despite locally installed FFmpeg and yt-dlp.
The successful 17-second build makes VidBee worth testing for work that begins with saved media and ends with notes tied to the recording. The missing test target keeps it from earning unattended trust. Try representative files and sites on the exact operating system, budget for 1,579 MB plus models and media, and decide whether transcript prompts may leave the computer before adding provider keys.

