Voice stays responsive while a separate agent does the work
Qwen Audio Agent separates the conversation from the job behind it. A realtime voice service listens and responds, while a backend agent can edit files, call tools, or run a longer task. You can ask for progress or cancel without waiting in silence for that task to finish. The distinction is useful for a desktop assistant because speaking and doing no longer share one blocking request.
Version 2.0 supports ACP and A2A backends, with adapters for coding agents named in the README. The front end can use Qwen Audio 3.0 Realtime, other cloud voice services, or a local Hugging Face speech-to-speech server. The same Gateway feeds a terminal interface, WebUI, and desktop orb. That breadth is the reason to consider it, and the source of most setup decisions.
The 397 MB install buys a full coordination layer
Our sandbox installed 748 npm packages in 31 seconds and used 397 MB on disk. The monorepo contained 1,424 files, roughly 177,972 lines of source, and four CI workflow files. It has a tests directory but no Dockerfile. This is a substantial application with desktop, web, terminal, Gateway, memory, and integration code, not a thin microphone wrapper.
The default quick start asks for a DashScope API key and starts the Gateway in one terminal, then the TUI or WebUI in another. A backend agent is optional for voice-only use. Once enabled, that backend brings its own authentication, model settings, tools, MCP servers, and project access. The one-click language in the README describes installation support, while working credentials and permissions remain your responsibility.
What happened when we ran it
Our run built commit f6dd0e3 in 11 seconds after the 31-second install. Node's test runner completed in 119 seconds with 1,581 passed and 0 failed. The unprivileged Debian container had 3 CPUs, 8 GB of RAM, Node 22, and no secrets. That is strong evidence that the supplied JavaScript packages agree with each other at this commit.
Npm audit found 7 known vulnerabilities: 3 high and 4 moderate, with no critical or low findings. The measurement does not say whether any advisory is reachable through the Gateway, desktop app, or build tooling, so we will not turn the count into an incident claim. It does make dependency triage a required evaluation step before the software receives microphone data, API keys, or access to an agent workspace.
Default audio leaves the machine, and LAN mode is unencrypted
The privacy guide is unusually direct about data flow. By default, microphone audio, realtime transcript context, and model response requests go to DashScope. Camera frames move only after the user enables realtime vision, and the Gateway says it does not write those frames into history or local files. Backend instructions and required attachments go to whichever task agent you configure. Each provider adds its own policy boundary.
Local state defaults to ~/.config/qwaudio/, including profiles, long-term memory, tasks, backend work directories, configuration, and logs. Uninstalling does not remove that directory. The Gateway listens locally by default, but --lan uses unencrypted HTTP and WebSocket traffic. The project advises Tailscale HTTPS Serve or a trusted reverse proxy for cross-network access, along with client pairing or access tokens.
Node and provider requirements narrow the easy path
Source and CLI installations require Node 22.22.2, Node 24.15.0, or version 26 and later, plus npm 10. Desktop installers bundle the Gateway for macOS and Windows, while Linux desktop users build from source. Installing a desktop package does not install or authenticate the backend agent, so packaged distribution removes the Node step without removing the integration work.
Provider choice affects features. Some realtime front ends support video, one local option has no tool calling, and cloud paths use different keys or endpoints. Issue 530 reports that the international Qwen plan lacks voice enrollment and that the client cannot select a cloned custom voice. Issue 131 still tracks browser acceptance work around camera permissions, disconnects, session changes, and multimodal privacy boundaries.
A two-day v2 release makes change control part of adoption
Release v2.0.0 arrived on September 23, 2026 with a rebuilt orchestration runtime, unified client protocol, new backend integrations, memory work, and expanded desktop features. Version 2.0.1 followed three days later with backend discovery, installation, and coordinator session fixes. GitHub recorded another push on September 28. This is fast maintenance around a fresh architecture, not a settled long-term branch.
GitHub showed 2,816 stars and 11 open issues and pull requests on September 29. Current work covers microphone state, custom voices, knowledge imports, Windows process handling, and speaker changes. The project earns a serious trial when its exact proposition matches yours: one continuing conversation in front, several agent tasks behind it. Our 1,581 passing tests support that trial, while the 7 audit findings and explicit privacy boundaries set the release gate.

