mrkeyoor.com_
Fri 02 Oct 15:00 UTC
AI Toolsevaluationupdated 26 Aug 2026

cua review

Cua is a collection of tools for agents that operate graphical computers: background desktop drivers, cross-OS sandboxes, evaluation environments, and macOS virtual machines. It gives developers APIs and an MCP server for seeing screens, clicking, typing, running shell commands, and collecting trajectories across Linux, macOS, Windows, and Android.

+1,130stars / 7d
Verdict

Our Cua root run installed only 2 packages in 11 seconds, then found no build or test target, so it validated a thin pnpm surface rather than this 4,255-file cross-OS platform. Cua is worth a focused trial when an agent must control native applications or when research needs repeatable desktop environments. Pick one package, one operating system, and one permission model first; adopting the whole monorepo at once makes verification harder than the product problem requires.

We ran it

Lab card: what happened when we ran cuaScreenshot of cua (cua.ai)
Install✓ · 14s2 packages · 9 MB
Buildn/ano build script
Testsn/ano test script
Repo5266 files~959,948 lines of source · 320 MB · 132 CI workflows · tests dir

Answers from our run

Does cua build from source?

Dependencies installed in 14 seconds (2 packages), and the project has no separate build step. We cloned commit e1c44bd into a clean Debian container with 3 CPUs and no project-specific setup.

Does cua have tests you can run?

Not through a standard command: the project exposes no test script or target that our harness could run.

Who should not use cua?

Developers expecting one small package to represent the whole project: the README describes separate drivers, SDKs, benchmarks, servers, and virtualization tools with different setup paths.

What are the alternatives to cua?

Browser Use, E2B Infrastructure, OmniParser. Our Cua root run installed only 2 packages in 11 seconds, then found no build or test target, so it validated a thin pnpm surface rather than this 4,255-file cross-OS platform.

Setup2/5Root install is tiny; real paths need OS permissions, VMs, or cloud
Docs4/5Clear product paths, examples, support matrix, and package links
Community5/5Pushed August 2026 with 102 workflows and heavy active development
Maturity3/5Broad platform coverage with component-specific releases and OS gaps

Discussed on

  1. hnApple Silicon and macOS VMs: Faster LLM Inference with llama.cpp307 points
  2. hnShow HN: Drive any macOS app in the background without stealing the cursor192 points
  3. hnLaunch HN: Cua (YC X25) – Open-Source Docker Container for Computer-Use Agents172 points
  4. hnShow HN: Lumier – Run macOS VMs in a Docker159 points
  5. hnShow HN: CUA-S1 – A System One Model for Computer Use95 points

Who it’s for

Teams building computer-use agents that need more than browser automation.
Researchers evaluating agents with OSWorld, ScreenSpot, Windows Arena, or custom tasks.
Claude Code, Codex, or MCP users who want controlled access to native desktop applications.
Apple-silicon developers needing repeatable macOS or Linux virtual machines through Lume.

Who it’s NOT for

Developers expecting one small package to represent the whole project: the README describes separate drivers, SDKs, benchmarks, servers, and virtualization tools with different setup paths.
Linux teams requiring unrestricted background control on native Wayland: the README says Wayland routes have explicit limits for raw background input.
Organizations that cannot review mixed licensing: the core is MIT, but the README says optional cua-agent[omni] includes an AGPL-3.0 dependency and OmniParser assets use CC-BY-4.0.
Cloud users needing their own qcow2 or ISO images today: the support table marks cloud BYOI as coming later.
Anyone exposing a default Lume guest unchanged: the documented presets create a lume user with lume as its password.

Setup reality

Our commit b296ec9 root pnpm run installed 2 packages in 11 seconds and used 9 MB. The root exposed no build script or target and no test script or target, so both steps were skipped rather than passed. The checkout was much larger than that install suggests: 4,255 files, about 694,015 source lines, and 264.6 MB.

Each product path has separate requirements. Drivers need OS permissions; local sandboxes need QEMU or platform virtualization; Lume needs Apple Silicon; benchmarks use Python tooling and base images; cloud work needs Cua credentials and remote capacity.

The repository had 102 CI workflows and a tests directory but no Dockerfile. The root harness did not exercise Rust, Python, Swift, Nix, VM images, MCP behavior, or platform-specific GUI control.

Cua is four products sharing one computer-use mission

The repository groups several related jobs. Cua Drivers controls native applications in the background. The sandbox SDK creates Linux, macOS, Windows, Android, or container environments through one API. Cua-Bench runs evaluation and reinforcement-learning tasks. Lume manages macOS and Linux virtual machines on Apple Silicon. Treating these as one install hides more than it explains.

The common thread is computer use outside a browser tab. An agent can request screenshots, click coordinates, type text, send mobile gestures, and run shell commands. Drivers expose a CLI and MCP server for Claude Code, Cursor, Codex, OpenClaw, or custom clients. Sandboxes make the target disposable, while the benchmark tooling records trajectories for evaluation or training.

Background desktop control reaches beyond browser automation

Cua Drivers aims to click, type, and verify in macOS, Windows, and Linux applications without stealing the user's cursor or focus. The same interface can address desktop tools that Playwright cannot see. This makes it relevant for coding agents working in terminals and editors, office workflows, legacy applications, and tests that cross browser and native windows.

Operating systems impose different boundaries. macOS requires Accessibility and Screen Recording permissions. Windows capture and input follow different APIs. Linux supports X11 and compositor-specific Wayland routes, and the README explicitly warns that raw background input has limits on Wayland. A cross-OS API can normalize commands; it cannot make every desktop security model behave alike.

Open issue #3389 shows the kind of ambiguity operators must handle. On a reported macOS Electron target, type_text said delivery was failed or unverifiable even though text appeared, while TextEdit produced confirmed readback. An agent that retries blindly could duplicate input. Verification should inspect the resulting UI state and treat "unverifiable" as uncertainty, not proof that nothing happened.

What happened when we ran it

We cloned commit b296ec9 into an unprivileged Debian container with 3 CPUs and 8 GB of RAM. The root pnpm install succeeded in 11 seconds, adding 2 packages and using 9 MB on disk. The harness found no root build script or target and no root test script or target, so both steps were skipped.

Those small numbers describe only the root package surface. The checkout contained 4,255 files, about 694,015 source lines, and 264.6 MB of checked-out data. It had 102 CI workflow files and a tests directory, but no Dockerfile. A two-package install plainly did not compile or exercise the platform-specific Rust, Python, Swift, Nix, VM, benchmark, and GUI paths.

We therefore cannot claim Cua built or passed tests from this run. The result instead reveals a monorepo-discovery problem for generic tooling: the root does not provide one command that represents project health. Evaluators should choose a component and follow its own checks on the target operating system, including real permission prompts, background focus behavior, screenshots, and cleanup after failure.

Sandboxes trade host setup for disposable operating systems

The Python SDK presents one Sandbox interface for Linux containers, Linux VMs, macOS, Windows, and Android. An ephemeral image can run shell commands and receive mouse, keyboard, screenshot, and mobile calls. The README's support table covers local QEMU and the hosted Cua service; cloud use of a custom qcow2 or ISO remains marked as coming later.

Local use still needs a virtualization backend and suitable host hardware. macOS guests carry Apple's platform restrictions, while Android, Windows, and Linux images have different licensing and initialization work. Hosted use moves that burden to Cua but adds an account, credentials, remote data handling, capacity, and service dependency. Sensitive teams should decide whether screenshots and trajectories may leave their network before selecting cloud convenience.

Lume has a particularly direct Apple-silicon path using Virtualization.Framework. Its unattended Sequoia and Tahoe presets create a lume user, enable SSH and autologin, and disable sleep and screen locking. The default password is also lume. That is useful for a disposable local lab, but the credential must change before the guest reaches an untrusted network.

Benchmarks provide structure, not a universal score

Cua-Bench supports OSWorld, ScreenSpot, Windows Arena, and custom tasks, with trajectory export for training. The quick start creates a Linux Docker base image and can run up to 4 tasks in parallel in its example. Those are configuration details, not performance results from our sandbox. We did not run a benchmark or validate any agent score.

Evaluation quality depends on image reproducibility, task reset, screen resolution, timing, model configuration, and the definition of success. A framework can organize those pieces while still producing incomparable results if two teams choose different environments. Pin the Cua commit, component release, base image, agent model, and task dataset when publishing a result.

Licensing also varies inside the optional stack. The repository is MIT, Kasm components are MIT, OmniParser is listed as CC-BY-4.0, and the optional Omni extra includes Ultralytics under AGPL-3.0. Organizations distributing a product should review the exact extras and assets they ship instead of relying on the top-level badge.

August 26 activity is high and release versions are component-specific

The repository was pushed on August 26, 2026, and GitHub listed 730 combined issues and pull requests. The latest GitHub release is sandbox-v0.4.3, published August 22, with a Cua Sandbox release and CLI update. Lume, Drivers, the agent SDK, and benchmark package follow their own versions, so that tag cannot describe compatibility across the entire monorepo.

The 102 workflow files and same-day work show unusually active platform testing across operating systems. The issue queue also contains platform-specific edge cases around capture, accessibility trees, background input, and VM behavior. That combination is appropriate for hard systems software: active does not mean uniform.

Cua deserves evaluation when native desktop control or cross-OS agent research is the actual requirement. Begin with one component and reproduce its operating-system checks, because our root run skipped the meaningful builds and tests. Browser-only agents should choose a browser-focused tool; Cua's extra machinery pays off only when the computer outside the browser matters.

Alternatives

ProjectWhat it isPick it when
Browser Use gh↗An agent framework focused on controlling web browsers.pick this instead when all target work happens in websites and desktop or VM control would add needless scope.
E2B Infrastructure gh↗Open infrastructure for isolated cloud sandboxes used by agents and code tools.pick this instead when disposable code sandboxes matter more than native GUI interaction across operating systems.
OmniParserA screen-parsing model that turns interface screenshots into actionable elements.pick this instead when visual element parsing is the missing layer and you already have capture and input control.

What people are saying

  1. [velocity-scout] trycua/cua

Sources

  1. Cua README
  2. Cua Sandbox v0.4.3 release
  3. Electron background typing issue
  4. Lume Sequoia setup issue

More ai tools reviews

xialingguo-ip · reelbench-skills · SoL-Pi · flybook · crypto-rag · anything2explainer · the whole board →