mrkeyoor.com_
Wed 30 Sept 06:10 UTC
Automationevaluationupdated 30 Sept 2026

huashu-mac-use review

huashu-mac-use is primarily a Chinese-language Agent Skill for controlling macOS applications that lack a normal API, with an English summary at the end of its README. It gives Claude Code, Codex, and other skill-aware agents a Swift command-line kernel for reading windows, choosing an automation channel, writing with safety checks, and saving screenshots as evidence.

Verdict

Our lab did not run commit 4dda98c because it requires a supported Swift-on-macOS environment and provides no Dockerfile, so its impressive safety design remains unverified by us. Try it on a non-sensitive Mac if you value background reads, explicit refusal states, and screenshot evidence more than pure visual clicking. Keep a person in charge of irreversible actions, and inspect the open screenshot-failure fix before trusting an exit code as proof.

We ran it

Screenshot of huashu-mac-use (github.com/alchaincyf/huashu-mac-use)

Answers from our run

Did you run huashu-mac-use yourself?

No. Its code is Swift, and it carries no manifest our lab installs from, and no Dockerfile, so there was nothing standard to install, build or test. This review is written from the repository's own documentation.

Who should not use huashu-mac-use?

Windows or Linux users: the merged code targets macOS, while Windows support remains in open pull request 5.

What are the alternatives to huashu-mac-use?

Claude Quickstarts, Interceptor, application-use. Try it on a non-sensitive Mac if you value background reads, explicit refusal states, and screenshot evidence more than pure visual clicking.

Setup2/5One installer, but macOS permissions and an unverified Swift build remain
Docs5/5Detailed control layers, safety gates, commands, and stop rules
Community3/5291 stars and recent detailed issues, but only 5 merged commits
Maturity2/5No release, no lab run, and a false-success fix is still open

Who it’s for

macOS power users who want a coding agent to operate native desktop applications.
Testers who need window-level screenshots and a record of what changed after each action.
Agent developers comparing AppleScript, CDP, Accessibility, coordinate, and pixel routes.
Claude Code and Codex users willing to supervise permissions and keep irreversible actions human-controlled.

Who it’s NOT for

Windows or Linux users: the merged code targets macOS, while Windows support remains in open pull request 5.
Anyone unwilling to grant Screen Recording and Accessibility access to the terminal application running the agent.
Fully unattended publishing, payments, deletion, overwrites, consent dialogs, keychain, Touch ID, or high-risk apps: the skill's own stop rules hand those actions back to a person.
Browser automation: the instructions explicitly route browser windows to a separate huashu-chrome skill.
Teams requiring a verified build in a fresh lab: we could not run this Swift project in our Debian sandbox, and open pull request 3 documents a screenshot path that can report success when no file was written.

Setup reality

We did not run commit 4dda98c in our sandbox. Its supported ecosystem is Swift on macOS, our Debian lab has no supported runner for it, and the repository has no Dockerfile, so we have no measured install, build, or test result.

Installation uses the Skills CLI or a manual clone, then scripts/build.sh compiles the Swift kernel. It needs Xcode Command Line Tools, Node.js for embedded-Chromium CDP work, and Screen Recording plus Accessibility permissions on the terminal app. No model API credential is required by the tool itself.

The README reports tests on macOS 26 and says macOS 14 or newer is likely, which is not the same as verified compatibility. Changing terminal apps requires new permissions, cross-Space coordinate writes are refused, and some applications filter synthetic input or cannot render background captures.

Four control layers avoid unnecessary visual clicking

huashu-mac-use starts by probing the target application. A native CLI, AppleScript dictionary, URL scheme, local port, or embedded Chromium debugging interface is preferred because it can act on structure. If those routes fail, the agent tries the macOS Accessibility tree, then window-relative coordinates, with pixels kept as the final control layer and a recurring verification source.

That ordering is the project's best idea. CDP can operate an embedded Chromium app across desktops without taking focus, while an Accessibility write can target a named control instead of a guessed pixel. The Swift kernel exposes window listing, background screenshots, element inspection, typing, clicking, scrolling, focus state, and a heads-up indicator. Its probe.sh script records bundle details, architecture, URL schemes, local ports, editable controls, and the app version before action begins.

Writes pass safety gates before borrowing focus

Reads are designed to stay in the background. For writes, the tool first posts an event to the target process and compares screenshots to see whether anything changed. Escalation to focus checks that the right process is frontmost, the point is not covered, the user has been idle for 2 seconds, and no other agent holds the machine-wide focus lock. It can wait up to 15 seconds before refusing.

Cross-desktop coordinate writes return exit code 2 instead of switching the user's Space. Terminal and IDE actions need special care because pressing Return can execute a shell command. A visible corner pulse tells the user when the agent borrows focus, while the overlay is excluded from captured evidence. The skill also asks for a dry run before writes so coordinate conversion and gate decisions can be examined without sending input.

The stop list is broad enough to matter. Publishing, submitting, payments, deletion, overwrite saves, consent, system permission dialogs, keychains, Touch ID, updates, and high-risk financial or medical apps stay with the user. Screen text is treated as data rather than a new instruction. These rules do not make desktop automation safe by themselves, but they define where the tool should return control instead of finding another click path.

What happened when we ran it

We did not run commit 4dda98c in our September 29 sandbox. The project language is Swift, the supported platform is macOS, and our Debian container has no supported ecosystem for compiling or exercising its window and Accessibility APIs. The repository provides no Dockerfile. We therefore have no measured installation time, build result, dependency count, vulnerability audit, or test result.

That limitation is load-bearing. We read the Swift build script and operating instructions, but our run did not grant macOS Screen Recording or Accessibility access, enumerate a window, inject input, or confirm a screenshot. The repository's Blender and desktop-client case studies belong to the author. They are useful demonstrations, not results reproduced on our box.

The build path itself is simple on paper: build.sh checks for swiftc and compiles mac.swift with optimization. The README says the kernel was tested on macOS 26 and describes version 14 or later as probable. Treat the latter as an expectation, not a compatibility promise. Apple permission behavior, window servers, application versions, and synthetic-event filtering can all change what works.

Success codes still need evidence from the application

The skill repeatedly warns that a tool call can succeed while the application ignores it. Its action commands classify the observed effect as confirmed, partial, suspected no-op, or unverifiable. Stronger proof comes from a changed app state or side effect, such as a task appearing or a file existing, rather than text merely appearing in an input field.

Open pull request 3 gives that principle an uncomfortable test. The report says a screenshot command could return exit code 0 when its destination directory did not exist, then mislabel the missing output as a blank background render. The proposed fix creates parent directories, checks that the file exists, and stops swallowing write errors. Until that fix is merged into the commit you install, verify evidence files independently.

Issue 2 documents another boundary on WeChat 4.1.13: several synthetic input routes returned success with no interface change, while a lower-level HID event worked. The issue also reports ignored modifier arguments and stale per-app notes. These are outside reports rather than our measurements. They show why the skill's per-application archive needs dates and version checks, and why a general macOS automation layer cannot promise identical behavior across apps.

Recent issue work is ahead of the merged branch

GitHub showed 291 stars and 4 combined open issues and pull requests on September 30, 2026. The merged branch was last pushed on September 8 and has 5 commits, with no published release. Detailed reports and a Windows port arrived later in September, but neither open pull request changes the fact that the reviewed commit is macOS-only.

The repository is young, yet its operating method is more thoughtful than many screen-clicking demos. Prefer structured routes, refuse uncertain coordinates, watch the user's activity, and demand a side effect before claiming success. Those ideas are worth borrowing. Whether this particular kernel belongs on your Mac depends on a real trial with non-sensitive apps, the exact terminal permissions, and a build where the open screenshot fix has been resolved.

Alternatives

ProjectWhat it isPick it when
Claude Quickstarts gh↗Anthropic's examples include computer-use demos and safety guidance around the Claude API.pick this instead when you want the model vendor's reference harness rather than a local Agent Skill for arbitrary runtimes.
InterceptorA CLI and MCP server spanning browsers, native macOS applications, and real iPhones.pick this instead when MCP integration and one tool across browser, desktop, and iPhone matter more than this skill's evidence workflow.
application-useA macOS-native command-line tool built specifically for agent-driven application automation.pick this instead when you want a smaller automation CLI rather than a large instruction and forensics method.

What people are saying

  1. [velocity-scout] alchaincyf/huashu-mac-use

Sources

  1. huashu-mac-use README
  2. huashu-mac-use skill instructions
  3. Screenshot false-success pull request
  4. WeChat automation findings
  5. Windows support pull request

More automation reviews

cloudflare-turnstile-solver · stop-stutter · warp-masque-actions · ansible · ffmpeg-skill · fable-orchestrator · the whole board →