The output is a reusable CLI, not a one-off agent action
CLI-Anything turns source access into an application-specific Python command line. Its method asks an agent to map GUI actions to callable APIs, design command groups and state, implement Click commands with JSON output, write unit and end-to-end checks, document the result, and package it. The generated harness can then be called by a person, a shell script, or another agent without repeating the source-analysis work for every task.
The companion CLI-Hub handles discovery, installation, updates, launching, and removal. It mixes project-built harnesses with selected public CLIs, and v0.4.0 added matrices that install several tools for a workflow. Many wrappers still depend on the real application or backend. Installing a GIMP, Blender, LibreOffice, or Obsidian entry does not install away the upstream program, its plugins, API keys, file formats, or platform quirks.
Claude Code gets a plugin and each harness gets a skill
Claude Code users add the GitHub marketplace and install the CLI-Anything plugin, then invoke /cli-anything against a path or repository. Codex, Cursor, Pi, OpenCode, OpenClaw, Hermes, and other agents have separate adapters. Each generated harness also receives a SKILL.md, which helps compatible agents discover its commands and constraints rather than guessing from --help alone.
That packaging has had real drift. Open issue 433 reports that an npx skills add installation omitted the references directory containing HARNESS.md and command specifications. The installed skill instructed agents to read files that were not present, leaving only condensed guidance. The repository's Codex installer vendors the reference material differently. Before generating anything, verify that the chosen agent installation includes the full method and mode-specific command instructions.
What happened when we ran it
Our sandbox tested commit 810c18b from the cli-hub subproject. Installation took 28 seconds, adding 35 Python packages and 36 MB on disk. The build succeeded in 11 seconds. The full checkout contained 1,872 files, about 321,103 source lines, and occupied 63 MB. We found 7 CI workflow files, no Dockerfile, and a tests directory. Pip-audit reported 0 known vulnerabilities.
Pytest exited with code 1 after 14 seconds: 157 tests passed and 7 failed of 164. Every listed failure was in TestAnalytics, and every traceback ended with RuntimeError: Could not determine home directory. The failures covered event sending, install and uninstall event names, launch, and human or agent visits. The log does not establish why a home directory was unavailable. It does establish that the measured commit's hub suite did not pass in our fresh Python 3.12 Debian container.
A successful command can still produce no artifact
The project's own method says export commands should verify output magic bytes, archive structure, pixels, audio levels, or duration rather than trusting exit status. Open issue 451 shows why. In the reported Blender wrapper, render execute generated a script and printed the command it would run, returned exit code 0, but never started Blender and never produced the requested PNG. Pull request 453 was opened to invoke Blender during execution.
Issue 454 is more serious because it concerns data loss. The reported Obsidian note open implementation sent a PUT request to the active-note endpoint with the target path as the request body. That endpoint replaces the current note's contents, so the active note became the literal path string instead of opening the requested file. A wrapper around a familiar application can make a wrong API call just as easily as any handwritten integration.
The generator depends on source access and model judgment
Python 3.10 or newer is the base requirement. A new harness also needs the target program or repository and a supported coding agent. The README explicitly says strong foundation models are needed for dependable generation, weaker models may emit incomplete or incorrect CLIs, closed-source binaries reduce coverage, and one pass may not cover the application. Those limitations are unusually candid and should shape the workflow.
Generated code should enter the same review process as a human contribution. Check authentication, path handling, command quoting, output verification, rollback, and what delete, overwrite, publish, or send really do. Run it first on disposable data. The 7-phase method provides useful structure, yet phase names cannot certify that a generated test asserted the right behavior against a current upstream application.
August bug activity shows a living but uneven catalog
GitHub recorded 48,322 stars, 83 combined issues and pull requests, and a last push on August 21, 2026. Current work includes fixes for Blender execution, preview bundle selection, skill output paths, and harness CI. Release v0.4.0 arrived on June 25 with new catalog entries, workflow matrices, security fixes, and contributions from many new authors.
Catalog size creates a maintenance burden because every upstream application can change independently. Open issue 403 argued that advertised harness results were not reproduced by CI; a later watchdog CI pull request addressed suite drift, but the issue remained open. CLI-Anything is most credible as a method and starting codebase. The installed wrapper still has to earn trust against the precise app version, operating system, files, and side effects that matter to you.

