mrkeyoor.com_
Mon 28 Sept 15:18 UTC
AI Toolsevaluationupdated 28 Sept 2026

SkillOpt review

SkillOpt improves the Markdown instruction files that guide an AI agent while leaving the model itself unchanged. It runs tasks, studies the results, proposes bounded edits, and keeps a candidate only when it passes validation; a separate preview called SkillOpt-Sleep can learn from supported coding-agent session histories.

Verdict

Our SkillOpt run passed 1,499 tests after a 116-package, 652 MB install, so the research engine earns a serious trial. Use it when you can define held-out tasks and score the behavior you want; that validation discipline is the product. Keep SkillOpt-Sleep off sensitive or unattended projects until you have reviewed its transcript boundary, open harvesting issues, and the gap between PyPI 0.2.0 and main.

We ran it

Lab card: what happened when we ran SkillOptScreenshot of SkillOpt (aka.ms/skillopt)
Install✓ · 85s116 packages · 652 MB
Build✓ · 1s
Tests✓ · 63s1499 passed · 0 failed · 9 skipped of 1499 (pytest)
Known vulns0(pip-audit)
Repo434 files~73,288 lines of source · 5.2 MB · 1 CI workflows · tests dir

Answers from our run

Does SkillOpt build from source?

Dependencies installed in 85 seconds (116 packages), and the build succeeded in 1 seconds. We cloned commit 79124b3 into a clean Debian container with 3 CPUs and no project-specific setup.

Do SkillOpt's tests pass?

Yes: 1499 of 1499 passed when we ran the project's own test command (pytest). Some failures need services or credentials a bare container does not have.

Does SkillOpt have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use SkillOpt?

Teams hoping one good chat transcript will improve an agent automatically: the research path expects benchmark splits, scored rollouts, and held-out validation.

What are the alternatives to SkillOpt?

DSPy, ART, Agent Lightning. Our SkillOpt run passed 1,499 tests after a 116-package, 652 MB install, so the research engine earns a serious trial.

Setup3/5Install passed, but real runs need data, a backend, and evaluation
Docs5/5Versioned guides explain backends, splits, Sleep, and privacy
Community4/517,760 stars with active September 2026 issues and PRs
Maturity3/51,499 tests pass, but Sleep is preview software with release lag

Who it’s for

Agent teams with repeatable tasks, a scoring method, and enough examples to keep training and validation separate.
Researchers comparing prompt or skill optimization without changing model weights.
Claude Code, Codex, Copilot, and Devin users willing to inspect proposed skill changes before adopting them.
Platform engineers who can budget provider calls and maintain benchmark-specific data and configuration.

Who it’s NOT for

Teams hoping one good chat transcript will improve an agent automatically: the research path expects benchmark splits, scored rollouts, and held-out validation.
Organizations that cannot send transcript-derived text to a model provider: the Sleep guide says outbound prompts are not guaranteed to be secret-free when a real backend is used.
Anyone who needs the PyPI package to match the current repository: version 0.2.0 omits later main-branch features and does not ship the agent plugin and MCP folders.
Operators ready to schedule unattended Sleep runs without auditing scope: open issue 294 reports home-directory sessions entering unrelated projects, and issue 286 reports Codex replay sessions being harvested as user evidence.
Small environments where a 652 MB Python install is already too large before benchmark data or local model services are added.

Setup reality

Our sandbox installed 116 Python packages in 85 seconds and used 652 MB on disk. The build succeeded in 1 second. Tests succeeded in 63 seconds: pytest reported 1,499 passed, 0 failed, and 9 skipped. Pip-audit found 0 known vulnerabilities.

Python 3.10 or newer is required. Research runs also need a configured hosted API, local model server, or authenticated execution CLI, plus benchmark data and YAML configuration. The Sleep mock backend needs no credentials, but it cannot perform a real provider-backed learning cycle.

The 5.2 MB checkout had 434 files and about 73,288 lines of source at commit 79124b3. It has a tests directory and 1 CI workflow file, but no Dockerfile. PyPI 0.2.0 installs the commands, while benchmark configs, data materializers, development tests, and agent integrations require a source checkout.

The trainable object is a 300 to 2,000-token Markdown file

SkillOpt's deployed result is usually a best_skill.md file of 300 to 2,000 tokens, not a new set of model weights. The target model attempts scored tasks, a separate optimizer reviews those attempts, and proposed add, delete, or replace edits face a held-out validation check. Accepted text can then travel with an otherwise unchanged agent. That is a narrower and more practical target than retraining a model when the recurring problem is an instruction set.

The approach is best understood as experimental prompt engineering with stricter bookkeeping. SkillOpt has epochs, batches, edit budgets, and validation gates, but those familiar names do not create a useful objective for you. A team still has to decide what success means, collect representative tasks, and protect examples the optimizer did not see. If you cannot score the work reliably, the loop can make a skill look different without proving that it became better.

Our 1,499-test pass makes the research engine worth trying

Our sandbox installed 116 packages in 85 seconds, taking 652 MB on disk. The build completed in 1 second. Pytest then finished successfully in 63 seconds with 1,499 passed, 0 failed, and 9 skipped. Pip-audit reported 0 known vulnerabilities in the installed environment. For an experimental AI repository with several backends and benchmark adapters, that is a much better starting signal than a polished diagram alone.

The checkout at commit 79124b3 contained 434 files and about 73,288 lines of source in 5.2 MB. Our scan found a tests directory and 1 CI workflow file, but no Dockerfile. The missing container recipe puts environment reproduction on the adopter, while the passing suite reduces the risk of changing Python dependencies or extending an adapter. Neither result says that your chosen model will learn a useful skill from your data.

A real run needs held-out tasks and a paid or local backend

SkillOpt supports 6 documented benchmark families and separates optimizer and target roles in YAML. Both roles may use the same backend, or a team can assign different models. Hosted APIs need credentials, local Qwen requires an OpenAI-compatible server, and CLI routes require the corresponding authenticated agent. The .env template is not loaded automatically, so the guide tells users to export it into the shell before training.

Data preparation is the larger commitment. A benchmark adapter needs a loader, scored rollout behavior, configuration, and a split that keeps validation useful. The default paper-style path accepts an edit only when the held-out score strictly improves. Turning that gate off force-accepts candidates and changes the meaning of the experiment. SkillOpt can enforce the decision rule, but it cannot rescue a weak judge, contaminated split, or task set that omits the failures users care about.

PyPI 0.2.0 trails the September 2026 main branch

The latest release is v0.2.0 from July 2, 2026, while the repository was pushed on September 5. The installation guide lists several later features that require a source checkout, including extra Sleep backends, reviewed subset adoption, and multi-skill fan-out. The wheel also omits benchmark configurations, data materializers, development tests, and the repository's agent integration shells and MCP servers.

That gap changes setup advice. pip install skillopt is enough to inspect the three installed commands and try the Sleep mock path. Reproducing paper experiments or installing Claude Code, Codex, Copilot, or Devin integrations means cloning the repository. Open issue 288 asks for a release cut because security hardening merged after 0.2.0 is not in the wheel. Teams should choose the package source deliberately instead of treating PyPI and main as interchangeable.

Sleep can send transcript excerpts outside your machine

SkillOpt-Sleep is labeled a preview and stages proposed changes for human review. Harvesting itself is local and read-only, and the mock backend makes no provider calls. A real backend sends truncated transcript excerpts and derived tasks to the selected provider for mining, replay, judging, and reflection. The guide says those outbound prompts are not guaranteed to be free of secrets, even with best-effort redaction.

Two open reports sharpen that warning. Issue 294 says --scope invoked can pull sessions started from a home directory into every project beneath it. Issue 286 says the Codex source at commit 79124b3 can ingest sessions that SkillOpt created during replay, turning evaluator instructions into later training evidence. These are specific provenance failures. Until they are closed in the version you deploy, inspect harvested task files and avoid scheduling Sleep over sensitive histories.

September issue traffic shows demand and unfinished edges

SkillOpt had 17,760 stars and 54 combined issues and pull requests when fetched. The last push was September 5, 2026, while harvesting reports and proposed fixes were still arriving later in September. That combination points to active outside scrutiny, but the default branch date also shows that recently opened fixes may not yet be available to users. Count merged code and release contents, not issue discussion alone.

The research engine and Sleep should receive different adoption decisions. Our 63-second successful test run supports trying the explicit benchmark workflow first, where tasks, gates, and artifacts are visible. Sleep touches personal coding history, spends provider calls, and proposes persistent instruction changes. Its review-before-adopt design is sensible, but the current scope and provenance reports make supervision part of the operating model, not an optional safety pass.

Alternatives

ProjectWhat it isPick it when
DSPy gh↗A framework for defining and optimizing multi-stage language-model programs.pick this instead when the unit you want to optimize is an LLM program or prompt pipeline rather than one deployable skill document.
ARTAn agent reinforcement-training framework built around rollouts and reward signals.pick this instead when you want reinforcement learning to change model behavior and can support the training infrastructure.
Agent Lightning gh↗A training system that separates agent execution from reinforcement-learning optimization.pick this instead when you need a broader agent-training layer rather than Markdown skills that run on a frozen model.

What people are saying

  1. [github-trending] microsoft/SkillOpt

Sources

  1. SkillOpt repository and README
  2. SkillOpt installation guide
  3. SkillOpt-Sleep data boundary and usage guide
  4. SkillOpt v0.2.0 release notes
  5. Project-scope harvesting issue 294
  6. Codex self-harvesting issue 286
  7. Release and staging-gap issue 288

More ai tools reviews

rizzo-pii · redamon · awesome-ai-agent-platforms · guizang-yingzao-skill · vdn-minimax-h3 · image-to-3d-pipeline · the whole board →