The trainable object is a 300 to 2,000-token Markdown file
SkillOpt's deployed result is usually a best_skill.md file of 300 to 2,000 tokens, not a new set of model weights. The target model attempts scored tasks, a separate optimizer reviews those attempts, and proposed add, delete, or replace edits face a held-out validation check. Accepted text can then travel with an otherwise unchanged agent. That is a narrower and more practical target than retraining a model when the recurring problem is an instruction set.
The approach is best understood as experimental prompt engineering with stricter bookkeeping. SkillOpt has epochs, batches, edit budgets, and validation gates, but those familiar names do not create a useful objective for you. A team still has to decide what success means, collect representative tasks, and protect examples the optimizer did not see. If you cannot score the work reliably, the loop can make a skill look different without proving that it became better.
Our 1,499-test pass makes the research engine worth trying
Our sandbox installed 116 packages in 85 seconds, taking 652 MB on disk. The build completed in 1 second. Pytest then finished successfully in 63 seconds with 1,499 passed, 0 failed, and 9 skipped. Pip-audit reported 0 known vulnerabilities in the installed environment. For an experimental AI repository with several backends and benchmark adapters, that is a much better starting signal than a polished diagram alone.
The checkout at commit 79124b3 contained 434 files and about 73,288 lines of source in 5.2 MB. Our scan found a tests directory and 1 CI workflow file, but no Dockerfile. The missing container recipe puts environment reproduction on the adopter, while the passing suite reduces the risk of changing Python dependencies or extending an adapter. Neither result says that your chosen model will learn a useful skill from your data.
A real run needs held-out tasks and a paid or local backend
SkillOpt supports 6 documented benchmark families and separates optimizer and target roles in YAML. Both roles may use the same backend, or a team can assign different models. Hosted APIs need credentials, local Qwen requires an OpenAI-compatible server, and CLI routes require the corresponding authenticated agent. The .env template is not loaded automatically, so the guide tells users to export it into the shell before training.
Data preparation is the larger commitment. A benchmark adapter needs a loader, scored rollout behavior, configuration, and a split that keeps validation useful. The default paper-style path accepts an edit only when the held-out score strictly improves. Turning that gate off force-accepts candidates and changes the meaning of the experiment. SkillOpt can enforce the decision rule, but it cannot rescue a weak judge, contaminated split, or task set that omits the failures users care about.
PyPI 0.2.0 trails the September 2026 main branch
The latest release is v0.2.0 from July 2, 2026, while the repository was pushed on September 5. The installation guide lists several later features that require a source checkout, including extra Sleep backends, reviewed subset adoption, and multi-skill fan-out. The wheel also omits benchmark configurations, data materializers, development tests, and the repository's agent integration shells and MCP servers.
That gap changes setup advice. pip install skillopt is enough to inspect the three installed commands and try the Sleep mock path. Reproducing paper experiments or installing Claude Code, Codex, Copilot, or Devin integrations means cloning the repository. Open issue 288 asks for a release cut because security hardening merged after 0.2.0 is not in the wheel. Teams should choose the package source deliberately instead of treating PyPI and main as interchangeable.
Sleep can send transcript excerpts outside your machine
SkillOpt-Sleep is labeled a preview and stages proposed changes for human review. Harvesting itself is local and read-only, and the mock backend makes no provider calls. A real backend sends truncated transcript excerpts and derived tasks to the selected provider for mining, replay, judging, and reflection. The guide says those outbound prompts are not guaranteed to be free of secrets, even with best-effort redaction.
Two open reports sharpen that warning. Issue 294 says --scope invoked can pull sessions started from a home directory into every project beneath it. Issue 286 says the Codex source at commit 79124b3 can ingest sessions that SkillOpt created during replay, turning evaluator instructions into later training evidence. These are specific provenance failures. Until they are closed in the version you deploy, inspect harvested task files and avoid scheduling Sleep over sensitive histories.
September issue traffic shows demand and unfinished edges
SkillOpt had 17,760 stars and 54 combined issues and pull requests when fetched. The last push was September 5, 2026, while harvesting reports and proposed fixes were still arriving later in September. That combination points to active outside scrutiny, but the default branch date also shows that recently opened fixes may not yet be available to users. Count merged code and release contents, not issue discussion alone.
The research engine and Sleep should receive different adoption decisions. Our 63-second successful test run supports trying the explicit benchmark workflow first, where tasks, gates, and artifacts are visible. Sleep touches personal coding history, spends provider calls, and proposes persistent instruction changes. Its review-before-adopt design is sensible, but the current scope and provenance reports make supervision part of the operating model, not an optional safety pass.

