Sixteen gates keep expensive work behind approval
Image Story Video Wizard is a set of agent instructions, templates, and one Python state helper. It moves an audio-first picture story through 16 stages: start, brief, benchmarks, writing package, script, voice, storyboard, visual style, character anchors, image prompts, image generation, asset checks, music, preview, final render, and feedback. The skill tells Codex or WorkBuddy what to do next and stops when the user must decide.
The gates are concrete. A benchmark direction must be approved before the writing package. The final script must be accepted before full narration. A 3 to 5 image pilot locks the look before bulk generation, and the preview must pass before the master render begins. Publishing is outside that approval and needs separate authorization. This sequence spends human attention where a wrong choice would multiply into dozens of assets.
What happened when we ran it
Our sandbox installed 35 Python packages in 18 seconds and used 37 MB on disk. The build completed in 5 seconds. Pytest then ran 9 tests in 7 seconds, and all 9 passed. Pip-audit reported 0 known vulnerabilities. For commit 6c6979f in a fresh Python 3.12 Debian container with 3 CPUs and 8 GB of RAM, the packaged local checks were clean.
The repository was tiny: 17 files, about 433 source lines, and a 0.1 MB checkout. It had a tests directory, no Dockerfile, and 0 CI workflow files. That size makes sense because most of the value is written guidance rather than an editor, model, renderer, or media pipeline. The missing CI file means the passing 9-test result came from our run, not a visible GitHub workflow on each change.
Those 9 tests check the state machine, prompt format, public files, gates, and initializer. They do not call the external writing, voice, image, or rendering services. The README warns against claiming those stages without running them.
Project state makes a long production resumable
The strongest piece is PROJECT_STATE.json. It records the current stage, status, next stage, pending user action, confirmed decisions, artifact paths, host capabilities, and history. Python commands initialize the folder tree and validate the state. When work moves between hosts, the skill checks the latest confirmed artifact and resumes from one pending request instead of restarting the interview.
There are 6 allowed stage statuses, written in Chinese, for work that has not started, is active, awaits confirmation, has been confirmed, needs rework, or was skipped. If an earlier decision changes, downstream artifacts stay on disk but are marked stale. That is safer than deleting expensive narration or images, and clearer than quietly treating them as trusted after the brief has changed.
The state file also records what the current host can really do: local files, model routing, a logged-in browser, TTS, image generation, HyperFrames, and rendering. An installed tool does not count as authorization. Missing capability should produce a bounded handoff, not a silent substitute. That matters when work crosses paid accounts.
The workflow is opinionated about tools and cadence
The tutorial route names WorkBuddy with Kimi K3 for long-form writing, Doubao Seed-TTS 2.0 for narration, and HyperFrames for assembly. Those choices make the instructions actionable for their intended audience. They also age faster than the state-machine ideas. The host-routing guide says to verify that named model controls are currently visible and describe the nearest verified substitute when they are not.
Production guidance gets specific. Voice selection uses the same roughly 20-second passage across 5 to 10 candidates before rate testing. A 10-minute narration maps to about 50 to 60 images at an economical 10 to 12 seconds per image, or 75 to 100 images at a faster 6 to 8 seconds. Those counts are planning rules in the skill, not measured output from our sandbox. They show why the confirmation gates matter: one cadence choice can add 40 images.
The main README and the user-facing turn labels are Chinese, while SKILL.md and the deeper workflow references are English. That split works for a Chinese creator using an English-readable agent, but it is awkward for an English-only production team. Translation would involve more than the README because stage statuses and required prompts are part of the tested contract.
It coordinates production; it does not supply production services
No model, voice engine, image generator, or renderer ships in the 0.1 MB repository. The skill checks whether each capability exists and guides a handoff if it does not. It warns against requesting plaintext secrets, asks the operator to confirm material costs before paid calls, and keeps uploads outside the default flow. These are good operating rules, though they leave the user responsible for every account and integration.
Image Story Video Wizard is therefore closer to a producer's runbook than a video application. A runbook can prevent premature batch generation and preserve decisions across a week-long project. It cannot guarantee that a character stays consistent, subtitles fit, music sits under narration, or a master file plays correctly. The workflow requires asset inspection and one complete human watch because local structural tests cannot answer those questions.
A young repository has clean tests and little history
The project was created September 2, 2026, and last pushed September 15. GitHub showed 352 stars and 1 open issue when fetched. There were no published releases. The sole open issue describes a third party using the skill to produce a brief and state file, but it reads partly as a partnership invitation and does not validate the full path through voice, images, preview, and final render.
For a new project, the documentation is unusually explicit about what has and has not been verified. The 9 passing tests make the state helper and written contract easier to trust, while 0 CI workflows and no tagged release weaken repeatability for future changes. Pin the commit when starting a real production, archive the skill with the project, and judge it by whether the first finished episode can resume cleanly after every handoff.

