mrkeyoor.com_
Fri 11 Sept 06:44 UTC
AI Toolsevaluationupdated 11 Sept 2026

reference-video-director review

Reference Video Director's primary README is Chinese, and a full English README plus English usage notes are included. It is an Agent Skill that turns reference images and a rough relationship-story idea into a timed prompt for video generators. Its specialty is short couple-POV scenes with identity, clothing, eyeline, hand, dialogue, and continuity instructions spelled out.

Verdict

The repository contains 6 Markdown files and no runnable ecosystem or Dockerfile, so Reference Video Director only supplies prompt instructions; it does not generate video. Try it when the brief is a 15 or 30-second everyday couple scene and you want more disciplined continuity language. Skip it for production reuse until the license is clarified, and do not treat its identity or hand constraints as proof that a model followed them.

We ran it

Screenshot of reference-video-director (github.com/sgyno09-source/reference-video-director)

Answers from our run

Did you run reference-video-director yourself?

No. GitHub reports no primary language for it, and it carries no manifest our lab installs from, and no Dockerfile, so there was nothing standard to install, build or test. This review is written from the repository's own documentation.

Who should not use reference-video-director?

Anyone expecting a finished video: the repository contains six Markdown files and no model API, renderer, browser automation, or media output.

What are the alternatives to reference-video-director?

Video Prompting Skill, Claude Remotion Skill, LTX-Video. The repository contains 6 Markdown files and no runnable ecosystem or Dockerfile, so Reference Video Director only supplies prompt instructions; it does not generate video.

Setup3/5Six files to import, but no client-specific verified installer
Docs4/5Full Chinese and English guides with concrete timing rules
Community2/5133 stars, one open issue, and no changes since 2026-08-28
Maturity2/5No license, releases, executable checks, or output validation

Who it’s for

Creators making 10, 15, 30, or 60-second couple-POV scenes from character, outfit, or room references.
Prompt writers who want a fixed checklist for identity, blocking, micro-expressions, dialogue timing, sound, and negative constraints.
Chinese or English users working with Seedance, Hailuo, Kling, Veo, Sora, or a model-neutral prompt.
Agent Skill users who are comfortable importing a folder and sending the resulting prompt to a separate video generator.

Who it’s NOT for

Anyone expecting a finished video: the repository contains six Markdown files and no model API, renderer, browser automation, or media output.
Creators working mainly on action, documentary, product, animation, or multi-character stories: the defaults are adult couple banter, close POV, tiny domestic conflict, and one usually off-screen boyfriend.
Teams that need verified identity or hand consistency: the Skill writes negative constraints, but it has no frame inspection or output-evaluation loop.
Organizations that need clear reuse rights before adoption: the repository has no license, and its only open issue asks the author to clarify that omission.

Setup reality

We did not run this repository because the lab found no supported ecosystem and no Dockerfile. It therefore has no install, build, test, dependency, timing, or vulnerability result from our sandbox.

Setup means downloading or cloning the whole six-file folder, then importing it into a client that supports Agent Skills. The README says to preserve SKILL.md and references/. No API key is configured here because the Skill only writes a prompt; the chosen video service is a separate step.

There is no executable runtime or client-specific installer to verify. Output quality depends on the agent following the instructions and the target video model honoring them. The repository also has no declared license or release tag.

Six Markdown files make a portable prompt playbook

The repository has 6 files, all Markdown: two READMEs, the core Skill, an editing guide, and two language references. The main README is Chinese, while README_EN.md and references/en.md give English users the same setup and usage path. There is no application hidden behind the documentation. A compatible agent reads SKILL.md, follows its rules, and returns text that the user can paste into a separate video generator.

The output template can contain 12 sections, including character and outfit consistency, setting, camera, timeline, performance, continuity, sound, negative constraints, and model notes. That structure is more useful than a loose request such as "make this cinematic" because each reference receives a job and each moment receives visible behavior. Minimal input is accepted: one reference image, a rough event, a duration, and an optional target model can become a detailed prompt.

Fifteen and 30-second couple scenes are the center of gravity

The timing rules are unusually specific. A 15-second result normally gets 1 to 3 shots and no more than 4 to 7 short dialogue lines. A 30-second scene gets 3 to 4 longer shots and usually 8 to 15 lines. If a 60-second idea meets a model limited to about 30 seconds, the Skill splits it into two parts and carries the final position, action, lighting, and emotional state across the join.

That precision comes with a narrow editorial taste. The default is a familiar adult couple, often a female lead facing a boyfriend's first-person camera. Stories favor teasing, a tiny misunderstanding, awkward comforting, suppressed laughter, and an everyday ending. The rules discourage breakup plots, major betrayal, melodramatic dialogue, rapid camera moves, and large gestures unless requested. Creators making product demonstrations, action scenes, documentaries, or ensemble work would spend much of the prompt overriding defaults.

What happened when we ran it

commit 60553c1 contains no supported language ecosystem and no Dockerfile; its 6 files are instruction documents rather than an executable project. There are consequently no sandbox figures for installation time, dependency count, disk use, build status, tests, or package vulnerabilities. That absence is the relevant finding. The README's download and import steps may be clear, but a code harness cannot verify whether a particular agent loaded the Skill or followed every direction.

Twelve sections add discipline without checking the result

The Skill anticipates common image-to-video failures. It asks the model to lock facial proportions, hair, clothing, room geometry, eyelines, mirror reflections, and hand contact. It translates emotions into observable changes such as looking away, pressing lips, or suppressing a smile. Negative constraints reject extra hands, face swaps, disappearing limbs, random scene changes, and visible camera equipment. These details can make a prompt easier for a video model to interpret.

None of the 12 output sections inspects the generated frames. The repository has no video API, task poller, image comparison, face matcher, media probe, or retry rule based on rendered evidence. A malformed hand can still appear after the prompt says "correct finger count." Identity can still drift after a strong reference mapping. Teams need a separate review loop that watches the clip, checks continuity, and decides whether to regenerate or edit the result.

Five named model families receive text-level adaptation

Seedance, Hailuo, Kling, Veo, and Sora are named in the model guidance. Seedance gets explicit time ranges and limited hand choreography; Hailuo gets stronger identity reminders and one main action at a time. Kling emphasizes spatial blocking, while Veo and Sora may receive richer sound and camera detail. These are prompt-writing branches only. The files contain no capability probe that confirms which model version, duration, reference count, or audio feature is available.

The folder import is similarly client dependent. The README tells users to select the complete Skill directory and refresh the conversation or Skill list, but it does not give verified paths for individual products. There is no package manager command, versioned release, or compatibility table. A user also needs access to a reference-capable video generator after the agent finishes writing. Any API key, credit cost, upload policy, or retention setting belongs to that separate service.

One unanswered license issue blocks safe reuse

GitHub showed 133 stars, 1 open issue, no release, and a last push on August 28, 2026. The only issue, updated September 10, asks what license governs the project because the repository has no LICENSE file. That is a serious adoption gap for a team that wants to copy, modify, bundle, or redistribute the Skill. Public visibility and readable source do not by themselves grant permission for those uses.

The closest substitute depends on the missing capability. Square-Zero-Labs' Video Prompting Skill covers more video models and character-sheet work. Claude Remotion Skill targets rendered motion graphics with an inspection loop, while LTX-Video supplies a runnable open model. Reference Video Director remains appealing for its exact niche: short, reference-led couple scenes with small reactions. Until licensing and output checks exist, keep it as an evaluated personal prompt aid rather than a dependency in a commercial production system.

Alternatives

ProjectWhat it isPick it when
Video Prompting SkillA broader Agent Skill for text-to-video, image-to-video, and character-sheet prompting across several models.pick this instead when you need many video genres or a character-sheet workflow rather than couple-POV direction.
Claude Remotion SkillA Claude-focused skill for creating and editing code-based motion graphics with Remotion.pick this instead when you need a rendered, editable composition with exact text and timing.
LTX-VideoAn Apache-2.0 video-generation model repository with local inference tooling.pick this instead when you need an open model and runnable generation stack rather than prompt guidance for hosted models.

What people are saying

  1. [velocity-scout] sgyno09-source/reference-video-director

Sources

  1. Reference Video Director repository
  2. Reference Video Director Chinese README
  3. Reference Video Director English README
  4. Reference Video Director Skill instructions
  5. Reference Video Director English usage notes
  6. Reference Video Director license question

More ai tools reviews

keras · whisper.cpp · agency-agents-zh · speech-to-speech · alphagenome · awesome-generative-ai-apps · the whole board →