Three visual inputs turn a photo into a directed edit
Yingzao does more than append a style prompt to a travel photo. For a high-style job, it requires 3 visual inputs in a fixed order: a corrected source photo, one dominant reference image, and a sparse layout guide. The agent first decides how the subject, background, lettering, and interaction should behave. A preparation script then checks that those decisions reached the actual image-model inputs.
That sequence is the project's best idea. It separates evidence from influence. The source image preserves roof slopes, eaves, signs, asymmetry, and other identity cues. The reference supplies one coherent visual mechanism. The layout guide marks text regions, shared axes, and overlaps without pretending that ordinary font outlines are the final lettering. A model still makes the edit, but the call arrives with a bounded job rather than a bag of style words.
Eight gates sit before the expensive image call
The preflight document defines 8 hard gates. They cover dependencies and paths, photo geometry, factual text, reference compatibility, visible model work, the layout guide, display-letter design, and final call arguments. The preparation script must return READY before generation. Output belongs under the caller's output/yingzao/<run-id>/ directory, keeping generated files out of the installed skill.
Several checks address mistakes that look fine at first glance. Place names and historical text must be verified, visible in the photo, or confirmed by the user. Reinterpreted Chinese display type gets a glyph brief, while smaller factual copy stays in an ordinary family. If the poster says the subject overlaps the title, the layout guide must use the real subject footprint. These rules cannot guarantee taste, but they make vague art direction inspectable.
What happened when we ran it
Our sandbox installed the yingzao/ Python project in 17 seconds. It pulled 39 packages and occupied 297 MB. The build completed in 4 seconds. Pytest ran for 17 seconds and reported 22 passed, 0 failed, and 2 skipped of 22. Pip-audit found 0 known vulnerabilities in the installed dependency set.
The checkout at commit 58c9b87 contained 121 files, about 4,228 lines of source, and used 18.2 MB before installation. It includes a tests directory, but our scan found 0 CI workflow files and no Dockerfile. Those numbers describe the deterministic Python helpers. We did not score poster quality or make an image-generation call, so the run says nothing about typography accuracy, architectural fidelity, or provider cost.
The image model still owns the hardest parts
Python handles measurements, font coverage, collision checks, alignment, image preparation, and call assembly. The chosen image model handles semantic extraction, scene reconstruction, materials, lighting, and the relationship between subject and text. That boundary keeps ordinary geometry testable. It also leaves the final artifact exposed to model errors, especially with Chinese characters and culturally specific building details.
The skill calls for one post-generation readback and up to 3 obvious problems in a note. It does not automatically regenerate, grade itself, or spend a second image credit. User feedback controls the next edit. Local lettering or edge faults can use the current image as the target, while changes to the subject, background, main layout, or reference direction return to the corrected original. That is a sensible cost boundary for supervised work.
One issue raises a context-cost question without answering it
Open issue 1 describes a Codex session compressing context from an early step while handling a roughly 600 KB uploaded photo. The reporter asks whether the cause is Codex, the skill, or a limit. There is no maintainer reply in the fetched issue and no demonstrated cause, so it would be careless to blame the skill. The report still matters because Yingzao has a long instruction file plus multiple references and generated analysis artifacts.
Test the workflow in the exact agent client you plan to use. Watch whether it reloads every reference, drops earlier facts, or compresses before the final call. A 39-package Python install can be healthy while the language-model session becomes awkward. If context use is tight, keep the user's verified text and the final generation manifest outside the chat so they survive any summarization.
Missing license terms are the adoption blocker
GitHub showed 464 stars, 1 open issue, and a last push on September 3, 2026. The latest-release endpoint returned no release, the repository had no CI workflows, and GitHub reported no license. The related Guizang Social Card Skill carries AGPL-3.0, but that separate repository's license does not apply here.
The missing license matters more than the 22 passing tests for organizations that want to copy, modify, redistribute, or bundle the skill. Ask the maintainer for explicit terms before client delivery or product integration. For personal evaluation, the workflow is worth studying: it turns poster generation into a sequence of visible decisions and gives the model one tightly prepared edit instead of repeated blind retries.

