Three passes put story structure ahead of word choice
Sepia's fiction route works in 3 passes: narrative architecture, discourse flow, then surface style. The first pass looks for explained themes, plots that resolve too neatly, narrow emotional description, sparse character networks, and endings built around acceptance. The next pass checks paragraph and sentence patterns. Familiar cleanup such as clichés and register comes last. That ordering is the project's main distinction from tools that swap vocabulary while leaving the story's shape intact.
The project also includes a 30-feature diagnosis rubric and model-specific fingerprints drawn from published vendor guidance and cited studies. It says those fingerprints apply only when the executing or writing model is known. Vendors without published guidance are recorded as consulted rather than assigned guessed traits. That restraint is useful, though every diagnosis still comes from instructions interpreted by a language model, so editors must decide whether a finding fits the passage in front of them.
Four operation modes set different editing boundaries
Four documented operations cover write, review, refactor, and recreate. Review reports problems without changing text. Refactor aims for minimal edits, while recreate permits a full rewrite from the source facts and intent. A fifth Hemingway entry applies the included voice profile to fiction writing or refactoring. These distinctions make Sepia easier to supervise than a single "humanize" command because the user chooses how much authority the agent receives before it touches the draft.
Professional prose uses 5 venue-specific overlays: release notes, pull request or issue replies, postmortems, tickets, and technical articles. The rules favor user impact in releases, file-and-line evidence in code review, timelines in incident reports, testable acceptance criteria in tickets, and a concrete problem in technical writing. This part is more practical than the fiction theory. It gives an agent a reason to treat a production incident differently from a blog post instead of applying one general ban list everywhere.
What happened when we ran it
Our 2026-08-31 measurement setup used a fresh unprivileged container with 3 CPUs and 8 GB of RAM. The sandbox did not run Sepia at commit 361e82e because it detected no supported ecosystem, and the checkout had no Dockerfile. There are therefore no measured install, build, test, dependency, timing, or vulnerability results for this project. That absence matters: the README's installation claims were not reproduced by our harness.
Sepia is mostly a collection of Markdown skills, reference files, research notes, packaging metadata, and evaluation material, so a package-manager build is not necessarily the right test. Even so, commit 361e82e supplied no target our sandbox knew how to execute. A useful adoption check is a before-and-after corpus from your own work, scored blind by editors who know the venue. The repository's badges and current test directory do not replace that local editorial test.
The evidence base is public, while intervention proof is limited
The Sepia README bases its fiction approach on StoryScope, a study covering 61,608 stories from humans and 5 frontier models. The project interprets that work as evidence that narrative structure can remain detectable after surface editing. Its research directory links the papers and separates measured findings from project inferences. Readers can follow the chain instead of accepting an unexplained checklist, which is a real advantage for a prompt-based editing tool.
The repository is equally direct about its limits. Voice-skill composition is described as experimental and grounded in 1 blind-review example rather than measured evidence. The v0.8.0 release notes say the new deletion and reversion tests are untested as an intervention. They also explain that an earlier content-word threshold was removed because a corpus-level effect could not justify a rule for one passage. That is sound editorial caution, though it leaves effectiveness for adopters to measure themselves.
Seventy-seven client targets do not mean verified behavior on 77 agents
The Skills CLI route advertises support for 77+ agents, while Sepia provides native plugin packaging for Claude Code, Codex, Grok Build, and Antigravity. The README narrows the claim: those 4 native installs were exercised only far enough to confirm installation and entry discovery. Runtime behavior was not checked platform by platform, and clients outside those four were not exercised by the maintainer. Teams should treat portability as a format claim, then verify prompt loading and file access in their chosen agent.
Installation has another sharp edge. The thin operation wrappers depend on the canonical sibling skill, so copying only sepia-review or sepia-refactor is unsupported. The full package keeps one canonical SKILL.md instead of separate platform forks, which reduces drift. Project-scoped users can commit the skill under .agents/skills or .claude/skills; user-scoped installs follow each client's plugin command. No hosted service or API credential is part of the documented path.
Version 0.8.0 is active, but age still limits the maturity case
GitHub records the repository's creation on August 28, 2026, its last push on September 5, and the v0.8.0 release on that same September date. It had 2,384 stars and 2 combined issues and pull requests when fetched. One open issue, issue 227, keeps automatic voice suggestions off professional routes until the report format has a closed vocabulary. That is specific, current maintenance activity, not proof of long-term stability.
Sepia is worth trying on copies of real drafts if its structural editing matches the problem you have. The 361e82e sandbox result cannot tell us whether it installs cleanly or produces better prose, and the project does not promise detector evasion. Its strongest case is narrower: clear editing modes, visible source material, and unusually frank limits around what has been verified. Keep a human editor responsible for the final judgment, especially when refactor or recreate can change meaning.
