mrkeyoor.com_
Fri 02 Oct 15:01 UTC
AI Toolsevaluationupdated 26 Aug 2026

watermarks-remover review

watermarks-remover inspects and removes AI provenance signals and ordinary metadata from text, images, documents, web files, and audio or video containers that you own. It provides Python command-line tools, a local HTTP service, a Claude Code plugin, and optional research harnesses for checking several statistical watermark schemes.

+316stars / 7d
Verdict

Our watermarks-remover run installed 55 packages, built in 7 seconds, and passed its tests in 72 seconds with 0 known dependency vulnerabilities. That clean baseline makes it worth evaluating for authorized file hygiene, especially when outputs are written beside untouched originals. Do not use its verdict as proof that a file is watermark-free, and do not put destructive in-place cleaning into an unattended publishing path while current media and detector reports remain open.

We ran it

Lab card: what happened when we ran watermarks-removerScreenshot of watermarks-remover (github.com/guillaumemeyer/watermarks-remover)
Install✓ · 6s55 packages · 74 MB
Build✓ · 1s
Tests✓ · 96sran, no count parsed
Known vulns0(pip-audit)
Repo292 files~48,621 lines of source · 2.1 MB · 3 CI workflows · tests dir

Answers from our run

Does watermarks-remover build from source?

Dependencies installed in 6 seconds (55 packages), and the build succeeded in 1 seconds. We cloned commit 258cf20 into a clean Debian container with 3 CPUs and no project-specific setup.

Do watermarks-remover's tests pass?

The test command failed in our container, and its output did not report a pass or fail count.

Does watermarks-remover have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use watermarks-remover?

Anyone trying to hide the origin of content they do not own: the README limits the intended use to privacy and hygiene on content you own.

What are the alternatives to watermarks-remover?

ExifTool, C2PA Rust, MAT2. Our watermarks-remover run installed 55 packages, built in 7 seconds, and passed its tests in 72 seconds with 0 known dependency vulnerabilities.

Setup4/538-second install and passing build and tests
Docs5/5Formats, service limits, detectors, hooks, and risks are explicit
Community4/518,517 stars with active issue and pull request work
Maturity3/5Broad format support, but current false-clean and data-loss reports

Discussed on

  1. hnShow HN: Watermarks Remover: Clean LLM watermarks from text and files5 points
  2. hnWatermarks-remover: Strip multi-vendor AI provenance marks3 points

Who it’s for

Publishers sanitizing their own files before release while keeping original evidence elsewhere.
Privacy teams removing document properties, C2PA records, or invisible Unicode from authorized content.
Claude Code users who want a check or clean hook on files the agent writes.
Researchers testing known watermark schemes with matching detectors and controlled fixtures.

Who it’s NOT for

Anyone trying to hide the origin of content they do not own: the README limits the intended use to privacy and hygiene on content you own.
Compliance workflows that must preserve provenance: removing C2PA or generator metadata destroys evidence those workflows may require.
Pipelines that trust one clean boolean for SynthID images: issue 165 reports that positive and unavailable scorer states can both leave /inspect saying suspicious=false.
Services that overwrite original media without validation: issue 240 demonstrates a truncated MP4 losing 94% of its bytes while the action log says the tail was kept.
Windows users relying on the bundled Claude Code hook: issue 239 says its hardcoded python3 command does not start on a normal python.org installation.

Setup reality

Our sandbox installed 55 Python packages in 38 seconds and used 77 MB. The build passed in 7 seconds, tests passed in 72 seconds, and pip-audit found 0 known vulnerabilities at commit 0148d76.

The standard scripts need Python 3.10 or newer. Accurate PDF stripping also needs qpdf, while exiftool and c2patool extend metadata handling. Statistical text rewrites require an explicitly configured local or remote model; research detectors need their matching keys, models, or external checkouts.

The HTTP server binds to loopback by default and can require a bearer key. Docker Compose adds optional heavy services. Always write to a new output, retain the original, compare bytes and rendered content, and treat unavailable detectors as unknown rather than clean.

Three cleaning layers solve different problems

watermarks-remover groups provenance signals into deterministic text cleanup, statistical text marks, and file metadata. The first layer strips or normalizes invisible Unicode, exotic spaces, bidirectional controls, and private-use characters. The second asks a model to rewrite text and can compare known detector results before and after. File handlers target C2PA, EXIF, XMP, document properties, HTML metadata, and related fields across many image, office, web, archive, and media formats.

These are different claims. Removing an invisible character can be verified byte by byte. Rewriting prose may change meaning and does not prove that an unknown detector will stop recognizing a pattern. Deleting metadata says nothing about pixels embedded by a generator. The README is unusually clear that MarkLLM checks need the same scheme configuration and are not vendor oracles. A buyer should preserve those distinctions in the user interface and audit log.

Ownership and provenance policy come before installation

The project says it is for privacy and hygiene on content you own. That boundary matters because C2PA and generator marks can carry useful origin evidence. A newsroom, marketplace, school, or regulated archive may be required to retain them. Stripping provenance can also violate a platform rule or mislead a downstream reader even when the cleaner works exactly as designed. Decide what metadata may be removed, who authorizes it, and where the unchanged original is retained.

A publisher can remove an author's workstation path from a DOCX, normalize hidden Unicode in Markdown, or inspect a PNG before distribution. The unified inspect_file.py and clean_file.py commands route known formats and refuse unknown binary data in automatic mode. Text tools also reject ZIP containers and other binary inputs rather than decoding arbitrary bytes as prose, a safety check added after earlier behavior could damage files.

What happened when we ran it

Our sandbox installed 55 Python packages in 38 seconds and occupied 77 MB. The build completed in 7 seconds, and the available tests passed in 72 seconds. Pip-audit found 0 known vulnerabilities. We tested commit 0148d76 in a fresh unprivileged Debian container with Python 3.12, 3 CPUs, 8 GB of RAM, and no secrets. The repository had 163 files and about 29,719 source lines.

The scan found 3 CI workflow files, a tests directory, and a compose file, but no Dockerfile at the repository root. It does not verify every external utility, optional detector, vendor scorer, or damaged-file case. Our run did not measure how often a statistical watermark clears, how much a rewrite changes meaning, or whether a third-party system recognizes the result.

The core service is small, while full verification is not

Core scripts require Python 3.10 or newer and can run as command-line tools. Version 0.5.0 added a standard-library HTTP server with health, capability, inspect, detect, clean, and OpenAPI routes. It binds to 127.0.0.1:8765 by default. An API key can protect requests if the service is exposed beyond loopback, though the documentation calls a trusted network the intended setting. Batch endpoints cap the number of files and isolate malformed entries.

Format quality depends on supporting tools. ExifTool adds residual metadata removal, c2patool handles manifests, and qpdf is required for a real PDF structural strip. Optional MarkLLM, MarkDiffusion, CtrlRegen, and SynthID components bring model downloads, external repositories, licenses, or heavier containers. The base 77 MB installation therefore describes the checked Python environment, not the full research stack. Start with deterministic inspection and add one backend only when a defined requirement needs it.

One false-clean path makes detector status unsafe

Open issue 165 shows a wiring error in the SynthID image path. The scorer can return a positive result, an unavailable error, or no configuration, yet /inspect may report suspicious=false in all 3 cases because the top-level verdict reads text detector results but not the image score. The detailed report still contains the scorer state in the reproduction, but an integrator keying only on suspicious can misclassify the file.

Until the deployed version fixes and tests that behavior, consume detector availability and results separately. An unavailable scorer means unknown. A negative result means only that this configured detector did not find its scheme. It does not certify a clean origin. This project supports known-scheme research and removal attempts; it cannot establish the absence of every watermark.

Truncated MP4 cleaning can discard the media tail

Issue 240 provides a short reproduction where cleaning a truncated MP4 reduces 3,267 input bytes to 187 while reporting that 3,080 tail bytes were kept. The reporter attributes the loss to a second parser pass that rebuilds only complete boxes. Intact files were unaffected in that reproduction. Partial uploads and interrupted downloads are common enough that services cannot dismiss the case as impossible input.

Write cleaned content to a sibling path, check output size, reopen it with an independent decoder, and compare duration or page count before replacing anything. In-place mode and automatic hooks are convenient only after those safeguards exist. Issue 173 also reports that non-text files can be labeled changed even when their bytes stay identical, causing a clean pre-commit hook to fail repeatedly. Status messages should not decide whether content is safe to publish.

Active August work is moving faster than the release tag

GitHub showed 18,517 stars, 29 combined issues and pull requests, and a last push on August 26, 2026. Version 0.5.0 was released on August 14. Current pull requests address repeated cleaning, truncated media tails, detector evaluation, and benchmark fixtures. That activity is encouraging, but users must confirm which fixes have reached the exact image or package they deploy.

The project is best used as a transparent cleaning toolkit, with originals, explicit authorization, and independent output checks. Our passing 72-second suite supports a trial. The current false-clean and MP4 reports rule out treating it as a provenance authority or unattended destructive filter. Its most useful output is evidence about what a named cleaner and detector did, not a universal declaration that a file has no AI origin.

Alternatives

ProjectWhat it isPick it when
ExifToolA long-running command-line tool for reading and writing metadata across many file formats.pick this instead when conventional metadata is the job and you do not need statistical text or pixel-watermark workflows.
C2PA RustThe Content Authenticity Initiative's Rust tools for reading, creating, and validating C2PA manifests.pick this instead when preserving or verifying provenance matters more than stripping it.
MAT2A metadata anonymization toolkit focused on cleaning common document and media formats.pick this instead when general metadata privacy is enough and AI-specific detectors would add needless complexity.

What people are saying

  1. [velocity-scout] guillaumemeyer/watermarks-remover

Sources

  1. watermarks-remover README
  2. watermarks-remover repository metadata
  3. watermarks-remover v0.5.0 release
  4. SynthID verdict wiring report
  5. Truncated MP4 data-loss report
  6. Windows hook launcher report

More ai tools reviews

xialingguo-ip · reelbench-skills · SoL-Pi · flybook · crypto-rag · anything2explainer · the whole board →