Five output formats cover pages, agents, and ingestion jobs
Webclaw takes a URL and returns Markdown, compact LLM text, plain text, JSON metadata, or cleaned HTML. Flags can keep the main content, include selected elements, remove navigation, follow same-origin links, map URLs, compare snapshots, and extract brand assets. That range covers documentation ingestion and monitoring without forcing every consumer to parse the same raw HTML. The extraction core is a separate Rust crate with no network input, which makes it usable inside another application.
The surrounding tools make the engine accessible. There is a terminal binary, MCP server, REST service, Docker image, and SDKs for TypeScript, Python, and Go. npx create-webclaw detects supported agent clients and writes MCP configuration. A separate skill exposes scraping, crawling, mapping, extraction, summaries, diffs, brands, and search to Claude Code and similar agents. Convenience is high, but the local and hosted tool lists are not identical.
Local extraction stops where rendering and bot defenses begin
Ordinary public pages can be fetched and cleaned locally with no Webclaw account. Search and research are hosted-only, while JavaScript rendering, protected-site handling, async jobs, watches, and production tracking also belong to webclaw.io. A local crawl against a static documentation site is therefore a different product experience from extracting an application shell or a site that challenges automated clients. Test representative URLs before choosing an architecture.
Structured extraction and summaries add another dependency. Webclaw can use Ollama or configured OpenAI-compatible, Anthropic-compatible, or OrcaRouter endpoints. Proxies can be supplied as one URL or a pool file. Those options provide escape routes for difficult pages, but the operator still owns crawl rate, proxy legality, credentials, robots policy, and the reliability of any external model. An API key does not turn every target into permissible data.
What happened when we ran it
Our Rust sandbox installed 384 packages in 17 seconds from a 7.5 MB checkout. commit 3af4704 had 182 files and about 43,355 source lines. The build succeeded in 171 seconds. That is a long compile relative to the repository size, but it is a one-time cost that produces native binaries rather than a large interpreted runtime. The repository includes 2 CI workflows, a Dockerfile, and a compose file.
Cargo tests completed in 105 seconds with 772 passed and 0 failed. The run used an unprivileged container with 3 CPUs, 12 GB of RAM, no secrets, and the supplied Rust lab image. Nothing in those results measures extraction accuracy, network success, bot bypass, crawl throughput, or hosted service latency. They establish a narrower and still useful fact: the checked-out engine built and its complete reported test command passed in our environment.
Scanned PDFs currently end without an OCR handoff
PDF text extraction exists, but issue 104 identifies a practical gap for image-only files. Webclaw buffers the response, attempts PDF extraction, then returns EmptyPdf and drops the bytes. A caller that owns OCR or vision fallback must fetch the source again, which can waste bandwidth or retrieve changed content. The request asks for a bounded, opt-in artifact result containing the exact bytes and metadata. It does not claim that Webclaw should become an OCR engine.
That distinction matters in document pipelines. A 772-test pass says the implemented behavior works as tested, while scanned contracts, brochures, and receipts may still yield no text. If PDFs are important, assemble a fixture set containing searchable text, scans, mixed pages, large files, redirects, and unusual encodings. Record final URLs and hashes so a fallback can prove it processed the same artifact. Until issue 104 is addressed, decide whether a second fetch is acceptable.
MCP works on the older revision, with 7 newest-draft failures reported
The MCP server gives agents a direct scrape and crawl interface through stdio. Release v0.6.21 upgraded its MCP library so one unreadable input line no longer ends the session. Issue 117 then tested the package against 2 protocol revisions. The report says the 2025-11-25 revision had 0 failed requirements, while the 2026-07-28 draft had 7 failures around server discovery and version-less tool listing.
This does not make the current MCP server unusable. It means client compatibility should be tested rather than inferred from the MCP label. Pin the Webclaw package and client version together, run initialization and tool listing in CI, and try one real scrape. The installer changes client configuration, so teams should review the generated command and package source before distributing it. Webclaw also offers a Claude Code skill, which deserves the same package review.
AGPL licensing and August maintenance favor internal tools
The engine is AGPL-3.0, a reasonable fit for internal use and open network services that comply with its source obligations. A proprietary hosted product needs legal review before embedding or modifying the server. The hosted webclaw.io service is separate from the open repository, and the README is clear about which capabilities depend on it. Avoid designing around hosted-only functions if a fully offline deployment is mandatory.
GitHub showed 2,305 stars, 3 combined open issues and pull requests, and a last push on August 26, 2026. Version 0.6.21 shipped 10 days earlier. The small queue and recent release are positive signals, though star and issue counts cannot substitute for extraction trials. Webclaw deserves a pilot for static pages and agent context because our build and all 772 tests passed. Keep a browser tool and OCR path available for the pages its local engine intentionally does not handle.

