Agent steps handle the fuzzy part, and code checks the result
In e2e@0.17.0, natural language gets a narrow job inside a normal test. You can ask an agent to upgrade a workspace, then use a role locator and an exact text assertion to prove that the status changed. A model can work through a UI whose path moves around, while the final claim stays deterministic and reviewable. Tests with no agent steps make no model calls.
The framework covers Chromium, Firefox, and WebKit through Playwright, plus iOS simulators and Android emulators through agent-device. Separate packages handle web, mobile, GitHub reporting, hosted browsers, and hosted devices. The repository is a 1,342-file pnpm monorepo with about 178,902 lines of source, so this is a sizeable testing platform rather than a thin prompt wrapper.
Verified actions replay with 0 model calls
A verified agent.act() recording can replay with 0 model calls. If a control or end state no longer matches, the agent takes over from that screen. Assertions, waits, and extraction still run live, so the cache reduces repeat calls without treating yesterday's result as today's proof.
This also creates work that ordinary Playwright users will not have: you need to inspect cache misses and run with --no-cache when diagnosing a failure. Each agent has separate cache entries keyed by its name and context. A buyer persona cannot silently reuse an admin persona's recorded path.
What happened when we ran it
Our sandbox installed commit f99e80e in 17 seconds. pnpm added 1,415 packages, and the installed tree occupied 1,861 MB. The build completed in 5 seconds. The fresh Debian container had 3 CPUs, 8 GB of RAM, Node 22, no secrets, and no elevated privileges. The checkout was 12.3 MB before installation.
Vitest ran for 647 seconds and reported 3,159 passed, 0 failed, and 2 skipped out of 3,161 tests. That is a strong result, although a nearly 11-minute suite affects contributor feedback and CI planning. Our measurements say nothing about model accuracy, token cost, browser download size, or reliability against a particular app.
The repository has seven CI workflow files, no Dockerfile, and no top-level tests directory. Package scripts run tests across the workspaces after a build. The documented environment is Node plus the browser or mobile platform tooling your targets need.
Four subscription choices sit beside API and local models
A deterministic web example can run without credentials. Agent steps can use a ChatGPT, GitHub Copilot, OpenCode Console, or SuperGrok subscription, an API key, or a local model. The docs require Node 24.8 or newer, or at least 22.22.3 on Node 22. Windows users are directed to WSL. Mobile adds Xcode and a simulator, or the Android SDK and an emulator.
The first run can provision a browser. Open issue 783 describes e2e 0.16.0 downloading full Chromium when only the headless shell was needed, then allowing Playwright's cleanup to remove older revisions in a shared cache. Teams with sealed CI runners or a shared cache should test this behavior before adoption.
Mobile support exists, but current tap defects can block a release gate
Issue 797 reports an Android locator resolving to a Continue button but tapping a different position. Issue 830 reports visible iOS controls above an Expo Router tab bar being rejected as covered; that report uses a coordinate offset to bypass the occlusion check. These concrete reports are reason to trial your real screens, not only the example app.
Both reports were open on October 4, 2026. They do not prove every mobile suite is unreliable, but accessibility snapshots, geometry, and native containers remain part of the failure surface. Reproduce your hardest navigation and form flows before migrating a release-gating suite.
Default telemetry needs an explicit team choice
CLI telemetry is on by default. The docs say it sends command and run metadata, machine class, versions, step counts, model usage totals, estimated cost, and error codes. It excludes test titles, paths, URLs, page content, screenshots, credentials, and environment-variable values. Disable it with a saved CLI setting, E2E_TELEMETRY_DISABLED=1, or DO_NOT_TRACK=1.
Debug mode prints the payload without sending it. A regulated team should encode its choice in CI. Model-backed tests also send screen evidence to the configured provider, so provider and test-data policies belong in the setup review.
The October 4 release is active work, not a stable 1.0 contract
GitHub showed 2,536 stars, 9 open issues, and 31 open pull requests on October 4, 2026. The repository was pushed that day, and e2e@0.17.0 was released the same day. The README states that APIs and configuration can change between minor releases on the way to 1.0.
The test result makes e2e credible enough for a serious trial. Keep locators for stable actions, use agent steps where the route truly changes, and preserve exact checks for money, permissions, and saved state. If Playwright or Maestro already expresses most of your suite cleanly, adding a model, replay cache, and 1,861 MB dependency tree buys complexity you may never recover.

