mrkeyoor.com_
Sun 04 Oct 15:18 UTC
Dev Toolsevaluationupdated 04 Oct 2026

e2e review

e2e is a TypeScript testing framework that lets an agent carry out a web or mobile task described in plain English, while ordinary locators and assertions check the result. It records verified agent actions and can replay them on later runs without another model call, which tackles the cost and repeatability problems of AI-driven tests.

Verdict

Our e2e run passed 3,159 of 3,161 tests, with 2 skipped, but its 1,415-package install occupied 1,861 MB, so the engineering is convincing and the local footprint is heavy. Use it when agent steps solve flows that are costly to encode with selectors and you will keep exact assertions around the result. Stay with Playwright or Maestro when deterministic automation already describes the job well, especially while e2e remains below 1.0.

We ran it

Lab card: what happened when we ran e2eScreenshot of e2e (tester.army/e2e)
Install✓ · 17s1415 packages · 1861 MB
Build✓ · 5s
Tests✓ · 647s3159 passed · 0 failed · 2 skipped of 3161 (vitest)
Repo1342 files~178,902 lines of source · 12.3 MB · 7 CI workflows

Answers from our run

Does e2e build from source?

Dependencies installed in 17 seconds (1415 packages), and the build succeeded in 5 seconds. We cloned commit f99e80e into a clean Debian container with 3 CPUs and no project-specific setup.

Do e2e's tests pass?

Yes: 3159 of 3161 passed when we ran the project's own test command (vitest). Some failures need services or credentials a bare container does not have.

Who should not use e2e?

Projects with a tight dependency budget: our pnpm install pulled 1,415 packages and occupied 1,861 MB.

What are the alternatives to e2e?

Playwright, Stagehand, Maestro. Our e2e run passed 3,159 of 3,161 tests, with 2 skipped, but its 1,415-package install occupied 1,861 MB, so the engineering is convincing and the local footprint is heavy.

Setup3/517-second install, but 1,415 packages and platform tooling
Docs5/5Detailed setup, model, cache, CI, security, and telemetry docs
Community4/52,536 stars, an October 4 push, and active issue work
Maturity3/53,159 tests passed, but APIs and packages remain below 1.0

Who it’s for

Web teams that want natural-language steps for changeable user journeys while keeping exact assertions in code.
Mobile teams prepared to run iOS simulators or Android emulators and investigate platform-specific failures.
Test engineers who can compare models, inspect agent transcripts, and keep deterministic checks around business outcomes.
Teams using Claude Code, Codex, or another coding agent that can consume the bundled testing skill and MCP setup.

Who it’s NOT for

Projects with a tight dependency budget: our pnpm install pulled 1,415 packages and occupied 1,861 MB.
Teams requiring a stable 1.0 API: the README says APIs and config may change between minor releases, and the current packages are still below 1.0.
Mobile release gates that cannot tolerate known tap workarounds: open issues 797 and 830 report incorrect Android targeting and false iOS occlusion errors.
Locked-down browser runners that forbid downloads during a test: open issue 783 reports an unnecessary Chromium download and shared Playwright cache cleanup.
Organizations unwilling to configure telemetry policy: anonymous CLI telemetry is on by default, although the docs provide three opt-out methods.

Setup reality

Our sandbox installed commit f99e80e in 17 seconds, pulling 1,415 packages and using 1,861 MB on disk. The build passed in 5 seconds. Vitest finished in 647 seconds with 3,159 passed, 0 failed, and 2 skipped out of 3,161 tests.

A deterministic web test needs Node 22.22.3 or newer, the e2e packages, an app URL, and a Playwright browser. Agent steps also need a supported subscription, API key, or local model. Mobile adds Xcode plus an iOS simulator, or the Android SDK plus an emulator.

The repository is a 1,342-file pnpm monorepo with seven CI workflows and no Dockerfile. Windows users are directed to WSL. The first browser run may download browser files, and telemetry is enabled unless you disable it.

Agent steps handle the fuzzy part, and code checks the result

In e2e@0.17.0, natural language gets a narrow job inside a normal test. You can ask an agent to upgrade a workspace, then use a role locator and an exact text assertion to prove that the status changed. A model can work through a UI whose path moves around, while the final claim stays deterministic and reviewable. Tests with no agent steps make no model calls.

The framework covers Chromium, Firefox, and WebKit through Playwright, plus iOS simulators and Android emulators through agent-device. Separate packages handle web, mobile, GitHub reporting, hosted browsers, and hosted devices. The repository is a 1,342-file pnpm monorepo with about 178,902 lines of source, so this is a sizeable testing platform rather than a thin prompt wrapper.

Verified actions replay with 0 model calls

A verified agent.act() recording can replay with 0 model calls. If a control or end state no longer matches, the agent takes over from that screen. Assertions, waits, and extraction still run live, so the cache reduces repeat calls without treating yesterday's result as today's proof.

This also creates work that ordinary Playwright users will not have: you need to inspect cache misses and run with --no-cache when diagnosing a failure. Each agent has separate cache entries keyed by its name and context. A buyer persona cannot silently reuse an admin persona's recorded path.

What happened when we ran it

Our sandbox installed commit f99e80e in 17 seconds. pnpm added 1,415 packages, and the installed tree occupied 1,861 MB. The build completed in 5 seconds. The fresh Debian container had 3 CPUs, 8 GB of RAM, Node 22, no secrets, and no elevated privileges. The checkout was 12.3 MB before installation.

Vitest ran for 647 seconds and reported 3,159 passed, 0 failed, and 2 skipped out of 3,161 tests. That is a strong result, although a nearly 11-minute suite affects contributor feedback and CI planning. Our measurements say nothing about model accuracy, token cost, browser download size, or reliability against a particular app.

The repository has seven CI workflow files, no Dockerfile, and no top-level tests directory. Package scripts run tests across the workspaces after a build. The documented environment is Node plus the browser or mobile platform tooling your targets need.

Four subscription choices sit beside API and local models

A deterministic web example can run without credentials. Agent steps can use a ChatGPT, GitHub Copilot, OpenCode Console, or SuperGrok subscription, an API key, or a local model. The docs require Node 24.8 or newer, or at least 22.22.3 on Node 22. Windows users are directed to WSL. Mobile adds Xcode and a simulator, or the Android SDK and an emulator.

The first run can provision a browser. Open issue 783 describes e2e 0.16.0 downloading full Chromium when only the headless shell was needed, then allowing Playwright's cleanup to remove older revisions in a shared cache. Teams with sealed CI runners or a shared cache should test this behavior before adoption.

Mobile support exists, but current tap defects can block a release gate

Issue 797 reports an Android locator resolving to a Continue button but tapping a different position. Issue 830 reports visible iOS controls above an Expo Router tab bar being rejected as covered; that report uses a coordinate offset to bypass the occlusion check. These concrete reports are reason to trial your real screens, not only the example app.

Both reports were open on October 4, 2026. They do not prove every mobile suite is unreliable, but accessibility snapshots, geometry, and native containers remain part of the failure surface. Reproduce your hardest navigation and form flows before migrating a release-gating suite.

Default telemetry needs an explicit team choice

CLI telemetry is on by default. The docs say it sends command and run metadata, machine class, versions, step counts, model usage totals, estimated cost, and error codes. It excludes test titles, paths, URLs, page content, screenshots, credentials, and environment-variable values. Disable it with a saved CLI setting, E2E_TELEMETRY_DISABLED=1, or DO_NOT_TRACK=1.

Debug mode prints the payload without sending it. A regulated team should encode its choice in CI. Model-backed tests also send screen evidence to the configured provider, so provider and test-data policies belong in the setup review.

The October 4 release is active work, not a stable 1.0 contract

GitHub showed 2,536 stars, 9 open issues, and 31 open pull requests on October 4, 2026. The repository was pushed that day, and e2e@0.17.0 was released the same day. The README states that APIs and configuration can change between minor releases on the way to 1.0.

The test result makes e2e credible enough for a serious trial. Keep locators for stable actions, use agent steps where the route truly changes, and preserve exact checks for money, permissions, and saved state. If Playwright or Maestro already expresses most of your suite cleanly, adding a model, replay cache, and 1,861 MB dependency tree buys complexity you may never recover.

Alternatives

ProjectWhat it isPick it when
Playwright gh↗A mature browser testing framework built around deterministic code and locators.pick this instead when predictable browser automation matters more than natural-language agent steps.
Stagehand gh↗An AI browser automation SDK for acting on and extracting from web pages.pick this instead when browser agents and extraction are the job, rather than a web and mobile test runner.
MaestroA mobile and web UI automation framework with declarative test flows.pick this instead when mobile automation is primary and you want flows that do not depend on a model.

What people are saying

  1. [github-trending] tester-army/e2e

Sources

  1. e2e README
  2. e2e quickstart
  3. Caching agent steps
  4. e2e telemetry documentation
  5. Browser provisioning issue 783
  6. Android locator issue 797
  7. iOS occlusion issue 830
  8. e2e releases

More dev tools reviews

nanoid · github-stars-history · BrokenPipe · FGOAC-scooby · plexo · window-sweaters · the whole board →