e2e is a free, Apache-2.0 end-to-end testing framework from TesterArmy that lets you write a browser or mobile test as a plain-English goal, has an AI agent drive the app to reach it, and then checks the result with ordinary locators and assertions. It picked up 1,390 GitHub stars in a single day this week and sits at about 7,500 total, and the setup below takes roughly ten minutes on a machine that already has Node.js 24.

  • npx e2e init writes a config, an example test and an agent skill for Claude Code and Cursor, then npm install pulls in the runner.
  • Tests without agent steps need no model and no API key. We ran two of them on Windows 10 with Node 24.17 and both passed in under two seconds.
  • Agent steps need a model: a ChatGPT, Copilot, OpenCode or SuperGrok subscription through npx e2e login, an API key such as AI_GATEWAY_API_KEY, or a local OpenAI-compatible server.
  • Once an agent.act step passes, its actions are cached and replayed on the next run with no model calls until the app changes.

The exact steps, start to finish

  1. Check your runtime. e2e needs Node.js 24.8 or newer, or 22.22.3 or newer on the 22 line.
    $ node --version
    $ npm --version
  2. Create the test project. Run this in your app's directory and pick Web (Playwright), then a model provider, or None for tests without AI.
    $ npx e2e init
  3. Install the dependencies the wizard added.
    $ npm install
  4. Start your app. Run your usual dev server so something answers at http://localhost:3000, or change app.url in e2e.config.ts to wherever it runs.
  5. Run the example test. No model calls happen yet, so no key is needed.
    $ npx e2e run
  6. Connect a model for agent steps. Either sign in with a subscription you already pay for, or set an API key. For Vercel AI Gateway, create a key from the AI Gateway docs and export it in the same terminal.
    # a ChatGPT Plus or Pro plan
    $ npx e2e login openai
    # or an API key for the default gateway() model
    $ export AI_GATEWAY_API_KEY=your-key-here
  7. Write your first agent step. Save a file like tests/agent.e2e.ts with one agent.act goal and one agent.assert check (the full file is further down), then run just that file.
    $ npx e2e run tests/agent.e2e.ts
  8. Watch it work. Add --headed to see the browser while the agent drives it.
    $ npx e2e run --headed

What npx e2e init drops into your repo

We ran the wizard in a throwaway folder holding a 20-line to-do page. With --yes it skips the questions and picks web plus Vercel AI Gateway. It did not install anything on its own. It added four dev dependencies to package.json (e2e, @e2e-dev/web, ai and zod), a test:e2e script, an e2e.config.ts, one example test, and an .agents/skills/e2e folder of nine files linked into .claude/skills. It also wrote .mcp.json and .cursor/mcp.json entries for an e2e MCP server, which is how a coding agent can later run and debug your suite for you.

RelatedStrix Setup: Run an AI Penetration Tester on Your Code

The generated config is short enough to read in one go:

import type { E2EConfig } from 'e2e';
import { web } from '@e2e-dev/web';
import { gateway } from 'ai';

export default {
  agents: { default: { model: gateway('openai/gpt-6-luna-fast') } },
  targets: [{ engine: web(), app: { url: 'http://localhost:3000' } }],
} satisfies E2EConfig;

That one line under agents is the only place a model appears. There is no default model and no shared key variable, so if you never call the agent, the config can point anywhere and nothing is billed.

How an e2e agent step runs the first time and on later runs A test file calls agent.act. On the first run a model drives the browser through Playwright and the verified actions are saved to a replay cache. On later runs the cache replays them without a model, and the agent takes over only if the app changed. E2E · ONE GOAL, TWO PATHS agent.act() plain-English goal Model (run 1) subscription, key, local Replay cache run 2+, no model calls Browser via Playwright expect() pass / fail verified steps saved If the app changes, replay hands off to the model from the current screen. genztech.blog
Fig 1 The model only drives an agent step until a later check verifies it. After that, the replay cache repeats the recorded clicks for free.

Does e2e run on Windows without WSL?

The quickstart says to run inside WSL on Windows. We tried the web engine natively first, in Git Bash on Windows with Node 24.17.0, and it worked. With a static page served on port 3000, npx e2e run printed this:

 RUN  e2e v0.18.0
      run 01a11948-... · targets: web
      model gateway/openai/gpt-6-luna-fast

 ✓  web  tests/example.e2e.ts (1 test) 1.99s
   ✓ app opens 1.99s

 Test Files  1 passed (1)
      Tests  1 passed (1)
   Duration  2.87s
     Report  .e2e\report.json

That is one data point, not a support promise, and the mobile engine is a separate story: iOS needs Xcode with a simulator runtime, Android needs the Android SDK with an emulator. If anything odd happens natively, WSL is the documented path. On macOS and Linux the commands in the checklist are the whole install.

The first time you run it, the CLI prints a notice that it collects anonymous usage data. The README says that covers commands, engines and where runs fail, never test content or credentials. Turn it off with npx e2e telemetry disable or by setting E2E_TELEMETRY_DISABLED=1.

Writing a to-do test with locators, then with an agent

A deterministic test reads a lot like Playwright or Testing Library code. This one fills a labelled input, taps a button and checks the new list item. It passed in 665 ms on our machine, with no model involved:

import { test } from '@e2e-dev/web';
import { expect } from 'e2e';

test('adds a todo', async ({ app, screen }) => {
  await app.open('/');
  await screen.getByLabel('New todo').fill('Ship the release');
  await screen.getByRole('button', 'Add').tap();
  await expect(screen.getByRole('listitem')).toHaveText('Ship the release');
});

The agent version swaps the three locator lines for a goal. params fills the placeholders, agent.assert asks the model to judge the screen, and you can still finish with a strict expect:

test('the agent adds a todo', async ({ app, agent }) => {
  await app.open('/');
  await agent.act('add a todo called {title}', { params: { title: 'Write the tests' } });
  await agent.assert('the list shows Write the tests');
});

The docs are firm on one point: give agent.act one goal per call. A step that tries to sign up, upgrade and log out in one sentence runs into the step budget and is harder to cache. If you do not want to write a test at all yet, npx e2e explore takes a goal like "explore the checkout flow like a first-time buyer" and writes findings to .e2e/report.json, which is a decent way to decide which flows deserve a real test.

RelatedOpenViking Setup: Long-Term Memory for Claude Code and Codex

Fixing "MODEL_PROVIDER_FAILED: Unauthenticated request to AI Gateway"

This is the error you will see first if you skip step 6. We hit it on purpose by running the agent test with no key set. The step failed in 246 ms with MODEL_PROVIDER_FAILED and the message "To authenticate, set the AI_GATEWAY_API_KEY environment variable with your API key." The fix is to export the key in the same terminal, switch to npx e2e login with a subscription, or replace the gateway() import with another AI SDK provider. Pick a model that supports tool calls and images, since the built-in agent needs both.

The other errors worth knowing before your first real suite:

  • APP_UNREACHABLE. Nothing answers at app.url. Start the dev server, check the port, or add app.command so the runner starts it for you with a log file.
  • APP_ALREADY_RUNNING. The runner wanted to start your app but the port is taken. Stop the other server or set reuseExisting: true for local runs.
  • LOCATOR_NOT_FOUND. Usually a wrong accessible name. Run with --headed and read the markup.
  • Next.js 16 pages that never hydrate. If the target opens 127.0.0.1 while next dev identifies as localhost, add that host to allowedDevOrigins in next.config.ts.

When a cached step misbehaves, npx e2e run --no-cache forces the agent to run live, and --debug prints timings and a step table.

e2e next to Playwright and Cypress

e2ePlaywright TestCypress
How you describe a stepLocators or a plain-English goalLocators and codeSelectors and code
Needs an AI modelOnly for agent stepsNoNo
Native mobile appsiOS and Android engineNo (mobile emulation in browsers)No
Browser enginesChromium, Firefox, WebKit via PlaywrightChromium, Firefox, WebKitChromium-family, Firefox, WebKit (experimental)
MaturityPre-1.0, APIs can changeMatureMature
LicenseApache-2.0Apache-2.0MIT

Under the hood the web engine is Playwright, so this is less a rival to Playwright than a layer on top of it. The interesting bit is the cache: most "AI testing" tools call a model on every run, which makes a big suite slow and expensive. Here the model only pays for the first verified pass, and the run summary tells you how many steps were replayed, handed off or missed.

Would we move a working Playwright suite over today? No. The README says e2e is still on the way to 1.0 and config can change between minor releases, and version 0.18.0 is not a long track record. Where it earns a try right now is the flows that keep breaking selector-based tests, like onboarding wizards and checkout pages that a designer reshuffles every sprint, and mobile apps where Playwright cannot follow. Start with one agent test beside your existing suite, look at the cache numbers after a week, and decide from there.

Primary sources