Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
29 commits
Select commit Hold shift + click to select a range
d7ff1f8
feat: add local native mobile testing
perixtar Sep 19, 2026
3efc612
fix: harden native mobile execution
perixtar Sep 19, 2026
42ea8c4
fix: close native safety gaps
perixtar Sep 19, 2026
75e556e
fix: fail closed on native privacy evidence
perixtar Sep 19, 2026
f946fc1
fix: reject private prose before planning
perixtar Sep 19, 2026
b275977
fix: close remaining native safety gaps
perixtar Sep 19, 2026
b1529c4
fix: reject hidden native secrets and unsafe captures
perixtar Sep 19, 2026
47062e4
Harden native execution evidence and privacy
perixtar Sep 19, 2026
be3d5e3
Align native safety checks with SDK contracts
perixtar Sep 19, 2026
c6f0e36
Fail closed on native prose and recording uncertainty
perixtar Sep 19, 2026
64785ca
Bind native planner goals and disclosed identities
perixtar Sep 19, 2026
1a2f7d1
Reject conditional native intent and unbound credential prose
perixtar Sep 19, 2026
b99a1e4
Keep safe app-specific login titles usable
perixtar Sep 19, 2026
183184f
Ignore decorative native ancestors in target identity
perixtar Sep 19, 2026
b39ff89
Keep native target identity exact and hide fixture decoration
perixtar Sep 19, 2026
fd1e3eb
Hide noninteractive fixture wordmark from native accessibility
perixtar Sep 19, 2026
7d60d48
Give native cancellation gate a visible catalog baseline
perixtar Sep 19, 2026
d46ad19
Settle native relaunch evidence and reject ambiguous multiline intent
perixtar Sep 19, 2026
1af55f7
Treat quoted native literals as data and stabilize live gate baseline
perixtar Sep 19, 2026
85b152c
Wait for exact saved native targets after relaunch
perixtar Sep 19, 2026
2456124
Fail closed on unbound native actions and private values
perixtar Sep 19, 2026
fb0b80f
Accept SDK install responses without device metadata when open verifi…
perixtar Sep 19, 2026
96a275e
Verify omitted iOS SDK quality with corroborated owned trees
perixtar Sep 19, 2026
738dc3c
Document native mobile results and publish demo video
perixtar Sep 19, 2026
dbcd27d
Clarify scope of published mobile trial cohort
perixtar Sep 19, 2026
f60f8db
Guard mobile intent, replay identity, and live previews
perixtar Sep 19, 2026
e980b45
Scope mobile benchmark provenance in public data
perixtar Sep 19, 2026
87236dc
Run mocked native tests across macOS and Linux
perixtar Sep 19, 2026
f539063
Keep mock native worker alive for deadline test
perixtar Sep 19, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -17,3 +17,4 @@ OPENROUTER_PLANNER_MODEL=openai/gpt-4.1-mini
# Models receive fixture references; the executor resolves the exact values.
E2E_TEST_EMAIL=
E2E_TEST_PASSWORD=
E2E_TEST_INVALID_PASSWORD=
4 changes: 4 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
@@ -1,4 +1,8 @@
node_modules/
Pods/
examples/mobile-app/ios/
examples/mobile-app/android/
examples/mobile-app/.expo/
dist/
coverage/
.env
Expand Down
44 changes: 44 additions & 0 deletions MOBILE_TEST_RESULTS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,44 @@
# Native mobile validation

Measured September 19, 2026 on implementation commit `96a275e8c452fbc40ad5b52d83c538664edc0b64`. **The local native alpha met the [prewritten mobile exit criteria](TEST_PLAN.md#native-mobile-alpha-exit-criteria) on the owned Jev Shop fixture.** This is evidence for those five accessible flows on dedicated virtual devices, not a reliability claim for arbitrary apps or physical phones.

The complete, sanitized [60-suite / 300-case first-attempt record](docs/benchmarks/native-mobile-2026-09-19.json) includes each trial's baseline, verdict, checked counts, timings, requests, model, and billed API cost. It comprises two isolated 30-suite runs; every first attempt within those runs is included without retry or replacement. Earlier diagnostic runs on other simulator and concurrent-device setups hit host/SDK failures and are not pooled into this cohort.

| Platform and mode | First attempts | PASS | Correct FAIL | BLOCKED | Per-case median / p95 | Jev requests | Billed API cost |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| iOS · healthy discovery | 50 | 50 | — | 0 | 27.21 / 30.24 s | 250 | $0.014207 |
| iOS · seeded faults | 50 | **0** | **50** | 0 | 29.17 / 31.66 s | 250 | $0.014238 |
| iOS · saved replay | 50 | 50 | — | 0 | 25.24 / 36.57 s | 30 | $0.001677 |
| Android · healthy discovery | 50 | 50 | — | 0 | 10.44 / 13.10 s | 259 | $0.012789 |
| Android · seeded faults | 50 | **0** | **50** | 0 | 12.36 / 14.94 s | 261 | $0.012835 |
| Android · saved replay | 50 | 50 | — | 0 | 8.70 / 11.12 s | **0** | **$0** |

Every row has ten runs for each of five cases. Healthy and replay reached 10/10 per case; all five seeded faults produced 10/10 observed, correct FAIL. The aggregate billed provider cost was **$0.055746222**. The median is per-case elapsed time; suite setup and aggregate artifact phases are retained in the trial record. Warm healthy medians were 27.20 s on iOS and 10.57 s on Android, below the 60-second gate.

Replay made **zero prose-planner requests** on both platforms. iOS made 30 Jev repair requests across 50 replay cases because an extra accessibility ancestor changed the saved control context; the runner refused to treat that target as an exact match and selected it again from fresh evidence. Android replay made zero model requests. Replay on iOS is therefore **not** model-free. The original evaluation script mistakenly also required zero Jev calls and marked its iOS summary failed. The [written gate](TEST_PLAN.md#native-mobile-alpha-exit-criteria) requires zero **planner** requests, while the [replay behavior](docs/MOBILE.md#reports-replay-limits-and-stop) allows Jev to repair a stale control. The raw script result is preserved; the published aggregate recomputes the written gate from the immutable trials, records `originalStrictEvaluatorPassed: false` and `strictZeroJevReplayDiagnosticPassed: false`, and the script's predicate has been corrected for future runs. The correction changes only evaluation reporting, not any mobile trial or runner behavior.

## Method and other gates

- Five authored cases in [examples/mobile.cases](examples/mobile.cases): invalid login rejection, filtered catalog search, quantity/total, cart removal, and cart persistence after relaunch. The fault run independently seeds invalid-login, wrong-filter, wrong-total, ineffective-removal, and lost-persistence variants. The local control service confirmed that every one of the 60 suites consumed its intended baseline before actions.
- One platform ran at a time: dedicated iOS Simulator with iOS 26.5, then Android Emulator with Android 16. Each run used the same app, cases, limits, `--planner off`, OpenRouter's `typesafe/jev-1.13-20260917`, and pinned `agent-device@0.21.6`. The runner used real native accessibility/actions and independent expectation checks. The 50 replay cases per platform reused only a complete healthy flow, after a fresh baseline reset.
- A separate interpretation cohort at the measured implementation SHA compiled **40/40** independently authored prose cases faithfully (20 per platform). It made 40 `openai/gpt-4.1-mini` requests via OpenRouter, zero Jev requests, and billed $0.0248056. These planner-only checks are separate from the planner-off 300-case matrix.
- Real SDK checks passed on both platforms: owned-device selection and competing lease, fresh references, Unicode text replacement, readable switch state, safe recording/discard, and cancellation during selection, snapshot, fill, and press/settle. Android's SDK check preceded the final iOS-specific quality-verdict change; Android's full 150-case matrix ran at the measured implementation SHA.
- A clean `npm pack` install with the optional SDK ran one healthy PASS (exit 0) and one seeded login FAIL (exit 1) on **each** device. CLI JSON and saved JSON/HTML agreed. A second install with `--omit=optional` compiled an explicit native plan offline and returned exit 2 with missing-SDK guidance from `doctor`. The local workbench rendered Android PASS and FAIL reports with matching JSON verdicts and no page errors. Type checks, the build, and all 53 contract/browser tests passed.

The [10-second video](docs/assets/mobile-demo-10s.mp4) and [real-time recorded segments](docs/assets/mobile-demo-real-time.mp4) show an additional, longer 12-check cart journey on both devices, captured from the same implementation SHA. Full case elapsed time was 54.535 s / $0.000603498 on iOS and 70.640 s / $0.000615846 on Android. Both passed 12/12. The 10-second edit uses **3.47×** iOS and **11.25×** Android playback for the eligible clips; full-run timers jump over private input and relaunch gaps. The 58.93-second companion plays only the shareable segments at 1×. Neither video is a claim that the full tests ran in 10 seconds. The recorded segments start after credential/search entry, and the artifact verifier plus manual frame inspection excluded input/known-secret screens.

The 300-case matrix, interpretation cohort, and filmed runs remain attributed to `96a275e8c452fbc40ad5b52d83c538664edc0b64`. Subsequent review fixes to preflight wording, saved build identity, and live preview publication have focused tests and separate final-candidate smoke checks; they do not retroactively change these measured trials.

To reproduce the controlled matrix after [building and installing the fixture](docs/MOBILE.md#reproduce-the-owned-shopping-fixture), use the opt-in command below. It makes paid calls and needs your own OpenRouter key, booted dedicated devices, and a running local fixture baseline. The linked public data retains the source SHA, timing phases, and every first attempt; exact device IDs and OS build details remain in private maintainer logs.

```sh
npm run build
node scripts/evaluate-mobile.mjs --live --platform ios \
--ios-device IOS_ID --rounds 10 --max-cost 2 \
--out .jev-e2e/mobile-evaluation/ios
node scripts/evaluate-mobile.mjs --live --platform android \
--android-device emulator-5554 --rounds 10 --max-cost 2 \
--out .jev-e2e/mobile-evaluation/android
```

Native accessibility is still a boundary: sparse or incomplete trees, ambiguous controls, unexposed switch state, and unsupported custom surfaces produce BLOCKED rather than a guessed PASS. Intermittent host simulator startup/lease trouble occurred while preparing the benchmark; the published 300-case run used healthy isolated devices and includes every first attempt in those registered rounds. No claims are made for hosted devices, physical phones, payments, biometrics, or complex WebViews.
43 changes: 33 additions & 10 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,19 +2,29 @@

# jev-e2e

**Test your website in plain English. Get evidence for every result.**
**Test websites and native apps in plain English. Get evidence for every result.**

[![Status: alpha](https://img.shields.io/badge/status-alpha-orange)](TEST_RESULTS.md)
[![Node: 22+](https://img.shields.io/badge/node-22%2B-339933)](https://nodejs.org/)
[![Node: 22.12+](https://img.shields.io/badge/node-22.12%2B-339933)](https://nodejs.org/)
[![License: MIT](https://img.shields.io/badge/license-MIT-blue)](LICENSE)

[Quick start](#quick-start) · [Write a test](#write-a-test) · [Benchmarks](#live-ebay-benchmark) · [CLI reference](docs/USAGE.md) · [Contribute](CONTRIBUTING.md)
[Quick start](#quick-start) · [Native mobile](docs/MOBILE.md) · [Write a test](#write-a-test) · [Results](MOBILE_TEST_RESULTS.md) · [CLI reference](docs/USAGE.md) · [Contribute](CONTRIBUTING.md)

</div>

Describe a flow and what should be true at the end. jev-e2e turns it into a test plan, uses Jev to select controls on the page, and runs the test with Playwright. Each result is **PASS**, **FAIL**, or **BLOCKED**, with an HTML report, JSON, and masked screenshots.
Describe a flow and what should be true. jev-e2e turns it into a test plan, uses Jev to select accessible controls, and runs it with Playwright or a local simulator/emulator. Each result is **PASS**, **FAIL**, or **BLOCKED**, with an HTML report, JSON, and privacy-aware evidence.

**Local alpha:** CLI and browser workbench for Chromium websites. Install from source; an npm release is not yet available.
**Local alpha:** CLI and workbench for Chromium websites, iOS Simulator, and Android Emulator. Install from source; an npm release is not yet available.

## Watch one test on iOS and Android

<a href="docs/assets/mobile-demo-10s.mp4"><img src="docs/assets/mobile-demo-poster.jpg" alt="Jev Shop native test on a large iOS simulator screen with elapsed time and billed Jev cost" width="420"></a>

[Watch the 10-second edit](docs/assets/mobile-demo-10s.mp4) · [Watch the real-time recorded segments](docs/assets/mobile-demo-real-time.mp4) · [Read the native setup](docs/MOBILE.md)

One authored case signs in, searches for a lamp, adds it to the cart, changes the quantity, checks the total, relaunches, verifies persistence, removes it, and enables notifications. The same React Native fixture ran on **iOS Simulator and Android Emulator**; all **12 explicit checks passed** on each recorded run. The timers and billed Jev costs come from those runs. The short edit speeds up the iOS and Android footage separately, while the second video plays the shareable recorded segments in real time. Screens with input fields (including sign-in and search) and transient relaunch screens are excluded for privacy, so the footage has capture cuts. These two runs are a demonstration; repeated reliability results are reported separately in [native validation](MOBILE_TEST_RESULTS.md).

In a separate controlled evaluation of five cases × ten first attempts per mode and platform, healthy runs passed **50/50** on each device, seeded faults produced **50/50 correct FAIL** on each, and saved replay passed **50/50** on each. iOS replay used Jev to repair 30 stale controls; Android replay used no model calls. See the [method, cost, timing, and limitations](MOBILE_TEST_RESULTS.md) and [all 300 sanitized case results](docs/benchmarks/native-mobile-2026-09-19.json).

## Watch the eBay comparison

Expand All @@ -29,14 +39,14 @@ This is a UI execution experiment with a shared human-authored plan and an exten
## Why jev-e2e?

- **Write cases in plain English.** Use an optional planner for prose, or explicit `Goal`, `Step`, and `Expect` templates without it.
- **Check the outcome.** Playwright verifies expectations independently. Missing evidence or unsupported requirements produce BLOCKED.
- **Check the outcome.** The runner verifies expectations independently of Jev's choices. Missing evidence or unsupported requirements produce BLOCKED.
- **See what happened.** Reports include expected and observed values, actions, timing, provider usage, and screenshots.
- **Replay successful flows.** Reuse saved controls and recheck assertions. An unchanged flow can replay with zero model calls; stale targets require Jev to repair them.
- **Run locally with limits.** Use your own OpenRouter key, authentication fixtures, request limits, deadlines, and cost budget. Stop execution from the workbench or with Ctrl+C.

## Quick start

Requires **Node.js 22+** and one **OpenRouter API key**.
Requires **Node.js 22.12+** and one **OpenRouter API key**.

### 1. Install from source

Expand Down Expand Up @@ -82,6 +92,18 @@ npm run ui

Open the printed workbench URL, normally **http://127.0.0.1:4007**. Point it at the demo on **http://127.0.0.1:4177**, review a plan, run it, and inspect the evidence. The demo command creates demo-only fixtures under `.jev-e2e/demo`. Use a fresh project name for repeated create tests; a fresh browser context does not reset application data.

### 5. Test a native app

```sh
node dist/cli.js doctor --platform ios
node dist/cli.js devices --platform ios
node dist/cli.js plan --platform ios --app com.example.app \
--cases mobile.cases --fixtures fixtures.json --out reviewed.json
node dist/cli.js run --plan reviewed.json --device EXACT_ID
```

Replace `ios` with `android` for an Android Emulator. Native tests use the same case, review, replay, and report flow; the local app needs accessible controls. Explicit `Step:` lines bypass the optional prose planner; `--planner off` makes that choice explicit. See the [native setup and case reference](docs/MOBILE.md).

## Write a test

Save a case in a `.cases` file. With `--planner on`, describe the flow and give a concrete expectation:
Expand Down Expand Up @@ -125,16 +147,17 @@ For product validation, see the separate [controlled-demo test results](TEST_RES
| --- | --- |
| Optional prose planner | Converts your case into semantic steps, input bindings, and explicit expectations. Saved plans need no new interpretation. |
| Jev | Selects from observed controls and navigation choices using OpenRouter's native Decisions API. |
| Playwright | Executes actions and checks expectations independently. |
| Playwright / agent-device | Executes browser or local native actions. |
| Evidence checker | Verifies authored browser or accessibility expectations independently of Jev's choice. |
| CLI + local workbench | Share the runner, budgets, cancellation, saved plans, and evidence reports. |

Jev uses `POST /api/alpha/decisions`; the optional planner uses `POST /api/v1/chat/completions`. Both use standard `fetch`. Direct TypeSafe transport is not implemented. [Provider setup](OPENROUTER_SETUP.md) · [Technical plan](TECHNICAL_PLAN.md)

## Scope and privacy

The alpha supports common Chromium forms, buttons, links, native selects, checkboxes, authenticated fixtures, and async results. Native desktop/mobile apps, CAPTCHA, canvas, complex frames, payment-provider flows, and subjective visual judgments are outside this release. Hosted infrastructure is a later stage.
The alpha supports common Chromium flows plus accessible local iOS and Android apps. Physical phones, hosted devices, native desktop apps, CAPTCHA, arbitrary canvas controls, complex WebViews/frames, payment-provider flows, biometrics, and subjective visual judgments remain outside this release. Hosted infrastructure is a later stage.

API keys stay in the server process. Known fixture/auth values are redacted from observations and reports; input fields are masked in screenshots. Visible application text is sent to OpenRouter, and unrelated page content can remain in reports. Inspect evidence before sharing it. Reports stay under your local `.jev-e2e/runs`; raw trace recording is not implemented. No telemetry is required.
API keys stay in the server process. Known fixture/auth values are redacted from observations and reports; input screens are omitted from screenshots. Visible application text is sent to OpenRouter, and unrelated app or page content can remain in reports. Native recording is explicit, local, and starts after credential-entry steps. If a later screen exposes an input or known secret, the whole report clip is discarded. Inspect evidence before sharing it. Reports and recordings stay under your local `.jev-e2e/runs`; nothing uploads automatically. No telemetry is required.

## Help shape the project

Expand Down
Loading
Loading