diff --git a/packages/web/src/components/bench-charts.tsx b/packages/web/src/components/bench-charts.tsx new file mode 100644 index 0000000..aae23bb --- /dev/null +++ b/packages/web/src/components/bench-charts.tsx @@ -0,0 +1,217 @@ +import { useState } from "react"; +import { TEXT } from "@/components/md"; +import { cn } from "@/lib/utils"; + +/** + * The two benchmark charts, drawn in HTML so they reflow like the rest of the + * page. reins do wears the site's violet and Claude an orange; the pair passes + * the colorblind and contrast checks on both themes. Figures come from + * docs/benchmarks/2026-09-reins-do.md. + */ +const DO = "bg-primary"; +const CLAUDE = "bg-[#eb6834] dark:bg-[#d95926]"; + +function Legend() { + return ( +
+ + + reins do + + + + Claude, step by step + +
+ ); +} + +const PASS: Array<{ set: string; bars: Array<[string, number, string]> }> = [ + { + set: "8 unseen tasks", + bars: [ + [DO, 87.5, "35/40"], + [CLAUDE, 87.5, "21/24"], + ], + }, + { + set: "30 tuned tasks", + bars: [ + [DO, 82.6, "124/150"], + [CLAUDE, 90, "81/90"], + ], + }, +]; + +/** Share of runs passed, per set. */ +export function PassChart() { + return ( +{group.set}
+
- The CLI is the entire interface: reins tabs, reins click,{" "}
- reins screenshot and the rest of the command set. Agents use it because they
- already have a shell: no MCP server to register, no per-agent setup. A skill (
- npx skills add karnstack/reins) teaches agents the loop.
+ The whole interface. Agents already have a shell, so there is no MCP server to register. A
+ skill (npx skills add karnstack/reins) teaches them the commands.
- You never run the daemon yourself. Any CLI command spawns it on demand. It exposes an HTTP{" "}
- /rpc endpoint for the CLI and holds the WebSocket that extensions dial into.
- One daemon serves any number of browsers. reins kill stops it, and logs live in{" "}
- ~/.reins/logs/.
+ Started by any CLI command. It takes HTTP /rpc calls from the CLI and holds the
+ WebSocket each browser's extension connects to. Logs are in ~/.reins/logs/.
- A Manifest V3 extension. Its service worker executes commands against tabs through{" "}
- chrome.debugger (the Chrome DevTools Protocol), and an offscreen document holds
- the persistent WebSocket to the daemon, because MV3 service workers are suspended when idle
- and cannot keep long-lived sockets.
+ Its service worker runs commands through chrome.debugger. An offscreen document
+ holds the WebSocket, because Chrome suspends idle MV3 service workers.
- The extension discovers the daemon by probing a small set of candidate localhost ports, and
- authenticates itself by its chrome-extension://<id> origin, a header the
- browser stamps itself, which web pages and other extensions cannot forge.
+ It finds the daemon by trying a few localhost ports, and proves who it is by its{" "}
+ chrome-extension://<id> origin, which pages cannot fake.
- Install the extension in several Chromium browsers (Chrome, Brave, Edge, Arc, Dia) and each
- connects to the same daemon. reins tabs lists every tab with a browser id. Pass{" "}
- --browser <id> only when more than one browser is connected. reins never
- guesses which browser you meant: it errors and names the ones it can see.
+ Each browser with the extension connects to the same daemon. With more than one connected,
+ pass --browser <id>. reins never guesses; it errors and lists the ones it
+ sees.
- reins snapshot assigns stable refs (e5: button "Submit") to
- interactive elements. Commands act by ref, which survives page repaints better than a
- hand-written selector, and a CSS --selector fallback exists for everything
- else.
+ reins snapshot gives each control a ref (e5: button "Submit").
+ Refs survive repaints better than hand-written selectors. --selector takes CSS
+ when you need it.
| - {h} - | - ))} -
|---|
| - {cell} - | - ))} -
- reins do hands a whole browsing goal to TypeSafe's Jev model. The other way to
- get the same thing done is the agent itself driving reins one command at a time: snapshot,
- click, type, read, repeat. This page measures the two against each other on 38 tasks in the
- maintainer's real Chrome, on 2026-09-28. Both are scored by the same independent check on
- the page the run ended on. Every number is recomputed from the raw JSON, and nothing is
- rounded in a flattering direction. The full report is on GitHub:{" "}
- 2026-09-reins-do.md.
-
- {" "}
- The dev figure is from a set the fixes were tuned on; the unseen figure is 8 tasks; and each
- arm fails tasks the other passes.
+ reins do hands a whole browsing goal to TypeSafe's Jev model. The alternative
+ is your agent driving reins itself: snapshot, click, type, repeat. We ran both on 38 tasks
+ in a real Chrome on 2026-09-28, and scored both with the same check on the page each run
+ ended on.
reins do passed
- 124 of 150 runs (82.6%). Claude step by step passed 81 of 90 (90%); 81 of 87 (93.1%)
- without fx-newtab, a task its rules made impossible.
- reins do passed 35 of 40 runs (87.5%) on its first and only run. Claude step
- by step passed 21 of 24 (87.5%).
- reins do (range 2.4x to 22.6x) and 111x cheaper in
- model spend (range 44x to 535x). Pooled over those tasks' passing runs, the median{" "}
- reins do run took 5.7 s and the median Claude run 24.5 s. The cost counts
- only Jev; the calling agent's own turn is not in it.
- reins do runs that ended{" "}
- done, 17 were wrong, and all 17 carried a self-check under 0.5 ("unsure"); 20
- of the 149 right ones (13%) were marked unsure too.
-
- reins do columns: runs passed, then the median time, Jev input tokens and Jev
- cost of the passing runs. Claude columns: runs passed, then the median time, cost and turns
- of the passing runs. Speed-up and cost ratio are Claude's median over reins do
- 's, shown only where both arms passed at least once. Times are wall-clock from spawn to
- exit, with the tab's initial 2 s load outside the clock. Numbers in brackets point at the
- notes below the tables.
+ About as accurate on tasks it had never seen, 4.6x faster and about 100x cheaper at the
+ median. Neither side wins everywhere: each one fails tasks the other passes.
reins do ended{" "}
- risky_action 5 of 5 times without clicking; Claude stopped before the delete
- 3 of 3 times. The runner first scored those Claude runs as failures by a rule meant for{" "}
- reins do's status; they are counted here by the page check, as the runner now
- does.
- reins do figures.
- left_site, not done: the search form on
- www.debian.org submits to packages.debian.org, and on the holdout's code{" "}
- reins do stopped there as a site change. The page it stopped on was the goal,
- so the check passes. Round 5 treats a leading www. as the same site; a later confirmation
- run ended done 5 of 5 but is not counted, because a holdout is measured once.
- - Per set, the median task speed-up is 4.2x on dev (24 tasks) and 5.4x on holdout3 (6 tasks); - the median cost ratio is 111x and 105x. On the same code as the holdout3 run (f1af855), the - dev set scored 122/150 (81.3%). + +
+ The 30 tuned tasks are the ones reins do was fixed against over five rounds, so
+ its score there flatters it. The 8 unseen tasks were run once, before anyone looked at them.
reins do, dev: 901 Jev calls, 6,256,057 input tokens, $0.263.
- reins do, holdout3: 205 Jev calls, 1,177,962 input tokens, $0.050.
-
- Every done is put back to Jev once as a yes/no question (is every requirement
- in the goal visibly satisfied on this page?), and the answer's probability comes back as{" "}
- doneConfidence. Below 0.5 the output reads{" "}
- done (unsure: self-check 0.43) and the next: line says to verify.
- Recomputed from every done run of the dev and holdout3 runs:
+ Median run on each of the 30 tasks both passed. reins do was faster on all 30,
+ from 2.4x to 22.6x.
- Read from each failed run's final reply. Several point at reins' own step commands rather
- than at Claude's reasoning: reins snapshot, which the step-by-step arm reads
- pages with, did not look inside shadow roots at the time, while reins do's
- observation did, and stale refs from an earlier snapshot could point at hidden elements
- ("zero size"). Both are fixed in 0.6.1. The fx-wizard page also opened with its modal
- already showing, for both arms, because of a CSS bug in the fixture (fixed since).
+
+ Jev's spend covers Jev only. Your agent still spends a turn sending the command and checking + the result.
-reins press / was rejected as an unknown
- key. Run 1 got there in 34 turns.
- - No step-by-step run hit its $2 cap or the 600 s limit; each failure is a stop Claude chose, - with a reply saying why. -
- -reins do: false on the start page and on near-miss
- states (searched but not sorted, the wrong edition, a date picked but not confirmed), true
- on a goal state reached another way. They are in tasks.mjs.
- reins do '<goal>' --fill … --timeout 60 --tab <id> --json, 90 s
- for four slow sites, at most 30 steps. 5 runs per task.
- claude -p with the same goal and fill values,
- allowed only the reins step commands (snapshot, click, type, fill, select, press, hover,
- scroll, wait, text, screenshot), one per Bash call, --restricted with an
- allowlist, no reins do, no navigation by URL, no other tabs, stop before
- anything irreversible, $2 cap. 3 runs per task, because each costs about a hundred times
- more. Its cost is Claude Code's reported total_cost_usd.
- reins do run on it; its one run is the unseen number here, measured before
- round 5.
- reins do times include that.
- reins do.
-
- reins do was tuned in five rounds on the dev set. Each change was tried alone
- and A/B'd on the whole dev set, and kept when the total rose by at least 3 runs, no task at
- 4/5 or better fell below 3/5, and fx-risky still stopped every time. Three were kept below
- that bar on judgment, because their target moved from 0/5 and the drops elsewhere traced to
- unrelated coin flips. Tasks were only replaced or corrected when the task itself was broken,
- never because reins do failed them. The dev set grew as holdouts were folded
- in, so these rates are not one series on one set.
+ reins do ends done when Jev thinks the goal is met, then asks Jev
+ once more to confirm. 17 of 166 done runs were wrong, and all 17 came back
+ marked unsure. So an unsure done is the one to look at before you
+ report success. It is not proof of failure: 20 correct runs were marked unsure too.
- Round 5's total rose from 122 to 124, which the traces put down to noise: neither round-5 - mechanism fired on a dev task. -
-Tried and reverted, each on its own A/B:
-
-
+ reins do never passed these:
done, self-check 0.32 to 0.49.
- done, self-check 0.12 to 0.14.
- done, self-check
- 0.38 to 0.45.
- blocked stop.
- stuck (×3,
- self-check 0.04) or said done there (×1, 0.08). The pass used the list's own
- "Filter by name" field.
- needs_text at step 0: xe labels its
- amount input "Receiving amount", the same as the converted output, so no field matches the
- amount fill. A re-run with the currency fills the task first lacked was still 0/5.
- Unsolved.
+ xe.com. The site labels its input and its output the same, so the amount never lands.
Claude driving reins step by step never passed these:
done was
- marked unsure, but that is 166 runs on 38 tasks, and the 17 wrong ones come from 5 tasks.
- Verify the page before reporting success.
- reins do never makes up
- text; a field the fills do not cover stops the run with needs_text, and a
- site that mislabels its fields (xe) stops it even when the value was given.
- window.open) can still bring Chrome to the
- front. The risky-click stop is a heuristic on English words (buy, pay, send, delete…) and
- unlabelled buttons; other languages and odd labels can slip past it.
- reins do, n=3 for Claude. A live site can change tomorrow. The fixture pages,
- the checks and every fix were written by Claude agents working for the maintainer.
- left_site (note 4).
- reins snapshot does not read
- shadow roots). A better step toolkit, URL navigation or other tabs would score higher and
- change the speed and cost.
- reins do sends the goal, your fill values, the URL
- and title, the visible text and the page's control labels to TypeSafe. The full list is
- under reins do and TypeSafe on the security page.
- Step by step sends page content to Anthropic. Either way it is opt-in.
+ A modal wizard. Clicks on a custom radio card reported "zero size". This and Open Library
+ were mostly gaps in reins snapshot, fixed in 0.6.1.
+ It is a small benchmark: 38 tasks, one machine, one day, from India to Jev's US servers.
+ Live sites change. Jev input costs $0.042 per million tokens,{" "}
+ TypeSafe's published price. reins do sends page text
+ to TypeSafe; here is exactly what.
+
- Other flags: --tier fixture|live, --tasks id1,id2,{" "}
- --claude-budget-usd, --out-dir and --dry (a self-test
- with no browser and no spend). It needs a TypeSafe key (reins key set typesafe)
- and claude on PATH. It costs money: about $0.32 of Jev for the 190 runs here
- and $21.83 of Claude for 114. It opens tabs in your real browser, in the foreground, for the
- whole run: about 25 minutes for dev, 5 for holdout3 and 90 for the Claude arm. Holdout3 is
- no longer unseen. The script is bench-do.mjs. Raw run files aren't kept
- in the repo; these commands regenerate them.
+ Needs a TypeSafe key (reins key set typesafe) and claude on your
+ PATH. It opens tabs in your real browser for about two hours and spends about $22, nearly
+ all of it on Claude. --dry checks the setup with no browser and no spend. The
+ script is bench-do.mjs.
- The first version (2026-09-27) ran 4 live tasks (flights, wikipedia, github, cookies) 5
- times per arm. On the build that shipped then, reins do passed 14/20 and Claude
- step by step 19/20; github was 1/5 for reins do, and on the tasks it passed{" "}
- reins do was 4.6x to 7.9x faster. It had no holdout and no fixture tier, which
- is why this version exists.
+ Per-task numbers, how each task is checked, every fix and what it moved, the ones we
+ reverted, and the traces behind each failure: 2026-09-reins-do.md on
+ GitHub.
- The CLI is the whole interface: agents shell out to it, and so can you. The commands that
- act on a page or a tab share three flags: --tab <id> (the active tab by
- default), --browser <id> (only needed when several browsers are
- connected, and the ids come from reins tabs) and --json for raw
- results. The management commands (status, doctor,{" "}
- kill, help) take none of them.
+ Page and tab commands share three flags: --tab <id> (default: the active
+ tab), --browser <id> (only when several browsers are connected) and{" "}
+ --json. Ids come from reins tabs.
- agent-browser and{" "} - dev3000 (Vercel Labs) and{" "} - playwright-mcp (Microsoft) are all - browser tooling for coding agents, and they all start from the same place: by default each - one launches and manages a browser for the agent. reins starts from the other end. It hands - the agent the browser you already have open. + The other browser tools for coding agents launch a browser of their own by default. reins + uses the one you already have open, so you are already logged in everywhere.
-- agent-browser is a fast, general automation CLI that owns its browser. reins puts an - extension inside the browsers you already run, so every session is authenticated by - definition, nothing new launches, and no debug port is ever exposed. The daemon only accepts - the extension's unforgeable origin on 127.0.0.1. + agent-browser (Vercel Labs) is a + fast automation CLI that launches its own Chrome for Testing. It can reuse a profile's + logins or attach to a running Chrome if you set that up. It adds HAR recording, request + mocking and web vitals.
- If you need headless fleets, request mocking or CI runs, agent-browser is the better fit. If - the task is "act as me, in my browser", that is reins. + Pick it for headless runs, CI and request mocking. Pick reins to act as you, in your + browser.
- dev3000 solves a different problem: it wraps your dev server, launches a monitored browser, - and merges server logs, console, network and screenshots into one timeline an AI can debug - from. That is dev-loop observability, not general browser control. + dev3000 (Vercel Labs) wraps your dev + server, launches a monitored browser, and merges server logs, console, network and + screenshots into one timeline for an AI to debug from.
- They compose. dev3000 watches the app you are building, and reins drives the rest of your - browser: dashboards, docs, the third-party service you are integrating. + Different job, and they combine well: dev3000 watches the app you are building, reins drives + everything else in your browser.
- The closest comparison: its extension mode can also drive existing tabs in your real browser - (Chrome and Edge only). The defaults differ. playwright-mcp launches a Playwright-managed - browser with its own persistent profile, and everything flows through an MCP server you - register in each client. reins is a plain CLI, so any agent with a shell drives your - everyday browsers with no per-agent setup, and one daemon serves them all at once. + playwright-mcp (Microsoft) is an + MCP server that launches a Playwright browser with its own profile. It has an opt-in + extension mode that drives your real Chrome or Edge tabs, the closest thing to reins. You + register it in each agent. +
++ Pick it for Firefox and WebKit, device emulation, or clean isolated sessions. Pick reins for + a plain CLI any agent can use with no setup, across all your Chromium browsers at once.
+ +- Pick playwright-mcp for cross-engine coverage (Firefox, WebKit), device emulation, or - clean-room isolated sessions. Pick reins when the point is acting as you, in the browser you - already work in. + An extension inside your browser instead of a debug port, per-site permission tiers, a + redacted audit trail, and raw CDP when the built-in commands are not enough.
> ); diff --git a/packages/web/src/routes/docs/faq.tsx b/packages/web/src/routes/docs/faq.tsx index 99d9ca3..d0e1d98 100644 --- a/packages/web/src/routes/docs/faq.tsx +++ b/packages/web/src/routes/docs/faq.tsx @@ -7,61 +7,60 @@ const FAQS = [ id: "remote", question: "Is anything ever sent to a remote server?", answer: - "Not by default. The extension talks to exactly one thing: the reins daemon on 127.0.0.1 on your machine. There is no analytics, no telemetry, and no remote code. The one exception is opt-in: once you save a TypeSafe API key, reins do sends page state (the goal, the tab's URL and title, visible text, interactive element labels and values, recent actions) from the daemon to TypeSafe. The security page lists exactly what is sent. Your browser still reaches the internet the way it always did, because reins drives the browser you already use rather than replacing it.", + "Not by default. The extension talks only to the reins daemon on 127.0.0.1. No analytics, no telemetry, no remote code. The one opt-in exception is reins do: once you save a TypeSafe key, it sends page state to TypeSafe while a run is working. The security page lists exactly what.", }, { id: "browsers", question: "Which browsers work?", answer: - "Any Chromium browser that supports Manifest V3 extensions. Chrome, Brave, Edge, Arc and Dia are all known to work. Install the extension in each browser you want agents to reach; one daemon serves them all.", + "Any Chromium browser with Manifest V3 extensions: Chrome, Brave, Edge, Arc, Dia. Install the extension in each one; one daemon serves them all.", }, { id: "agents", question: "Which agents work?", answer: - "Anything with a shell: Claude Code, Cursor, Codex, GitHub Copilot, Gemini CLI, and plain scripts. Agents with skill support learn the commands via npx skills add karnstack/reins; everything else can read reins help.", + "Anything with a shell: Claude Code, Cursor, Codex, Copilot, Gemini CLI, plain scripts. Teach it with npx skills add karnstack/reins, or point it at reins help.", }, { id: "banner", question: 'Why does Chrome show an "is being debugged" banner?', answer: - "reins executes commands through chrome.debugger, the same Chrome DevTools Protocol that powers DevTools. Chrome shows its native banner whenever a debugger is attached. That is deliberate transparency: you always know when an agent is acting on a tab.", + "reins drives tabs through chrome.debugger, the same protocol DevTools uses, and Chrome shows that banner whenever a debugger is attached. It is how you know an agent is acting.", }, { id: "daemon", question: "Do I need to run or configure the daemon?", answer: - "No. Any reins command starts the daemon on demand, and the extension finds it on its own through localhost port discovery. reins kill stops it; reins status shows what is connected.", + "No. Any reins command starts it, and the extension finds it on its own. reins status shows what is connected; reins kill stops it.", }, { id: "update", question: "How do I update reins?", answer: - "Run npm i -g @karnstack/reins@latest. The next tool command (reins tabs, say) notices the running daemon is older than the CLI and restarts it on the new version; reins restart does it right away. reins status and reins doctor only report the mismatch. The Chrome Web Store extension updates itself.", + "Run npm i -g @karnstack/reins@latest. The next command restarts the daemon on the new version, or run reins restart. The Web Store extension updates itself.", }, { id: "mcp", question: "How is this different from an MCP browser server?", answer: - "There is nothing to register per agent. reins is a plain CLI, so any tool that can run shell commands can drive the browser. And it drives your real, logged-in profile rather than a separate automation browser.", + "Nothing to register per agent. reins is a plain CLI, so anything that runs shell commands can use it, and it drives your real logged-in browser rather than a separate one.", }, { id: "stop", question: "How do I stop an agent while it is running?", answer: - "Click the reins toolbar icon and press Disconnect. The connection is cut at once. reins kill stops the daemon entirely.", + "Click the reins toolbar icon and press Disconnect. reins kill stops the daemon entirely.", }, { id: "store", question: "Can I install the extension without the Chrome Web Store?", answer: - "Yes. reins extension stages the bundled extension for Chrome's Load unpacked, with no reins allow step. The npm package carries a full copy, so it works with no store access at all. The docs page Install without the store has the walkthrough.", + "Yes. reins extension stages the copy bundled in the npm package for Chrome's Load unpacked. See Install without the store.", }, { id: "dev-builds", question: "Does reins work with unpacked dev builds of the extension?", - answer: - "Yes. Load the unpacked extension, then allow its ID once with reins allow- Quick answers about how reins works. Anything missing? Ask on{" "} - GitHub. + Missing something? Ask on GitHub.
{FAQS.map((faq) => (- reins gives coding agents full control of your actual, logged-in Chromium browser, through a - CLI and a Manifest V3 extension. Claude Code, Cursor, Codex, Copilot, anything with a shell. - This page takes you from nothing to an agent driving a tab. -
+Four steps from nothing to your agent driving your logged-in browser.
- This installs the reins command and the daemon it manages. You never run the
- daemon yourself: any command starts it on demand, it binds 127.0.0.1, and{" "}
- reins kill stops it. There is nothing to configure and nothing to keep running.
+ This installs reins. Its daemon starts on its own when you run any command, and{" "}
+ reins kill stops it. Nothing to configure.
- The skill teaches agents the command set and the loop below. This is the step people skip,
- and skipping it is why an agent with reins installed still says it cannot open a browser.
- Agents without skill support can run reins help; the CLI is self-describing.
+ This teaches your agent the commands. Skip it and your agent will still say it cannot open a
+ browser. No skill support? Point it at reins help.
- Install the reins extension from the Chrome Web Store in - every Chromium browser you want agents to reach. Chrome, Brave, Edge, Arc and Dia all work. - The extension finds the daemon on its own through localhost port discovery, and the toolbar - popover turns green when it is connected. -
-
- Prefer to skip the store? reins extension stages the bundled copy for Chrome's
- Load unpacked. The walkthrough is on Install without the store.
+ Add the reins extension to each Chromium browser you want
+ agents to reach: Chrome, Brave, Edge, Arc, Dia. Its icon turns green once it finds the
+ daemon. No store access? See Install without the store.
Working from a dev build instead? Load the unpacked extension and allow its ID once:
-Every page interaction is the same three beats: look, act, check.
+Look, act, check.
- The commands that act on a page or a tab share three flags: --tab <id>{" "}
- (the active tab by default), --browser <id> (only needed when several
- browsers are connected) and --json for raw output.
+ Page commands act on the active tab unless you pass --tab <id>. Add{" "}
+ --json for raw output.
- Every site resolves to one of three tiers, in order of power: deny <{" "}
- read < full. The extension checks the tier before it runs any
- command against a tab. The check lives in the extension itself, the one place a process on
- your machine cannot reach around, so even a misbehaving agent (or a compromised daemon)
- cannot skip it.
+ Every site gets one of three tiers: deny, read or{" "}
+ full. The extension checks it before any command touches a tab, so neither the
+ agent nor the daemon can skip it.
- The shipped default is full everywhere, so a fresh install behaves exactly as
- before. The policy model is opt-in hardening: tighten the sites you care about, or flip the
+ The default is full everywhere. Tighten the sites you care about, or flip the
default and grant sites back one by one.
- Navigation is checked on both ends: reins nav and reins open need{" "}
- full on the destination host as well as the current one, so a read-only page
- cannot be steered somewhere permissive.
+ reins nav and reins open need full on both the
+ current and the destination site.
- Grants happen only in the extension popup: click the reins icon, and the Site permissions - section offers a tier control for the current tab, the rules list, and the default. That is - deliberate. The popup is a user gesture, and an agent in your shell cannot perform one. From - the CLI you can inspect the policy and tighten it, never loosen it: + Grant access in the extension popup, under Site permissions. An agent in your shell cannot + click it, which is the point. The CLI can only look and tighten:
- A blocked command fails with a policy_denied error. It names the host, its
- current tier, and what to do about it:
+ A blocked command fails with policy_denied and says what to do:
- The CLI prints the message and exits nonzero, so agents relay the instruction instead of - retrying. -
+It exits nonzero, so agents pass the message on instead of retrying.
- A tool that drives your logged-in browser has to be careful with it. reins keeps the attack - surface small by having no cloud half at all: the pieces only ever talk to each other, on + reins has no cloud half. The CLI, the daemon and the extension only talk to each other, on your machine.
-127.0.0.1. Neither the daemon nor the extension is reachable
- from the network.
+ Everything binds 127.0.0.1. Nothing is reachable from the network.
/rpc and the other daemon endpoints validate the Host header, so
- web pages cannot reach the daemon even through rebound DNS.
+ The daemon checks the Host header, so a web page cannot reach it through DNS
+ rebinding.
chrome-extension://<id> origins. The browser stamps that header itself,
- so pages and other extensions cannot forge it. Dev builds are added explicitly with{" "}
- reins allow <id>.
+ It accepts only allowlisted chrome-extension://<id> origins. Chrome
+ sets that header itself, so pages and other extensions cannot fake it.
deny, read or full.
- The extension enforces it before any command touches a tab. The check runs inside the
- extension, so nothing that speaks the protocol can skip or loosen it. That includes the
- CLI, the daemon, and any other local client.
- reins policy) can view and tighten the
- policy, never loosen it.
- full everywhere, which is today's behavior, so
- tightening is opt-in. deny also redacts the site's tabs from{" "}
- reins tabs.
-
+ Every site is deny, read or full, checked inside the
+ extension before any command runs. Only a click in the popup can grant more; the CLI can
+ only tighten. The default is full everywhere, so this is opt-in.
+
- The tiers contain the agent you invited in. They are not a defense against other software on - your machine. Anything already running as your OS user sits inside the trust boundary: it - could talk to the daemon or rewrite the policy store directly, and no browser automation - tool's permission model survives local malware. The honest write-up covers what the tiers - protect against, what they do not, prompt injection, and a hardening checklist. It is the{" "} - - threat model (SECURITY.md) - - . + The tiers contain the agent you invited in. They do not stop other software running as you, + which could talk to the daemon or edit the policy directly. The{" "} + threat model{" "} + covers this, prompt injection, and a hardening checklist.
~/.reins/logs/audit-YYYY-MM-DD.jsonl. reins audit renders the
- trail, and --denied shows only what policy blocked.
- eval code and CDP payloads. The trail never stores what the agent typed, only
- that it typed.
-
+ Every command, and every one policy blocks, adds a line to{" "}
+ ~/.reins/logs/audit-YYYY-MM-DD.jsonl. Typed text, fill values, eval code and
+ CDP payloads are redacted, so the log shows that the agent typed, never what. Files are
+ deleted after 30 days. reins audit shows the trail; --denied shows
+ only blocks.
+
- reins do is off until you save a TypeSafe API key (
- reins key set typesafe, or the Jev section of the extension popup). Once a key
- is saved, and only while a reins do run is working, the daemon (not the
- extension) sends this to api.typesafe.ai, under your own TypeSafe account:
+ reins do is off until you save a TypeSafe key (
+ reins key set typesafe). While a run is working, the daemon sends this to{" "}
+ api.typesafe.ai, under your account:
--fill names and values
+ your goal and your --fill values
- Never sent: password, file and hidden inputs. The extension itself makes no remote requests. -
-
- Before every click or keystroke, reins re-checks the chosen element (still present, visible,
- not covered, not moving) and refuses to act when the check fails. The page still controls
- its own DOM, so what sits under a chosen element can change between the check and the
- action; keep --confirm for anything you would not click blind.
+ Password, file and hidden inputs are never sent. The key lives in{" "}
+ ~/.reins/credentials.json (mode 0600) and is never printed, logged or sent to a
+ page.
- reins do hands page state to TypeSafe's Jev model, which answers typed
- multiple-choice questions: which operation, and which observed element. Jev can only choose
- among elements reins actually read from the page; its output never becomes a selector,
- coordinate or code. Page text can still try to steer it (prompt injection), so:
+ Jev can only pick from elements reins read off the page. To limit what a page can talk it
+ into:
--confirm for that label or the goal
- names it word for word. Unlabeled buttons stop too. This is a heuristic, not a guarantee:
- other languages and odd labels can slip past.
+ Clicks labelled with words like buy, pay, send or delete stop the run unless you pass{" "}
+ --confirm or the goal names that label. Unlabelled buttons stop too. It
+ matches English words, so it is not a guarantee.
left_site), and site permissions
- still apply on every step (full required).
+ Leaving the site stops the run (left_site). Site permissions apply on every
+ step.
--timeout or a daemon restart stop the run before its
- next action.
+ Ctrl-C, --timeout or a daemon restart stops it before the next action.
~/.reins/credentials.json (0600). The key is never returned
- by any command, never logged, and never sent to a page.
+ reins re-checks each element right before acting, but a page can still swap what sits
+ under it. Use --confirm for anything you would not click blind.
- What a run costs in time and tokens, measured against an agent driving the step commands - itself, is on the Benchmarks page. -
-reins do with a TypeSafe key (see above).
- chrome.storage on
- your device.
-
- The full policy is at reins.tech/privacy. The code is MIT-licensed
- and auditable at github.com/karnstack/reins
- .
+ No analytics, no telemetry, no remote code. Page content goes only to your local daemon,
+ unless you opt in to reins do. The extension stores its settings and your site
+ policy in chrome.storage. Full policy:{" "}
+ reins.tech/privacy. Source:{" "}
+ github.com/karnstack/reins.
reins snapshot lists what you can click, with short refs to act on. CSS
+ selectors work too.
reins eval for JavaScript, reins cdp for any raw DevTools
+ Protocol call.
reins snapshot lists the interactive elements with
- stable refs, and commands act by ref. A CSS --selector is there when you need
- it.
- reins eval runs JavaScript in the page.{" "}
- reins cdp sends a raw Chrome DevTools Protocol command when the curated set
- is not enough.
- ~/.reins/logs, with the values redacted.
+ reins do hands a whole task to Jev, TypeSafe's action model, instead of going
+ click by click.
- Every page interaction is the same three beats: look, act, check. The commands that act on a
- page or a tab share three flags: --tab <id> (the active tab by default),{" "}
- --browser <id> (only when more than one browser is connected) and{" "}
- --json for raw output.
-
Look, act, check.
- Three pieces with one narrow contract between them, and all three run on your machine. The - daemon ships inside the CLI and starts on demand, so there is nothing to keep running and - nothing to register per agent. + Three pieces, all on your machine. The daemon ships inside the CLI and starts on its own. + Chrome shows its "is being debugged" banner while reins is attached to a tab.
{PATH}
-
- The extension finds the daemon by probing a small set of localhost ports, and authenticates
- by its chrome-extension://<id> origin, a header the browser stamps itself
- and a page cannot forge. Chrome shows its native debugging banner the whole time it is
- attached.
-
- Every site your agent touches resolves to one of three tiers. The check lives in the - extension, the one place no process on your machine can reach around, so a misbehaving agent - cannot skip it. + Every site gets one of three tiers, checked inside the extension where an agent cannot skip + it.
- Granting more access takes a click in the extension popup. That is a user gesture, and an
- agent in your shell cannot perform one. From the CLI, reins policy can inspect
- the policy and tighten it. It can never loosen it.
+ Only a click in the extension popup grants more access. From the shell,{" "}
+ reins policy can tighten, never loosen.