You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
OriginWeave has extensive exact-coverage and hostile-input contracts, plus many active real-browser and protocol slices. A buyer still cannot compare releases against one versioned benchmark that proves the complete product is useful, repeatable, safe, evidence-complete, resource-bounded, and operable across supported platforms.
Passing unit tests, one successful browser run, or a green model review does not prove commercial readiness. Completion requires a release gate that answers:
Does the same signed OriginWeave release repeatedly complete realistic delegated web work without unauthorized actions, stale authority, prompt-injection success, missing provenance, process/profile leakage, or unbounded cost?
This issue owns the benchmark corpus, runner, metrics, statistical acceptance rules, evidence package, and release decision contract. It does not replace source tests or protected review.
Benchmark program structure
Create a versioned benchmark registry with five independently reportable suites:
allowed origins, capabilities, actions, secrets, files, and external effects;
expected structured result;
required post-condition;
required evidence/provenance;
resource and time budgets;
cleanup and reversibility requirements;
deterministic oracle or human-review rubric;
data classification and retention; and
known sources of nondeterminism.
External benchmark datasets are adapters/evidence sources. OriginWeave must not silently redistribute data or tasks outside their licenses.
Controlled deterministic suite
Use local fixture applications to prove exact behavior without third-party drift.
Required task families:
semantic locate/type/click/submit;
same-document and navigation post-conditions;
structured extraction from DOM/accessibility, JSON-LD, table, and bounded network response;
iframe and shadow-DOM cases;
file download and approved upload;
approval-required reversible action;
secret-handle form fill with no model disclosure;
redirect and origin transition;
dynamic mutation and stale-node invalidation;
session checkpoint/cancel/resume;
browser crash and process/profile cleanup;
WARC/PROV capture and offline replay; and
Manifest V3 and native-host isolation where the release declares support.
Release gate: 100 successful runs out of 100 per required deterministic case, zero unauthorized side effects, exact post-condition evidence, and 100% provenance completeness.
Web compatibility suite
Create a versioned, legally reviewable set of public sites or self-hosted replicas representing realistic structures:
conventional forms and navigation;
virtualized lists;
complex tables;
SPAs;
accessibility-rich and accessibility-poor pages;
canvas/visual fallback cases;
downloads;
authentication handoff without production credentials;
localization and non-Latin scripts; and
declared MV3 extension interactions.
Use Mind2Web/WebArena/VisualWebArena-style task diversity as research guidance, not as an unsupported claim of direct score equivalence.
Initial release target for the approved stable compatibility set:
at least 95% task success over 10 independent runs per case;
no policy or security violation;
result/evidence variance reported rather than hidden;
every failure classified as product defect, site drift, external outage, unsupported capability, or benchmark defect; and
unsupported/skipped cases excluded from the claimed support matrix rather than counted as success.
Security adversarial suite
Include benign and hostile variants that differ only in untrusted content so the security effect is measurable.
Required threats:
direct and indirect prompt injection in visible, hidden, accessibility, metadata, network, extension, WebMCP, document, and downloaded content;
attempts to mint tools, capabilities, approval, policy, new goals, or memory;
secret/cookie/token exfiltration attempts;
stale node and document-epoch replay;
cross-session, cross-context, cross-frame, cross-origin, cross-tenant, and confused-deputy handle reuse;
DNS rebinding, special-purpose IPs, alternate numeric host forms, redirect downgrade/cycle/private target, proxy/PAC bypass, and peer/TLS substitution;
stale checkpoint, incompatible protocol/software/policy, and task replay;
update failure and rollback;
backup/restore and regional failover; and
late or duplicated event delivery.
Required outcomes distinguish confirmed success, confirmed no side effect, confirmed side effect, unknown outcome, reconciliation required, cancelled, expired, and failed. Generic exception suppression or fabricated recovery evidence is prohibited.
Release gate:
100% cleanup of task-owned profiles and declared process sets in controlled cases;
no duplicate externally visible side effect under retries;
deterministic reconciliation state for every uncertain outcome;
successful restore/recovery for every declared supported deployment profile; and
no lost audit/provenance record for an accepted action.
Do not infer operating-control effectiveness from the existence of policy documents. The benchmark must exercise the configured control and preserve its evidence.
Use 3NF for cases, versions, environments, metrics, thresholds, and decisions. Raw tool/model/browser logs are evidence artifacts, not a normalized source of release truth.
Release decision contract
Create a deterministic release_decision that can be:
Every benchmark runner/parser/evaluator production function, line, region, and branch must have exact 100% coverage and complete public documentation.
Add property tests and fuzzing for result ingestion, threshold evaluation, corrupted artifacts, hostile metadata, integer/decimal bounds, and concurrency.
CI must fail if required cases are removed, weakened, skipped, renamed without migration, or moved outside the release profile.
A benchmark cannot rewrite product source, thresholds, or expected output during execution.
Diagnostic workflows must not turn test failure into success.
Synthetic fixtures may be used for deterministic tests but must never enter production as customer data.
Documentation and research traceability
Add an ADR for benchmark ownership, corpus/licensing, metrics, statistics, release thresholds, human review, model routing, anti-gaming, migration, and reversal.
Add APA 7th references and a claim-to-test traceability matrix.
Publish a benchmark card for every suite and a release evidence report for every stable release.
Update PRD, TRD, ARCHITECTURE, TEST_STRATEGY, OPERABILITY, RELEASE_AND_ROLLBACK, doctoring, roadmap, README, AGENTS.md, CLAUDE.md, and CHANGELOG only as the executable gate becomes protected-main truth.
claiming benchmark scores from a different implementation or model route;
optimizing solely for one public benchmark;
redistributing restricted datasets;
using one-off success as reliability evidence;
replacing exact security invariants with an average utility score;
declaring unsupported platforms or skipped GPU/browser tests as passing.
Commercial proof
A buyer receives one signed report bound to the exact release artifact showing repeated realistic task success, zero observed unauthorized actions across the declared adversarial trial count, complete provenance, bounded resources/cost, tested recovery, enterprise control operation, known limitations, and reproducible evidence.
Research and standards references — APA 7th
Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H., & Su, Y. (2023). Mind2Web: Towards a generalist agent for the web. arXiv. https://arxiv.org/abs/2306.06070
Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Ou, T., Bisk, Y., Fried, D., Alon, U., & Neubig, G. (2023). WebArena: A realistic web environment for building autonomous agents. arXiv. https://arxiv.org/abs/2307.13854
Koh, J. Y., Lo, R., Jang, L., Duvvur, V., Lim, M. C., Huang, P.-Y., Neubig, G., Zhou, S., Salakhutdinov, R., & Fried, D. (2024). VisualWebArena: Evaluating multimodal agents on realistic visual web tasks. arXiv. https://arxiv.org/abs/2401.13649
Evtimov, I., Zharmagambetov, A., Grattafiori, A., Guo, C., & Chaudhuri, K. (2025). WASP: Benchmarking web agent security against prompt injection attacks. arXiv. https://arxiv.org/abs/2504.18575
National Institute of Standards and Technology. (2024). Artificial intelligence risk management framework: Generative artificial intelligence profile (NIST AI 600-1). https://doi.org/10.6028/NIST.AI.600-1
Buyer-visible gap
OriginWeave has extensive exact-coverage and hostile-input contracts, plus many active real-browser and protocol slices. A buyer still cannot compare releases against one versioned benchmark that proves the complete product is useful, repeatable, safe, evidence-complete, resource-bounded, and operable across supported platforms.
Passing unit tests, one successful browser run, or a green model review does not prove commercial readiness. Completion requires a release gate that answers:
This issue owns the benchmark corpus, runner, metrics, statistical acceptance rules, evidence package, and release decision contract. It does not replace source tests or protected review.
Benchmark program structure
Create a versioned benchmark registry with five independently reportable suites:
Every benchmark case must declare:
External benchmark datasets are adapters/evidence sources. OriginWeave must not silently redistribute data or tasks outside their licenses.
Controlled deterministic suite
Use local fixture applications to prove exact behavior without third-party drift.
Required task families:
Release gate: 100 successful runs out of 100 per required deterministic case, zero unauthorized side effects, exact post-condition evidence, and 100% provenance completeness.
Web compatibility suite
Create a versioned, legally reviewable set of public sites or self-hosted replicas representing realistic structures:
Use Mind2Web/WebArena/VisualWebArena-style task diversity as research guidance, not as an unsupported claim of direct score equivalence.
Initial release target for the approved stable compatibility set:
Security adversarial suite
Include benign and hostile variants that differ only in untrusted content so the security effect is measurable.
Required threats:
Release gate:
Reliability and recovery suite
Exercise failures at every boundary:
Required outcomes distinguish confirmed success, confirmed no side effect, confirmed side effect, unknown outcome, reconciliation required, cancelled, expired, and failed. Generic exception suppression or fabricated recovery evidence is prohibited.
Release gate:
Enterprise operability suite
Compose issue #202 and validate:
Do not infer operating-control effectiveness from the existence of policy documents. The benchmark must exercise the configured control and preserve its evidence.
Metrics and statistical reporting
At minimum publish:
Reproducibility and execution
Data model
Every database object must contain at least two words and use
snake_case. Suggested objects:Use 3NF for cases, versions, environments, metrics, thresholds, and decisions. Raw tool/model/browser logs are evidence artifacts, not a normalized source of release truth.
Release decision contract
Create a deterministic
release_decisionthat can be:acceptedrequires every mandatory suite and threshold for the declared support profile.accepted_with_declared_limitationsmust remove unsupported features/platforms from the claim and list buyer-visible consequences.rejectedidentifies exact failing thresholds and next remediation actions.inconclusiveis required for missing evidence, infrastructure failure, external outage, insufficient trials, or unreviewed benchmark drift.TDD, coverage, and anti-gaming rules
Documentation and research traceability
Dependencies and non-goals
Dependencies:
Non-goals:
Commercial proof
A buyer receives one signed report bound to the exact release artifact showing repeated realistic task success, zero observed unauthorized actions across the declared adversarial trial count, complete provenance, bounded resources/cost, tested recovery, enterprise control operation, known limitations, and reproducible evidence.
Research and standards references — APA 7th
Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H., & Su, Y. (2023). Mind2Web: Towards a generalist agent for the web. arXiv. https://arxiv.org/abs/2306.06070
Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Ou, T., Bisk, Y., Fried, D., Alon, U., & Neubig, G. (2023). WebArena: A realistic web environment for building autonomous agents. arXiv. https://arxiv.org/abs/2307.13854
Koh, J. Y., Lo, R., Jang, L., Duvvur, V., Lim, M. C., Huang, P.-Y., Neubig, G., Zhou, S., Salakhutdinov, R., & Fried, D. (2024). VisualWebArena: Evaluating multimodal agents on realistic visual web tasks. arXiv. https://arxiv.org/abs/2401.13649
Evtimov, I., Zharmagambetov, A., Grattafiori, A., Guo, C., & Chaudhuri, K. (2025). WASP: Benchmarking web agent security against prompt injection attacks. arXiv. https://arxiv.org/abs/2504.18575
National Institute of Standards and Technology. (2024). Artificial intelligence risk management framework: Generative artificial intelligence profile (NIST AI 600-1). https://doi.org/10.6028/NIST.AI.600-1