A vendor-neutral, enterprise-grade quality and safety governance framework covering the full lifecycle of AI Agents.
Table of Contents
Runtime evaluation of an Agent's output quality and detection of sophisticated logical attacks fall into an "impossible triangle" — business real-time requirements, judgment complexity, and long context — which is a fundamental constraint that cannot be solved through engineering alone. As a result, the industry consensus is to shift checks left into CI/CD. But once shifted left, what exactly should be checked, what criteria should be checked against, and what standard should be used to grant release? The high-frequency, hot-iteration nature unique to Agents causes the same class of problems to recur every time a Prompt, Skill, or Tool Schema is modified. SanityOps is designed precisely to answer these three questions.
The output quality and logical vulnerabilities of an Agent are closely tied to the rigor of each logical artifact (Prompt, Skill, Tool Schema, Permission), manifesting as a wide variety of defects as well as permission requests disproportionate to the assigned responsibilities.
In our inspection of 500+ logical artifacts from 50+ agents, over 90% contained defects — including LLM-generated ones.
Therefore, we start from the inspection and remediation of defects and permissions in logical artifacts, and extend to evaluating Agent output quality and detecting logical risks.
Through near-real-time synchronization with Git and shadow sandboxes that mirror production, SanityOps weaves defect inspection, permission checks, explicit/implicit risk audits, and quality evaluation into a continuous governance loop that runs in lockstep with rapid iteration.
SanityOps has two parts:
-
Framework (the specification) — the complete methodology across Inspect, Risk, Quality, Relevance, and Core, published as open specification documents (CC BY-SA 4.0).
-
Platform (the tooling) — the set of tools that implement the specifications above:
-
🟢 Open source — the standalone Inspector CLI (Apache 2.0, supporting L1/L2/L3).
-
🔵 Commercial — the full feature suite (Defect Inspector, Risk Scanner, Quality Evaluator), offered as SaaS / self-hosted.
-
- Like a "compiler" for AI logic, it establishes a complete inspection framework targeting defects in logical artifacts (→ Inspect).
- Establishes mapping relationships among defects, permissions, quality, and risk, aiding problem localization and remediation (→ Relevance / Inspect Permission).
- Establishes methods for generating test cases and mock data based on logical artifacts, making Red Team exercises more targeted and effective (→ Risk Implicit / Quality).
- Establishes dedicated quality evaluation dimensions and metric systems tailored to the distinct service forms of RAG-Agents and Tool-Agents respectively (→ Quality RAG-Agent / Quality Tool-Agent).
- Performs dynamic verification in a shadow environment equivalent to production, more closely reflecting real operating conditions than traditional sandboxes (→ Risk Implicit).
- A ratchet mechanism synchronized with logical artifact versions provides a simple yet effective production-line admission method for hot iterations (→ Quality / Core: Baseline & Gate).
SanityOps Framework
│
├─ Inspect (Defect Inspection)
│ ├─ Inspect Prompt ← System prompt defect inspection
│ ├─ Inspect Skill ← Skill defect inspection
│ ├─ Inspect Tool ← Tool schema defect inspection
│ ├─ Inspect Cross ← Cross-artifact defect inspection
│ └─ Inspect Permission ← Permission-responsibility proportionality inspection (QD-PM)
│
├─ Risk (Risk Scanning)
│ ├─ Risk Explicit ← Explicit logic artifact risk detection
│ └─ Risk Implicit ← Implicit agent runtime vulnerability scanning
│
├─ Quality (Service Quality)
│ ├─ Quality RAG-Agent ← RAG-agent service quality assessment
│ └─ Quality Tool-Agent ← Tool-agent service quality assessment
│
└─ Relevance (Impact Mapping)
└─ Relevance ← defect → risk/quality diagnostic mapping
Inspect performs static defect and permission checks on logical artifacts; Risk Explicit audits explicit risks in artifacts; Relevance maps discovered defects to Risk attack surfaces and Quality failure modes; Risk Implicit and Quality complete dynamic verification and evaluation in the shadow sandbox; finally, Gate makes the release decision based on the baseline and ratchet mechanism. This closed loop automatically reruns with every logical artifact version change.
| Who | What you get |
|---|---|
| Leaders (CTO/CIO) | A standardized, quantified quality-and-safety scaffold for enterprise Agent programs — decisions backed by scores and evidence, not demos |
| Builders (Agent Dev/Test, Platform Eng, FDE) | Catch logic defects before they ship, get guided fixes, and wire the governance loop into CI/CD for hot iteration |
| Assurers (QA, Security Audit) | Version-synced quality evaluation and explicit/implicit risk baselines, with ratchet gates that block regressions on every artifact change |
The framework specification is fully open; the Inspect CLI is open source (Apache 2.0); Risk and Quality are offered as SaaS / self-hosted.
- Read the framework — a guided reading path: Framework Home
- Try the Inspect CLI — install and run static defect checks locally or in CI/CD: Try the Tools ·
pip install sanityops-cli - Try the live demo — the full Inspect + Risk + Quality loop: demo.sanityops.org. Sign-up is invite-only (LLM inference costs) — click Get Code on the registration form.
- Browse sample logical artifacts — System Prompts, Skills, and Tool Schemas for three example agents (energy, freight logistics, medication): samples/artifacts/
- Review sample audit reports — formal PDF reports for a "Voyager" sample agent, spanning Inspect, Risk (Explicit & Implicit), and Quality, each with a full audit and an executive summary: public/samples/report/
| Activity | What it answers | Docs |
|---|---|---|
| Inspect | Does my Prompt / Skill / Tool Schema / cross-artifact contract contain logic defects? Are granted permissions proportional to responsibilities? | Inspect |
| Risk | What explicit risks are in the artifacts, and what implicit vulnerabilities surface at runtime in a production-equivalent shadow environment? | Risk |
| Quality & Gate | Is user-visible quality (RAG-Agent / Tool-Agent) good enough, and did this version regress? | Quality |
| Relevance | Which risks and quality failures may this defect be associated with — and what should be re-verified as a priority after the fix? | Relevance |
Full scenario list (with tool mapping)
| Scenario | Typical Example | Reference Document | Tool |
|---|---|---|---|
| Prompt quality self-check | Before submitting a new System Prompt, check for contradictions, unclear boundaries, missing permissions, and other defects | Inspect Prompt | 🟢 Defect Inspector |
| Skill definition audit | Check whether the Skill's trigger conditions, permission declarations, and failure strategy contain logical flaws | Inspect Skill | 🟢 Defect Inspector |
| Tool Schema compliance check | Verify whether Tool Schema parameter constraints and side-effect declarations comply with specifications | Inspect Tool | 🟢 Defect Inspector |
| Cross-artifact consistency verification | Verify whether authorization, parameters, and contracts among Prompt–Skill–Tool are consistent | Inspect Cross | 🟢 Defect Inspector |
| Permission proportionality check | Check whether permissions granted to an Agent exceed its responsibility boundaries, evaluating permission-responsibility proportionality | Inspect Permission | 🟢 Defect Inspector |
| Explicit risk audit | Audit and quantitatively grade dangerous expressions, scripts, and dangerous authorizations in logical artifacts | Risk Explicit | 🔵 Risk Scanner |
| Implicit risk attack verification | Simulate sophisticated logical attacks in the shadow environment to detect runtime vulnerabilities | Risk Implicit | 🔵 Risk Scanner |
| Tool-Agent reliability evaluation | Evaluate the task success rate of tool-based Agents, establishing risk-driven testing rigor | Quality Tool-Agent | 🔵 Quality Evaluator |
| RAG-Agent quality evaluation | Evaluate knowledge-based Agents' accuracy, completeness, relevance, traceability, and timeliness through four dimensions and 12 metrics | Quality RAG-Agent | 🔵 Quality Evaluator |
| Enterprise Agent go-live admission | Before an enterprise's Agent goes live, pass three Gates — Inspect + Risk + Quality — and generate a release report | Core, Relevance | 🔵 Full tool suite |
| Defect→Risk/Quality diagnosis | Link defects found by Inspect to Risk attack surfaces and Quality failure modes to locate root causes | Relevance | 🔵 Full tool suite |
| Quality regression | After a logical artifact is modified, regression testing reveals quality degradation, forming the basis for release decisions | Quality RAG-Agent | 🔵 Quality Evaluator |
| Compliance audit evidence generation | Provide auditors with traceable version records, evaluation reports, and Gate decision evidence | Core | 🔵 Full tool suite |
See how SanityOps relates to other tools and frameworks:
- SanityOps Positioning — Framework positioning and ecosystem relationships
- SanityOps and FDE — Relationship with Forward Deployed Engineers
- SanityOps and Harness — Relationship with Harness runtime layer
- SanityOps vs Promptfoo — Red teaming and adversarial testing
- SanityOps vs RAGAS — RAG evaluation and quality metrics
- SanityOps vs NVIDIA SkillSpector — Skill security scanning
Terms appearing in this document such as logical artifact, Prompt / Skill / Tool Schema / Permission, ratchet mechanism, shadow sandbox, explicit/implicit risk, Gate, Baseline, and Harness are all standardized against the Core Appendix A terminology glossary as the baseline.
v1.0 · September 2026
- ✅ Framework 1.0 core specification released
- ✅ A complete methodology system for Inspect, Risk, Quality, and Relevance has been formed
- ✅ Inspect CLI fully open-sourced, with a free CLI provided
- ✅ Commercial SaaS / self-hosted services and reports available
v1.1 · November 2026
- 🔜 DMC 1.0 derivative model capability evaluation and optimization method
- 🔜 Inspect A2A 1.0, A2A Service Contract Defect Inspection Specification
You are welcome to participate in improvements via Issues, Pull Requests, or Discussions:
- Submit new logical artifact defect patterns
- Share risk audit and attack verification cases
- Contribute quality evaluation benchmarks and test sets
- Discuss the evolution of terminology, rules, grading, and Gate
The Framework specification of this project is licensed under CC BY-SA 4.0; the open-source Inspect CLI is separately licensed under Apache 2.0.
- Website: www.sanityops.org
- GitHub: sanityops-org/sanityops-framework
- Discussions: GitHub Discussions
- Email: hello@sanityops.org