Skip to content

chore: InputGuard v0.3.0 release — tested docs, integration recipes, measured FP benchmark, changelog - #14

Merged
obvious-autobuild[bot] merged 19 commits into
mainfrom
release/v0.3.0
Sep 18, 2026
Merged

obvious-autobuild[bot] merged 19 commits into
mainfrom
release/v0.3.0

Conversation

@obvious-autobuild

Copy link
Copy Markdown
Contributor

Why

InputGuard v0.3's implementation is complete and merged (PRs #3-#13), but the release artifacts were not: the README still carried v0.1-era examples that had already drifted from the code once, there were no integration recipes, the false-positive benchmark doc did not state the labeling-guide target contrast, and the 0.3.0 changelog was missing the docs wave. This PR is the release vehicle that completes those items on release/v0.3.0 so main receives the finished 0.3.0 tree.

(Release PR #2 was already squash-merged before the docs work landed, so it could not be reopened or marked ready; this PR carries the remaining release commits, rebased onto main's history by a verified merge -s ours reconciliation that keeps the release tree as a strict superset.)

What changed

  • README rewritten for the v0.3 surface with an anti-drift gate. New sections: public exports (including the extension API), Policy customization, custom rule and domain registration, follow-up questions, multilingual degradation, and CLI usage. tests/test_readme_examples.py executes every README python block in document order, diffs the embedded to_dict() JSON against live analyzer output, and checks the export list, rule tables, gap vocabulary, and CLI output against the registry — the drift that produced the v0.2 README bug now fails CI instead of shipping.
  • Integration recipes: docs/recipes/langchain.md and docs/recipes/litellm.md — copy-paste patterns over the to_dict() serialization boundary (preflight-and-gate; clarify-then-complete). Framework packages stay out of core and out of the dependency tree; recipe blocks are example-tested with framework stubs in tests/test_recipes.py.
  • False-positive benchmark doc refreshed from a measured run of eval/measure_fp.py on this exact tree: 116/121 overall, FP 3.3% / FN 0.8%, zero false positives on true negatives and degradation rows, degradation notes on 14/14 rows, PF rows 19-20 ms. It now states the labeling-guide target honestly: fixture-style boundary FP measured 4/6 against the 0/6 target (v0.2 baseline 6/6), driven by the matcher's deliberate inflection tolerance — documented as an eval-driven trade-off for a future release, never a label edit.
  • CHANGELOG 0.3.0 completed with the docs wave; version 0.3.0 and Development Status :: 4 - Beta were already in pyproject.toml and __init__.__version__ (verified, unchanged).

Tradeoffs

  • The recipe tests run framework imports through minimal stubs, so they verify the inputguard-side contract, not the frameworks themselves — no framework dependency enters the dev tree either.
  • The 4/6 boundary result is documented as measured rather than forced to the target: the residual rows come from the deliberate inflection-tolerance choice in matching.py (recall for natural word forms), and changing it is a dedicated eval-driven decision, out of scope for the release docs.

How to review

  • Anti-drift gate: tests/test_readme_examples.py — try editing a README value and watch the suite fail.
  • Measured rates: python3 eval/measure_fp.py and compare with docs/false-positive-benchmark.md.
  • The hard constraints hold: the 89 original baseline tests pass unmodified, zero runtime dependencies, additive-only API.

Verification

  • Full suite: 467 passed (includes 89 baseline tests unmodified, README anti-drift gate, recipe gate).
  • ruff check inputguard tests eval clean; mypy --strict inputguard clean (20 source files).
  • python3 eval/measure_fp.py re-run on this tree matches the documented numbers.

Human author: Kalisetti Nihanth Naidu (nihanthnaidu007@gmail.com)

🔗 Obvious Project · 🧵 Obvious Thread

ObviousApp and others added 18 commits September 17, 2026 18:53
…res (#3)

* feat(rule-protocol-registry): typed Rule protocol and in-process registry

Open the engine (spec art_bTvdPdJS §1): rules and domains become
registered data instead of code paths.

- inputguard/registry.py: runtime-checkable Rule protocol (id, domain,
  severity, gap + check(text, intent)) and RuleRegistry with
  register_rule (decorator and imperative forms) and register_domain
  (priority-ordered intent signals, single empty-terms fallback intent).
  Duplicate rule ids, unknown severities, and — once a domain exists —
  undeclared rule domains raise at registration time.
- detector.py: the coding priority chain becomes INTENT_SIGNALS data;
  detect_intent gains an optional signals parameter so registered
  domains run the same chain; normalize() becomes public so analyze()
  can honor the "rules receive normalized text" contract.

Nothing consumes the registry yet; behavior is unchanged (89 tests
green). Zero new dependencies.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* refactor(built-ins-via-registry): dispatch analyze() through the registry

Replace the hardcoded domain whitelist and the if/elif intent dispatch
with registry lookups (spec art_bTvdPdJS §1):

- analyzer.py: analyze() resolves domain signals via
  REGISTRY.get_domain_signals (unknown domain still raises ValueError,
  now naming every registered domain) and runs rules via
  REGISTRY.rules_for_intent — the deduped, deterministic-finding-order
  contract of the v0.2 runners is preserved by the registry's
  registration-order iteration.
- rules/*: each of the 19 built-in rules gains a thin registry adapter
  class exposing the v0.2 check function through the Rule protocol,
  registered with @register_rule — the built-ins dogfood the exact
  extension path users get. The coding catch-all adapter re-runs the
  built-in build rules internally to preserve its fires-only-when-empty
  semantics under the frozen check(text, intent) contract.
- rules/__init__.py: registers the coding domain (INTENT_SIGNALS chain
  plus all 19 rule adapters) through register_domain; the v0.2 runner
  exports are unchanged.

Behavior-identical: the 89 existing tests pass unmodified.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* fix(scorer-loud-failures): unknown severities raise; unknown gaps keep advice

Close the two silent-failure edges the survey pinned (art_CnghyDxp §2.3,
§2.4) per spec art_bTvdPdJS §6:

- scorer.py: calculate_score validates every finding's severity up front
  and raises ValueError on an unknown one — the silent zero penalty from
  .get(severity, 0) is gone (rank comparisons and penalty lookups go
  through validated paths).
- recommender.py: get_recommendations no longer drops gaps without a
  curated entry; unknown gaps get a documented four-key generic
  fallback (_fallback_recommendation), keeping advice complete for rule
  authors who ship a new gap.

Existing behavior for known severities and gaps is unchanged: the 89
existing tests pass unmodified.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* test(registry-invariants): registry contract and parity tests

New tests/test_registry.py pins the foundation's acceptance points:

- dogfooding: all 19 built-in rules registered through the registry,
  coding domain signals in priority order with a single fallback intent,
  deterministic build-intent rule order;
- validation: duplicate rule id, unknown severity, and unknown domain
  raise ValueError at registration; decorator form registers and
  satisfies the Rule protocol;
- loud failures: unknown severity raises in calculate_score (including
  alongside known severities); unknown gaps get the documented four-key
  fallback, never a silent drop;
- end to end: a custom registered rule changes a result (finding, gap,
  score -15, fallback recommendation); a registered custom domain
  analyzes; duplicate domain registration raises;
- invariants: the registry-walking completeness check (every built-in
  gap has a complete recommendation entry), dispatch parity with the
  v0.2 runners across nine inputs, and 65 parallel analyze() calls
  across 8 threads matching sequential results.

Full suite: 114 passed (89 pre-existing, unmodified).

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

---------

Co-authored-by: Obvious <obvious@obvious.ai>
Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
…ult field (#4)

Per-gap templated clarifying questions (one or two per gap, deduped, gap-ordered), slot fills for function/dataset names, a documented fallback for unknown gaps, the additive AnalysisResult.follow_ups field serialized in to_dict(), and the registry-walking completeness invariant extended to require follow-ups for every built-in gap. 139/139 tests green; existing 114 unmodified.

- feat(follow-up-templates): per-gap clarifying question engine
- feat(result-field): additive follow_ups on AnalysisResult
- test(follow-up-completeness): extended invariant + engine behavior tests

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
…ath for uncovered scripts (#5)

* feat(script-probe): add unicodedata script-histogram language probe

Pure, thread-safe script classification built on unicodedata only —
each letter's Unicode name mentions its script, so token scanning
classifies scripts with no hardcoded range tables. Classifies the
input's dominant script, a coarse script-derived language guess, and
the English-heuristic coverage band (full / partial / none / unknown).
A deterministic stride sample bounds probe cost on arbitrarily long
input.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* feat(degraded-path): explicit degradation for scripts without heuristic coverage

Wire the script probe into analyze(): uncovered scripts skip the
English-only rules outright, take a 20-point confidence penalty (score
80 — never 'ready' in either mode), and report detected_intent
'undetermined' instead of an unearned fallback. Results carry the three
additive fields detected_language, heuristic_coverage, and
degradation_note; to_dict() includes them additively. Covered scripts
(full/partial coverage) run the unchanged pipeline — partial adds a
note without a penalty.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* test(non-english-honesty): probe-P2 regression, Unicode samples, thread parity

The probe-P2 input (Chinese) must return a degraded result with a
degradation_note — never ready/100. Also covers: degraded shape and
mode behavior, per-script degradation, English parity (probe fields on
covered input, byte-parity scores: 'fix my code' -> 35, repeated build
-> 50), partial-coverage rules-run-with-note, validation order,
additive to_dict() keys, systematic no-crash Unicode samples with the
note-consistency invariant, bounded long-input probing, and 64-worker
parallel analyze() consistency across English/Chinese inputs.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* fix(tests): adapt to_dict key-set expectations to the merged additive contract

PR #4's parity test asserted the exact to_dict key set with follow_ups
as the only additive key; the three language-probe fields are additive
per the v0.3 spec, so the expectation now covers the merged additive
set. The key-order test documents the merged 11-key contract.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

---------

Co-authored-by: Obvious <obvious@obvious.ai>
Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
… eval/ (#7)

* feat(eval-prompt-set): version the 121-case clarity-evaluation set as eval/

Checkpoint the labeled evaluation set (workbook wb_q30ippCW + labeling
guide art_XPvHhPeZ) into the repo: cases.csv (121 rows, QUOTE_ALL/CRLF),
README with the labeling criteria, per-case_type rubric, gap-vocabulary
calibration contract, and a stdlib-only measure_fp.py that runs the
analyzer over every row and reports per-case_type and overall FP/FN rates.
The tool is a measurement, not a gate: it always exits 0 and treats
label/behavior disagreement as the data.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* test(eval-loader): smoke-test the eval dataset mechanics

Three smoke tests: the CSV parses with all nine required columns and
non-empty texts, case ids are unique, and measure_fp.py runs end to end
on a 10-case sample. Deliberately no label-vs-analyzer assertions — that
agreement is what the measurement tool reports over time.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

---------

Co-authored-by: Obvious <obvious@obvious.ai>
Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
…s, registration guards (#8)

* fix(registry-contract-spec-b1): revert Rule.check to the spec's check(self, text) signature

The merged foundation (PR #3) dispatches rule.check(text, intent), but the
pinned spec (art_bTvdPdJS section 1) defines check(self, text). The intent
parameter is redundant by construction: analyze() dispatches
REGISTRY.rules_for_intent(detected_intent), which filters on
rule.domain == intent, so every rule invoked already knows the intent as its
own domain member — and all 19 built-in adapters ignore it.

The mismatch is not cosmetic: registration only checks callable(rule.check),
so a rule written per the spec registers cleanly and then crashes analyze()
mid-run with "TypeError: check() takes 2 positional arguments but 3 were
given" (review probe P1). Spec and code must agree before wave-2/3 rule
authors copy either version.

- inputguard/registry.py: Rule protocol back to check(self, text); module
  and member docstrings updated to match.
- inputguard/analyzer.py: dispatch rule.check(normalized).
- inputguard/rules/: all 19 built-in adapters drop the unused intent
  parameter; InsufficientContextRule docstring no longer cites (text, intent).
- tests/test_registry.py: helper and decorator-form test rules updated.

Part of release PR #2. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* fix(registry-intent-scoping-b2): globally-unique intents, registration guards, loud rule attribution

Closes the cross-domain leak and the registration holes the adversarial
review proved against the merged foundation (art_hC18m78C). All guards
fire at registration time, before any mutation.

- B2: register_domain raises ValueError when an intent name is already
  declared by another registered domain, naming both domains. Rules
  dispatch on intent name alone (rules_for_intent filters on
  rule.domain), so a shared intent name ran one domain's rules inside
  the other's analysis (review probe P6b: a legal rule fired inside a
  coding analysis). Duplicate intent names within one signals argument
  raise too — the ownership map must be unambiguous.
- C4: the self-contradictory error ("declares domain 'coding', which no
  registered domain declares as an intent. Registered domains: 'coding'")
  now states the actual constraint and enumerates the valid intent
  names; the Rule protocol documents that domain holds the intent name
  the rule is registered under, never a domain name.
- C3: duplicate rule ids within one register_domain call raise — the
  old code silently dropped the second rule (review probe P2).
- N3: _validate checks the check() signature via inspect.signature
  binding; a rule whose check cannot accept the single positional text
  argument is a registration error, not a mid-analyze TypeError (review
  probe P1's crash shape).
- C2-lite: analyze() wraps rule dispatch and re-raises RuntimeError with
  the rule id and its registration origin ("at <file>:<line>", captured
  at _add time and exposed via RuleRegistry.rule_origin), chaining the
  original exception. Blast-radius contract documented on the Rule
  protocol and RuleRegistry: a rule exception aborts analyze() by
  design, registration is permanent, REGISTRY is process-global.
- C6: the extension API (register_rule, register_domain, REGISTRY, Rule)
  is exported at package top level, additive on the untouched v0.2
  four-export surface.

Part of release PR #2. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* test(registry-contract-guards): pin every remediated registry behavior in the suite

Permanent tests for the review's findings so the gates enforce the chosen
contract (the review found none of these covered):

- spec-compliant check(text) rule registers and fires end to end (P1)
- wrong-arity and zero-argument check rejected at registration, nothing
  enters the registry (N3)
- cross-domain intent collision raises naming both domains (B2/P6b),
  leakage probe cannot recur
- duplicate intent within one signals argument raises (N1 registry half)
- same-call duplicate rule ids raise and leave no partial registration (C3/P2)
- rule exception aborts analyze() with rule id + registration origin and
  the original cause chained (C2-lite/P4)
- extension API importable at top level; v0.2 four-export surface intact (C6)

Suite: 123 passed (114 baseline + 9 new).

Part of release PR #2. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* test: adapt PR #4 follow-ups test rule to the spec check(text) signature

Rebase adaptation: PR #4's unknown-gap test registered a rule with the
retired check(text, intent) signature; the N3 registration guard now
correctly rejects it. The test's subject (unknown-gap fallback follow-up)
is orthogonal to the signature — the rule moves to the spec's
check(self, text). Suite: 202 passed.

Part of release PR #2. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

---------

Co-authored-by: Obvious <obvious@obvious.ai>
…les (#6)

* fix(word-boundary-generalization): shared word-boundary matcher across detector and all rule modules

Raw substring containment (detector.py:57-58) read "fixture" as the debug
signal "fix" (probe P1) and flagged innocent input at 35/needs_clarification.
All term lookups now route through inputguard/matching.py, which matches
terms as standalone words using the boundary class the coding matcher
already used, extended with inflectional endings so suffix-shaped hits
(bugs, errors, debugged, refactoring, profiling, stack traces) keep firing.
Coding rules keep their v0.2 multiword token fallback via token_fallback=True;
other callers get strict adjacent-phrase matching. No-alnum terms (/, =>, ->)
keep substring semantics. Finders compile once per term-set (lru_cache).

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* test(fp-regressions): pin word-boundary matcher contract and probe-P1 fix

Regression coverage for the shared matcher: the P1 "fixture" input no
longer reads as debug and is not hard-flagged; embedded words (fixture,
refix, prefix, praised) stop matching; inflected forms (bugs, errors,
debugged, refactoring, profiling, stack traces) keep firing so the
boundary tightening adds no false negatives; phrases match adjacently
with final-word inflections; the coding token fallback stays opt-in;
symbol terms keep substring semantics.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

---------

Co-authored-by: Obvious <obvious@obvious.ai>
Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
…ing, input cap, and score breakdown (#9)

* feat(policy-object): frozen validated Policy with rule filters and word floor

Every scoring tunable moves into a frozen, validated Policy dataclass whose
defaults are the v0.2 constants: ready_at 85, usable_at 60, strict_clarify_at
65, penalties 5/15/25, min_words 3, max_chars 10000, borderline_at 74.
Validation rejects mis-ordered bands, out-of-range values, uncompilable
allow_patterns, and disabled rule ids that are not registered. InputGuard and
analyze() take an optional policy (a per-call policy overrides the guard's).

disabled_rules skips registered rules by id; allow_patterns regex matches skip
flagging entirely; min_words gates the insufficient_context vague rule below
its floor. Penalties flow from Policy through calculate_score(findings,
policy=None).

Closes review concern C5 (art_hC18m78C): one severity vocabulary
(policy.SEVERITIES) shared by registration and a new single choke point,
scorer.require_known_severity, which the analyzer applies to every emitted
finding; registry.KNOWN_SEVERITIES is cross-pinned to it by test.

Zero new dependencies; stdlib only.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* feat(two-layer-banding): policy-driven status bands and a borderline near-miss signal

get_status(score, mode, policy=None) now reads its bands from Policy
(ready_at 85, usable_at 60, strict_clarify_at 65 — the v0.2 constants, pinned
byte-exactly in tests/test_banding.py), replacing the four hard-coded scorer
thresholds. Severity still decides what fires; bands decide what happens.

AnalysisResult gains an additive borderline: bool field (default False, in
to_dict) set when the score lands in [policy.borderline_at, ready_at) — the
distinct near-miss "worth one more pass" signal from the spec's result states.
interpretation_note is untouched, so existing serialized results are
byte-identical under default policy.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* feat(input-cap): enforce Policy.max_chars with visible truncation

analyze() now slices input to policy.max_chars before intent detection and
rule execution — analysis cost is bounded regardless of input size (v0.2 took
~3.8 s on a 1.6 MB input; capped analysis of the same input runs in
milliseconds). Truncation is never silent: AnalysisResult gains an additive
truncated: bool field (default False, in to_dict), and the allowlist is
evaluated against the same capped input so no scan escapes the budget.

Default cap 10_000 matches the Policy default, pinned by test.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* feat(score-breakdown): additive score_breakdown audit field on the result

calculate_score_with_breakdown(findings, policy=None) scores findings and
returns the audit trail — {"base": 100, "penalties": [{"code", "severity",
"points"(negative)}...], "final"} — one entry per distinct gap (or code when
gap is None) in first-occurrence order, matching the result's gaps order.
calculate_score keeps its v0.2 signature and delegates to it, so both stay in
lockstep by construction.

AnalysisResult gains an additive score_breakdown field (default None, deep-
copied in to_dict so a serialized snapshot can never be rewritten by later
mutation). With default policy the penalties mirror the v0.2 deductions
exactly, so serialized defaults differ only by the new keys.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* test(pin-v02-defaults): pin every Policy default byte-exactly to v0.2

Dedicated pinning tests guard the calibration-drift risk the spec flags:
each Policy default (ready_at 85, usable_at 60, strict_clarify_at 65,
penalties 5/15/25, min_words 3, max_chars 10000, borderline_at 74) is
asserted three ways — dataclass field defaults, constructed instance values,
and the literal source line (underscore separators normalized) — plus a
no-unpinned-fields guard so future fields cannot ship uncalibrated.

Behavioral pins reproduce the observed v0.2 reference outputs under default
policy (75/usable_with_warnings, 35/needs_clarification, 50, strict
blocked/needs_clarification) and the v0.2 result and recommendation key
shapes.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* test: integrate policy keys into sibling additive-contract tests

The followups parity test and the language ordered-key-list test pin the
exhaustive to_dict key set as of their own merge; the policy-calibration
fields (borderline, truncated, score_breakdown) are additive per the compat
contract, so the expected sets extend rather than the fields disappearing.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

---------

Co-authored-by: Obvious <obvious@obvious.ai>
Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
…e advice (#10)

* feat(writing-domain): first-party writing domain — six rules, compose intent, complete advice

Register the writing domain through the same registry path user domains
take: one globally unique fallback intent ("compose") and six rules
covering the eval vocabulary's six gaps — audience, purpose,
structure/format, source material, context, completeness — at the
vocabulary-pinned severities (high/high/medium/high/medium/low).

All term matching goes through the shared word-boundary matcher
(inputguard.matching, PR #6). Rules follow the trigger-and-satisfy shape;
the source-material rule fires only when existing text is referenced but
not provided. Recommendations and follow-up entries ship for all six gaps
in the same change, keeping the registry completeness invariant whole.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* test(writing-domain): per-rule, completeness, integration, no-leakage, and eval-label coverage

Five layers: per-rule fire/silent behavior; the writing-domain
completeness invariant (every writing gap has a four-key recommendation
and one or two follow-up questions); analyze() integration end to end
(underspecified -> needs_clarification at score 30, fully specified ->
ready, partially specified -> usable_with_warnings); no-leakage in both
directions at registry and analyze level; and the 14 labeled English
writing rows from eval/cases.csv pinned as a calibration net.

The exact-set registry expectations extend from 19 coding rules to 25
rules across two domains — the guards are strengthened, not weakened.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

---------

Co-authored-by: Obvious <obvious@obvious.ai>
Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
…complete advice (#11)

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
Function-word layer (fr/es/pt stop-word lexicons with English margin) plus uncovered-run layer (3+ consecutive uncovered-script letters) route accented-Latin and mixed English+Han input to the explicit degraded path; DG-012's silent ready and DG-011/013/014's unnoted rule findings are gone. Degradation eval FP 3/14 -> 0/14, note honesty 10/14 -> 14/14, all 121 probe verdicts otherwise byte-identical.
- Bump pyproject.toml and inputguard.__version__ to 0.3.0; write the full
  0.3.0 CHANGELOG section from the merged release-wave history.
- Report the literal "degraded" status on the degraded path in both modes;
  add the mixed-script (>=5% uncovered remainder) and non-English-Latin
  (<10% English evidence, filler-resistant denominator) gates so French/
  Spanish/Portuguese and mixed English+Han prompts degrade with zero gaps
  instead of spurious ones; preserve the truncated flag on that path.
- Linear dataset-filename extraction (follow-ups + data-analysis rule)
  with 1K/10K timing tests; benchmark doc records the before/after.
- README: rewritten Non-English section, full result-object field table
  (follow_ups, probe fields, borderline, truncated, score_breakdown),
  writing and data-analysis rule tables.
- eval: docs/false-positive-benchmark.md from measured results.
- CI: minimal GitHub Actions workflow running pytest on push/PR (3.9+3.13).

Eval: 116/121 matching (all 14 degradation rows, notes on 14/14); the 5
residual mismatches are pre-existing coding boundary/intent rows.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
Human author: Kalisetti Nihanth Naidu (nihanthnaidu007@gmail.com)
#13)

* test(registry-contract-guards): CI-enforced guards for the three review findings

Dedicated suite for the adversarial-API-review gaps (art_hC18m78C), as the
named acceptance artifact:

- C3/P2: duplicate rule ids inside one register_domain call raise
  ValueError naming the id; a rejected batch mutates nothing.
- B1/N3/P1: a rule whose check cannot be called as check(self, text)
  (extra param, zero param, keyword-only) is rejected at registration with
  an actionable message; defaulted and spec signatures accepted and fired.
- C2/P4: an exception inside check() aborts analyze() as a RuntimeError
  carrying the rule id, its registration origin, and the original
  exception chained -- via the register_domain path (register_rule is
  covered in test_registry.py).

The registry-isolation fixture moves to tests/conftest.py so the guard
suite and the policy suite share one snapshot/restore implementation
instead of importing across test modules.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* build(typing): annotate for mypy --strict without behavior change

Pre-work for the CI strict-typing gate: missing Iterable[str] on the five
rule modules' _contains_any helpers, Dict[str, Any] / Dict[str, str] generics
on types.py and recommender.py, a pinned Rule-typed local in
registry._instantiate, and the writing signals Tuple[str, ...] type argument.
Also removes a literal duplicate "filter by"/"sort by" set entry in the
feature scope signals (ruff B033); set semantics make this a no-op at
runtime, and rules/__init__ marks its v0.2 signal re-exports as intentional
via __all__ (ruff F401).

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* ci(actions-matrix): GitHub Actions CI with lint, typed, tested, benchmark, wheel gates

The v0.2 failure mode was "every change ships on trust" — no CI, lint,
typing, or coverage tooling at all. This adds the production gate set from
the spec (§8) as three jobs:

- test: Python 3.9-3.13 matrix running ruff, pytest with a 90%
  branch-coverage floor (--cov-branch --cov-fail-under=90 --strict), and
  mypy --strict (finally checking the shipped py.typed).
- benchmark: latency budgets as merge gates, not notes. Budgets are derived
  from measured numbers (build-intent p50 21.94 ms at 10k chars including
  the C1 InsufficientContextRule double-run, debug 6.59 ms, ready 5.62 ms;
  1.6 MB input bounded by the Policy cap at 17.1 ms vs the v0.2 unbounded
  3.8 s scan). Hard gate is p50 with CI-noise headroom; p50/p99 print to
  the job log per spec. Budgets and derivation documented in
  tests/test_benchmarks.py.
- wheel-zero-dep: the zero-dependency promise verified against the built
  artifact — wheel METADATA carries no Requires-Dist, a bare venv install
  pulls no third-party package, and the public API works from the
  installed wheel. Smoke steps run from a scratch directory so the repo's
  local inputguard/ package can never shadow the installed wheel.

pyproject gains the matching tool config ([tool.ruff] E/F/W/B with E501
ignored, [tool.mypy] strict, explicit pytest testpaths) and the expanded
[dev] extras the CI jobs install; all dev tooling stays out of runtime
dependencies.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* feat(cli): inputguard analyze with argparse, to_dict JSON, CI-friendly exit codes

The developer surface the adoption research flagged (spec §5): a
zero-dependency argparse CLI shipped as a console script.

- "inputguard analyze TEXT [--stdin] [--domain D] [--mode M]
  [--format text|json] [--min-score N]"; JSON output is the result's own
  to_dict() contract, byte-equal to the library's (asserted by test).
- Exit codes documented in the module docstring and --help: 0 = analysis
  ok and at/above the --min-score floor when given; 1 = below the floor
  (the commit-hook / CI gate); 2 = usage error (missing text, unknown
  domain) reported cleanly on stderr, never a traceback.
- Text rendering follows the spec §5 example shape: status, clarity,
  intent, domain, missing gaps, ask lines, and the degradation note when
  the language probe marks the result degraded (probe P2 input shows the
  note, never a silent ready).

Entry point added in pyproject [project.scripts]; stdlib argparse and json
only — the zero-dependency promise is untouched.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* chore(packaging-metadata): 0.3.0, Beta status, 3.13 classifier, project URLs

The packaging drift the survey flagged is closed: classifiers extend to
Python 3.13 (the runtime the project is developed and tested on),
Development Status moves to 4 - Beta per the release plan (1.0 deliberately
deferred until the extension API sees real use), [project.urls] points at
the repository, changelog, and issue tracker, and inputguard.__version__
syncs with the pyproject version. The changelog's Unreleased v0.3 section
is promoted to [0.3.0] with the CI and CLI entries from this PR.

Also amends the wheel-gate assertion to what the zero-dependency promise
actually means: no UNCONDITIONAL Requires-Dist entries (setuptools emits
extra-gated `Requires-Dist: ...; extra == "dev"` lines for the optional
tooling — verified against the built wheel, and the bare-venv install
proves none of them install at runtime).

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* fix(benchmark-budgets): gate latency only in the untraced benchmark job, budgets from measured CI data

The first CI run (35277175054) failed the latency asserts inside the
coverage pass: coverage tracing roughly triples per-call cost (debug
p50 21.14 ms vs 6.59 ms untraced; ready 20.35 ms vs 5.62 ms), and runner
hardware is slower than the dev sandbox. Gating the same budgets under
coverage in the test job AND untraced in the benchmark job was a design
error — two different measurement conditions cannot share one threshold.

- The coverage pass now ignores tests/test_benchmarks.py (with the
  observed numbers in a comment); latency gates belong exclusively to
  the dedicated benchmark job, which runs untraced and reflects
  real-world analysis cost.
- Budgets re-derived from measured environments documented in the test
  module: sandbox untraced (build 21.94 / debug 6.59 / ready 5.62 ms
  p50) and the CI-traced observation above. New p50 budgets: build 65
  ms (~3x), debug 25 ms (~3.8x), ready 25 ms (~4.4x); p95 sustained
  guards at 2.5x unchanged. Still order-of-magnitude gates: they catch
  algorithmic regressions (accidental O(n^2), a lost cap) that blow
  past by 10x+, while absorbing CI hardware variance.

Local verification of the exact split: coverage pass 441 passed, 96.29%
branch (floor 90); benchmark job untraced 4 passed; ruff + mypy --strict
clean.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* fix: give the extraction linearization bounds CI headroom

test (3.10) flaked on the rebased head (run 35277834661): the dotted-filler
linearization test's absolute bound asserted a single wall-clock sample
under 50 ms, and the coverage-traced 3.10 runner measured 50.61 ms. The
meaningful guard is the 40x growth-rate assertion — the regression it
protects against measured 1.3 s at 10 K, ~26x the bound. Both timing tests
now use a 250 ms absolute bound (5x+ detection margin, no CI-hardware
flake); the growth-rate assertion is unchanged.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

---------

Co-authored-by: Obvious <obvious@obvious.ai>
Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
…se tree is a strict superset, verified by tree-hash equality
…ple executed by tests

The v0.2 README drifted from recommender.py within two releases. The
rewritten README documents the v0.3 surface (extension API, Policy,
follow-ups, degradation, CLI) and every example now runs: a new
tests/test_readme_examples.py executes all python blocks in document
order in a subprocess, diffs the embedded to_dict() JSON against live
analyzer output, and checks exports, rule tables, gap vocabulary, and
CLI output against the live registry. Example values carry real
assertions sourced from runtime probes.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
…_dict()

Copy-paste integration patterns over the to_dict() serialization
boundary — preflight-and-gate and clarify-then-complete. The recipes
are documentation only: framework packages stay out of core and out of
the dependency tree. tests/test_recipes.py executes every recipe block
in document order, stubbing the single framework idiom each wiring
block uses, so the inputguard-side contract cannot drift silently.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
…get honestly

Re-ran eval/measure_fp.py on the release branch: overall 116/121 match,
FP 3.3% / FN 0.8%, zero false positives on true negatives and
degradation rows, notes on 14/14 degradation rows, PF rows 19-20 ms.
Adds the labeling-guide target contrast the doc lacked: fixture-style
boundary FP measured 4/6 (v0.2 baseline 6/6, target 0/6) — driven by
the matcher's deliberate inflection tolerance, documented as an
eval-driven trade-off for a future release, not a label edit.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
Version 0.3.0 and the Development Status :: 4 - Beta classifier were
already in pyproject.toml and inputguard.__version__ — verified, no
bump needed. This commit completes the 0.3.0 entry with the docs wave:
tested README rewrite, integration recipes, and the measured
false-positive benchmark with the boundary-target contrast.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
@obvious-autobuild
obvious-autobuild Bot marked this pull request as ready for review September 17, 2026 22:04
…ample

The custom-rule example annotated check() with PEP 604 union syntax
(-> RuleFinding | None), which is evaluated at definition time and
raises TypeError on Python 3.9 — caught by the CI matrix's 3.9 job
running the README anti-drift gate. Use typing.Optional instead; an
AST scan of every executed README block confirms no other
runtime-evaluated 3.9-incompatible construct remains.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
@obvious-autobuild
obvious-autobuild Bot merged commit eff7a20 into main Sep 18, 2026
14 checks passed
@obvious-autobuild
obvious-autobuild Bot deleted the release/v0.3.0 branch September 18, 2026 12:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants