Skip to content

feat: CI matrix, latency budgets, CLI, and packaging metadata for v0.3 - #13

Merged
obvious-autobuild[bot] merged 7 commits into
release/v0.3.0from
feat/ci-packaging
Sep 17, 2026
Merged

obvious-autobuild[bot] merged 7 commits into
release/v0.3.0from
feat/ci-packaging

Conversation

@obvious-autobuild

@obvious-autobuild obvious-autobuild Bot commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Closes the v0.3 spec's §8 engineering baseline: until now the package shipped on trust — no CI, lint, typing, or coverage tooling at all. This PR adds the production gate set as three GitHub Actions jobs, the inputguard analyze CLI, and the release packaging metadata. Rebased onto release/v0.3.0 @ 70f41ab (the release-sweep commit; its minimal 21-line CI workflow is superseded by this PR's full matrix, whose jobs are a strict superset).

CI (.github/workflows/ci.yml), three jobs:

  • test — Python 3.9–3.13 matrix: ruff check, pytest -q --cov=inputguard --cov-branch --cov-fail-under=90 --strict (latency asserts excluded from the traced pass — they gate only in the dedicated benchmark job; coverage tracing roughly triples per-call cost, observed on run 35277175054), and mypy --strict inputguard (the shipped py.typed is finally checked).
  • benchmark — untraced latency budgets as merge gates; p50/p99 print to the job log.
  • wheel-zero-dep — the zero-dependency promise verified against the built artifact: wheel METADATA has zero unconditional Requires-Dist entries, a bare-venv install pulls no third-party package, and the public API + CLI work from the installed wheel alone (smoke steps run from a scratch directory so the repo's local inputguard/ can never shadow the wheel).

Registry contract guards (tests/test_registry_contract_guards.py) — the three tests the adversarial API review (art_hC18m78C) found missing, CI-enforced: same-call duplicate rule ids in register_domain raise with nothing mutating; a check signature that cannot be called as check(self, text) is rejected at registration with an actionable message; dispatch exceptions surface with rule id, registration origin, and the original exception chained — via the register_domain path (register_rule was already covered).

CLI (inputguard/cli.py) — inputguard analyze console script: stdlib argparse, --format json emitting byte-equal to_dict() output (asserted by test), exit codes 0 ok / 1 below --min-score / 2 usage error, degradation note surfaced in text output.

Packaging — version 0.3.0, Development Status → 4 - Beta, 3.13 classifier, [project.urls], dev tooling confined to [dev] extras (zero runtime dependencies preserved), explicit [tool.ruff]/[tool.mypy]/[tool.pytest.ini_options] config. The release sweep already bumped the version; this PR contributes the Beta status, classifier, URLs, and changelog entries (woven into the sweep's [0.3.0] section). Two post-rebase CI flakes in sweep-authored timing tests were fixed by giving single-sample absolute bounds CI headroom (250 ms; the guarded regression is 1.3 s).

Suite + coverage evidence

  • 455 tests pass locally on the rebased head; the 89 v0.2 baseline tests remain untouched and green.

  • Coverage 96.28% branch (--cov-branch --cov-fail-under=90 --strict — the exact CI command, run locally).

  • Latency, measured on the rebased head (untraced, Python 3.13, 10,000-char inputs = the default cap, 60 samples after warmup):

    Path Measured p50 Budget (hard gate, benchmark job)
    build-intent, vague (C1 InsufficientContextRule double-run) 17.83 ms p50 ≤ 65 ms (~3×)
    debug-intent, with findings 7.48 ms p50 ≤ 25 ms (~3.8×)
    ready path, score 100 6.21 ms p50 ≤ 25 ms (~4×)
    1.6 MB input under the Policy cap 18.7 ms, truncated=True < 1000 ms

    Budget derivation (documented in the test module): untraced sandbox measurements multiplied ~3–4× to absorb CI hardware variance, calibrated against the first CI run's traced observations (debug p50 21.14 ms, ready 20.35 ms under coverage tracing — the reason latency asserts were moved out of the coverage pass). The budgets still gate by an order of magnitude: algorithmic regressions (accidental O(n²), a lost cap) blow past them ~10×, as the v0.2 unbounded 3.8 s scan would. Hard gate is p50; p95 sustained-regression guards sit at 2.5×.

  • Wheel gate verified locally end to end: inputguard-0.3.0-py3-none-any.whl builds, installs into a bare venv (pip list → inputguard + pip only), imports and runs analyze() from outside the source tree, CLI JSON + exit-code smoke passes, and METADATA shows 0 unconditional / 6 extra-gated dev-only Requires-Dist entries.

Part of release PR #2.

Backend-only infrastructure change; visual evidence not applicable.

Human author: Kalisetti Nihanth Naidu (nihanthnaidu007@gmail.com)

🔗 Obvious Project · 🧵 Obvious Thread

@obvious-autobuild
obvious-autobuild Bot marked this pull request as ready for review September 17, 2026 21:31
ObviousApp and others added 6 commits September 17, 2026 21:34
…ew findings

Dedicated suite for the adversarial-API-review gaps (art_hC18m78C), as the
named acceptance artifact:

- C3/P2: duplicate rule ids inside one register_domain call raise
  ValueError naming the id; a rejected batch mutates nothing.
- B1/N3/P1: a rule whose check cannot be called as check(self, text)
  (extra param, zero param, keyword-only) is rejected at registration with
  an actionable message; defaulted and spec signatures accepted and fired.
- C2/P4: an exception inside check() aborts analyze() as a RuntimeError
  carrying the rule id, its registration origin, and the original
  exception chained -- via the register_domain path (register_rule is
  covered in test_registry.py).

The registry-isolation fixture moves to tests/conftest.py so the guard
suite and the policy suite share one snapshot/restore implementation
instead of importing across test modules.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
Pre-work for the CI strict-typing gate: missing Iterable[str] on the five
rule modules' _contains_any helpers, Dict[str, Any] / Dict[str, str] generics
on types.py and recommender.py, a pinned Rule-typed local in
registry._instantiate, and the writing signals Tuple[str, ...] type argument.
Also removes a literal duplicate "filter by"/"sort by" set entry in the
feature scope signals (ruff B033); set semantics make this a no-op at
runtime, and rules/__init__ marks its v0.2 signal re-exports as intentional
via __all__ (ruff F401).

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
…mark, wheel gates

The v0.2 failure mode was "every change ships on trust" — no CI, lint,
typing, or coverage tooling at all. This adds the production gate set from
the spec (§8) as three jobs:

- test: Python 3.9-3.13 matrix running ruff, pytest with a 90%
  branch-coverage floor (--cov-branch --cov-fail-under=90 --strict), and
  mypy --strict (finally checking the shipped py.typed).
- benchmark: latency budgets as merge gates, not notes. Budgets are derived
  from measured numbers (build-intent p50 21.94 ms at 10k chars including
  the C1 InsufficientContextRule double-run, debug 6.59 ms, ready 5.62 ms;
  1.6 MB input bounded by the Policy cap at 17.1 ms vs the v0.2 unbounded
  3.8 s scan). Hard gate is p50 with CI-noise headroom; p50/p99 print to
  the job log per spec. Budgets and derivation documented in
  tests/test_benchmarks.py.
- wheel-zero-dep: the zero-dependency promise verified against the built
  artifact — wheel METADATA carries no Requires-Dist, a bare venv install
  pulls no third-party package, and the public API works from the
  installed wheel. Smoke steps run from a scratch directory so the repo's
  local inputguard/ package can never shadow the installed wheel.

pyproject gains the matching tool config ([tool.ruff] E/F/W/B with E501
ignored, [tool.mypy] strict, explicit pytest testpaths) and the expanded
[dev] extras the CI jobs install; all dev tooling stays out of runtime
dependencies.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
…y exit codes

The developer surface the adoption research flagged (spec §5): a
zero-dependency argparse CLI shipped as a console script.

- "inputguard analyze TEXT [--stdin] [--domain D] [--mode M]
  [--format text|json] [--min-score N]"; JSON output is the result's own
  to_dict() contract, byte-equal to the library's (asserted by test).
- Exit codes documented in the module docstring and --help: 0 = analysis
  ok and at/above the --min-score floor when given; 1 = below the floor
  (the commit-hook / CI gate); 2 = usage error (missing text, unknown
  domain) reported cleanly on stderr, never a traceback.
- Text rendering follows the spec §5 example shape: status, clarity,
  intent, domain, missing gaps, ask lines, and the degradation note when
  the language probe marks the result degraded (probe P2 input shows the
  note, never a silent ready).

Entry point added in pyproject [project.scripts]; stdlib argparse and json
only — the zero-dependency promise is untouched.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
…ct URLs

The packaging drift the survey flagged is closed: classifiers extend to
Python 3.13 (the runtime the project is developed and tested on),
Development Status moves to 4 - Beta per the release plan (1.0 deliberately
deferred until the extension API sees real use), [project.urls] points at
the repository, changelog, and issue tracker, and inputguard.__version__
syncs with the pyproject version. The changelog's Unreleased v0.3 section
is promoted to [0.3.0] with the CI and CLI entries from this PR.

Also amends the wheel-gate assertion to what the zero-dependency promise
actually means: no UNCONDITIONAL Requires-Dist entries (setuptools emits
extra-gated `Requires-Dist: ...; extra == "dev"` lines for the optional
tooling — verified against the built wheel, and the bare-venv install
proves none of them install at runtime).

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
…ob, budgets from measured CI data

The first CI run (35277175054) failed the latency asserts inside the
coverage pass: coverage tracing roughly triples per-call cost (debug
p50 21.14 ms vs 6.59 ms untraced; ready 20.35 ms vs 5.62 ms), and runner
hardware is slower than the dev sandbox. Gating the same budgets under
coverage in the test job AND untraced in the benchmark job was a design
error — two different measurement conditions cannot share one threshold.

- The coverage pass now ignores tests/test_benchmarks.py (with the
  observed numbers in a comment); latency gates belong exclusively to
  the dedicated benchmark job, which runs untraced and reflects
  real-world analysis cost.
- Budgets re-derived from measured environments documented in the test
  module: sandbox untraced (build 21.94 / debug 6.59 / ready 5.62 ms
  p50) and the CI-traced observation above. New p50 budgets: build 65
  ms (~3x), debug 25 ms (~3.8x), ready 25 ms (~4.4x); p95 sustained
  guards at 2.5x unchanged. Still order-of-magnitude gates: they catch
  algorithmic regressions (accidental O(n^2), a lost cap) that blow
  past by 10x+, while absorbing CI hardware variance.

Local verification of the exact split: coverage pass 441 passed, 96.29%
branch (floor 90); benchmark job untraced 4 passed; ruff + mypy --strict
clean.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
test (3.10) flaked on the rebased head (run 35277834661): the dotted-filler
linearization test's absolute bound asserted a single wall-clock sample
under 50 ms, and the coverage-traced 3.10 runner measured 50.61 ms. The
meaningful guard is the 40x growth-rate assertion — the regression it
protects against measured 1.3 s at 10 K, ~26x the bound. Both timing tests
now use a 250 ms absolute bound (5x+ detection margin, no CI-hardware
flake); the growth-rate assertion is unchanged.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
@obvious-autobuild
obvious-autobuild Bot merged commit 850401f into release/v0.3.0 Sep 17, 2026
7 checks passed
@obvious-autobuild
obvious-autobuild Bot deleted the feat/ci-packaging branch September 17, 2026 21:43
obvious-autobuild Bot added a commit that referenced this pull request Sep 18, 2026
…measured FP benchmark, changelog (#14)

* chore: release branch for InputGuard v0.3 production upgrade

* feat: typed Rule protocol, in-process registry, and loud scorer failures (#3)

* feat(rule-protocol-registry): typed Rule protocol and in-process registry

Open the engine (spec art_bTvdPdJS §1): rules and domains become
registered data instead of code paths.

- inputguard/registry.py: runtime-checkable Rule protocol (id, domain,
  severity, gap + check(text, intent)) and RuleRegistry with
  register_rule (decorator and imperative forms) and register_domain
  (priority-ordered intent signals, single empty-terms fallback intent).
  Duplicate rule ids, unknown severities, and — once a domain exists —
  undeclared rule domains raise at registration time.
- detector.py: the coding priority chain becomes INTENT_SIGNALS data;
  detect_intent gains an optional signals parameter so registered
  domains run the same chain; normalize() becomes public so analyze()
  can honor the "rules receive normalized text" contract.

Nothing consumes the registry yet; behavior is unchanged (89 tests
green). Zero new dependencies.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* refactor(built-ins-via-registry): dispatch analyze() through the registry

Replace the hardcoded domain whitelist and the if/elif intent dispatch
with registry lookups (spec art_bTvdPdJS §1):

- analyzer.py: analyze() resolves domain signals via
  REGISTRY.get_domain_signals (unknown domain still raises ValueError,
  now naming every registered domain) and runs rules via
  REGISTRY.rules_for_intent — the deduped, deterministic-finding-order
  contract of the v0.2 runners is preserved by the registry's
  registration-order iteration.
- rules/*: each of the 19 built-in rules gains a thin registry adapter
  class exposing the v0.2 check function through the Rule protocol,
  registered with @register_rule — the built-ins dogfood the exact
  extension path users get. The coding catch-all adapter re-runs the
  built-in build rules internally to preserve its fires-only-when-empty
  semantics under the frozen check(text, intent) contract.
- rules/__init__.py: registers the coding domain (INTENT_SIGNALS chain
  plus all 19 rule adapters) through register_domain; the v0.2 runner
  exports are unchanged.

Behavior-identical: the 89 existing tests pass unmodified.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* fix(scorer-loud-failures): unknown severities raise; unknown gaps keep advice

Close the two silent-failure edges the survey pinned (art_CnghyDxp §2.3,
§2.4) per spec art_bTvdPdJS §6:

- scorer.py: calculate_score validates every finding's severity up front
  and raises ValueError on an unknown one — the silent zero penalty from
  .get(severity, 0) is gone (rank comparisons and penalty lookups go
  through validated paths).
- recommender.py: get_recommendations no longer drops gaps without a
  curated entry; unknown gaps get a documented four-key generic
  fallback (_fallback_recommendation), keeping advice complete for rule
  authors who ship a new gap.

Existing behavior for known severities and gaps is unchanged: the 89
existing tests pass unmodified.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* test(registry-invariants): registry contract and parity tests

New tests/test_registry.py pins the foundation's acceptance points:

- dogfooding: all 19 built-in rules registered through the registry,
  coding domain signals in priority order with a single fallback intent,
  deterministic build-intent rule order;
- validation: duplicate rule id, unknown severity, and unknown domain
  raise ValueError at registration; decorator form registers and
  satisfies the Rule protocol;
- loud failures: unknown severity raises in calculate_score (including
  alongside known severities); unknown gaps get the documented four-key
  fallback, never a silent drop;
- end to end: a custom registered rule changes a result (finding, gap,
  score -15, fallback recommendation); a registered custom domain
  analyzes; duplicate domain registration raises;
- invariants: the registry-walking completeness check (every built-in
  gap has a complete recommendation entry), dispatch parity with the
  v0.2 runners across nine inputs, and 65 parallel analyze() calls
  across 8 threads matching sequential results.

Full suite: 114 passed (89 pre-existing, unmodified).

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

---------

Co-authored-by: Obvious <obvious@obvious.ai>
Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* feat: per-gap follow-up questions engine with additive follow_ups result field (#4)

Per-gap templated clarifying questions (one or two per gap, deduped, gap-ordered), slot fills for function/dataset names, a documented fallback for unknown gaps, the additive AnalysisResult.follow_ups field serialized in to_dict(), and the registry-walking completeness invariant extended to require follow-ups for every built-in gap. 139/139 tests green; existing 114 unmodified.

- feat(follow-up-templates): per-gap clarifying question engine
- feat(result-field): additive follow_ups on AnalysisResult
- test(follow-up-completeness): extended invariant + engine behavior tests

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* feat: multilingual degradation — script probe and explicit degraded path for uncovered scripts (#5)

* feat(script-probe): add unicodedata script-histogram language probe

Pure, thread-safe script classification built on unicodedata only —
each letter's Unicode name mentions its script, so token scanning
classifies scripts with no hardcoded range tables. Classifies the
input's dominant script, a coarse script-derived language guess, and
the English-heuristic coverage band (full / partial / none / unknown).
A deterministic stride sample bounds probe cost on arbitrarily long
input.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* feat(degraded-path): explicit degradation for scripts without heuristic coverage

Wire the script probe into analyze(): uncovered scripts skip the
English-only rules outright, take a 20-point confidence penalty (score
80 — never 'ready' in either mode), and report detected_intent
'undetermined' instead of an unearned fallback. Results carry the three
additive fields detected_language, heuristic_coverage, and
degradation_note; to_dict() includes them additively. Covered scripts
(full/partial coverage) run the unchanged pipeline — partial adds a
note without a penalty.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* test(non-english-honesty): probe-P2 regression, Unicode samples, thread parity

The probe-P2 input (Chinese) must return a degraded result with a
degradation_note — never ready/100. Also covers: degraded shape and
mode behavior, per-script degradation, English parity (probe fields on
covered input, byte-parity scores: 'fix my code' -> 35, repeated build
-> 50), partial-coverage rules-run-with-note, validation order,
additive to_dict() keys, systematic no-crash Unicode samples with the
note-consistency invariant, bounded long-input probing, and 64-worker
parallel analyze() consistency across English/Chinese inputs.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* fix(tests): adapt to_dict key-set expectations to the merged additive contract

PR #4's parity test asserted the exact to_dict key set with follow_ups
as the only additive key; the three language-probe fields are additive
per the v0.3 spec, so the expectation now covers the merged additive
set. The key-order test documents the merged 11-key contract.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

---------

Co-authored-by: Obvious <obvious@obvious.ai>
Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* feat(eval-prompt-set): version the 121-case clarity-evaluation set as eval/ (#7)

* feat(eval-prompt-set): version the 121-case clarity-evaluation set as eval/

Checkpoint the labeled evaluation set (workbook wb_q30ippCW + labeling
guide art_XPvHhPeZ) into the repo: cases.csv (121 rows, QUOTE_ALL/CRLF),
README with the labeling criteria, per-case_type rubric, gap-vocabulary
calibration contract, and a stdlib-only measure_fp.py that runs the
analyzer over every row and reports per-case_type and overall FP/FN rates.
The tool is a measurement, not a gate: it always exits 0 and treats
label/behavior disagreement as the data.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* test(eval-loader): smoke-test the eval dataset mechanics

Three smoke tests: the CSV parses with all nine required columns and
non-empty texts, case ids are unique, and measure_fp.py runs end to end
on a 10-case sample. Deliberately no label-vs-analyzer assertions — that
agreement is what the measurement tool reports over time.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

---------

Co-authored-by: Obvious <obvious@obvious.ai>
Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* fix: enforce registry contract — check(text) signature, unique intents, registration guards (#8)

* fix(registry-contract-spec-b1): revert Rule.check to the spec's check(self, text) signature

The merged foundation (PR #3) dispatches rule.check(text, intent), but the
pinned spec (art_bTvdPdJS section 1) defines check(self, text). The intent
parameter is redundant by construction: analyze() dispatches
REGISTRY.rules_for_intent(detected_intent), which filters on
rule.domain == intent, so every rule invoked already knows the intent as its
own domain member — and all 19 built-in adapters ignore it.

The mismatch is not cosmetic: registration only checks callable(rule.check),
so a rule written per the spec registers cleanly and then crashes analyze()
mid-run with "TypeError: check() takes 2 positional arguments but 3 were
given" (review probe P1). Spec and code must agree before wave-2/3 rule
authors copy either version.

- inputguard/registry.py: Rule protocol back to check(self, text); module
  and member docstrings updated to match.
- inputguard/analyzer.py: dispatch rule.check(normalized).
- inputguard/rules/: all 19 built-in adapters drop the unused intent
  parameter; InsufficientContextRule docstring no longer cites (text, intent).
- tests/test_registry.py: helper and decorator-form test rules updated.

Part of release PR #2. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* fix(registry-intent-scoping-b2): globally-unique intents, registration guards, loud rule attribution

Closes the cross-domain leak and the registration holes the adversarial
review proved against the merged foundation (art_hC18m78C). All guards
fire at registration time, before any mutation.

- B2: register_domain raises ValueError when an intent name is already
  declared by another registered domain, naming both domains. Rules
  dispatch on intent name alone (rules_for_intent filters on
  rule.domain), so a shared intent name ran one domain's rules inside
  the other's analysis (review probe P6b: a legal rule fired inside a
  coding analysis). Duplicate intent names within one signals argument
  raise too — the ownership map must be unambiguous.
- C4: the self-contradictory error ("declares domain 'coding', which no
  registered domain declares as an intent. Registered domains: 'coding'")
  now states the actual constraint and enumerates the valid intent
  names; the Rule protocol documents that domain holds the intent name
  the rule is registered under, never a domain name.
- C3: duplicate rule ids within one register_domain call raise — the
  old code silently dropped the second rule (review probe P2).
- N3: _validate checks the check() signature via inspect.signature
  binding; a rule whose check cannot accept the single positional text
  argument is a registration error, not a mid-analyze TypeError (review
  probe P1's crash shape).
- C2-lite: analyze() wraps rule dispatch and re-raises RuntimeError with
  the rule id and its registration origin ("at <file>:<line>", captured
  at _add time and exposed via RuleRegistry.rule_origin), chaining the
  original exception. Blast-radius contract documented on the Rule
  protocol and RuleRegistry: a rule exception aborts analyze() by
  design, registration is permanent, REGISTRY is process-global.
- C6: the extension API (register_rule, register_domain, REGISTRY, Rule)
  is exported at package top level, additive on the untouched v0.2
  four-export surface.

Part of release PR #2. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* test(registry-contract-guards): pin every remediated registry behavior in the suite

Permanent tests for the review's findings so the gates enforce the chosen
contract (the review found none of these covered):

- spec-compliant check(text) rule registers and fires end to end (P1)
- wrong-arity and zero-argument check rejected at registration, nothing
  enters the registry (N3)
- cross-domain intent collision raises naming both domains (B2/P6b),
  leakage probe cannot recur
- duplicate intent within one signals argument raises (N1 registry half)
- same-call duplicate rule ids raise and leave no partial registration (C3/P2)
- rule exception aborts analyze() with rule id + registration origin and
  the original cause chained (C2-lite/P4)
- extension API importable at top level; v0.2 four-export surface intact (C6)

Suite: 123 passed (114 baseline + 9 new).

Part of release PR #2. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* test: adapt PR #4 follow-ups test rule to the spec check(text) signature

Rebase adaptation: PR #4's unknown-gap test registered a rule with the
retired check(text, intent) signature; the N3 registration guard now
correctly rejects it. The test's subject (unknown-gap fallback follow-up)
is orthogonal to the signature — the rule moves to the spec's
check(self, text). Suite: 202 passed.

Part of release PR #2. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

---------

Co-authored-by: Obvious <obvious@obvious.ai>

* fix: match terms at word boundaries across detector and all rule modules (#6)

* fix(word-boundary-generalization): shared word-boundary matcher across detector and all rule modules

Raw substring containment (detector.py:57-58) read "fixture" as the debug
signal "fix" (probe P1) and flagged innocent input at 35/needs_clarification.
All term lookups now route through inputguard/matching.py, which matches
terms as standalone words using the boundary class the coding matcher
already used, extended with inflectional endings so suffix-shaped hits
(bugs, errors, debugged, refactoring, profiling, stack traces) keep firing.
Coding rules keep their v0.2 multiword token fallback via token_fallback=True;
other callers get strict adjacent-phrase matching. No-alnum terms (/, =>, ->)
keep substring semantics. Finders compile once per term-set (lru_cache).

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* test(fp-regressions): pin word-boundary matcher contract and probe-P1 fix

Regression coverage for the shared matcher: the P1 "fixture" input no
longer reads as debug and is not hard-flagged; embedded words (fixture,
refix, prefix, praised) stop matching; inflected forms (bugs, errors,
debugged, refactoring, profiling, stack traces) keep firing so the
boundary tightening adds no false negatives; phrases match adjacently
with final-word inflections; the coding token fallback stays opt-in;
symbol terms keep substring semantics.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

---------

Co-authored-by: Obvious <obvious@obvious.ai>
Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* feat(policy): frozen validated Policy with pinned v0.2 defaults, banding, input cap, and score breakdown (#9)

* feat(policy-object): frozen validated Policy with rule filters and word floor

Every scoring tunable moves into a frozen, validated Policy dataclass whose
defaults are the v0.2 constants: ready_at 85, usable_at 60, strict_clarify_at
65, penalties 5/15/25, min_words 3, max_chars 10000, borderline_at 74.
Validation rejects mis-ordered bands, out-of-range values, uncompilable
allow_patterns, and disabled rule ids that are not registered. InputGuard and
analyze() take an optional policy (a per-call policy overrides the guard's).

disabled_rules skips registered rules by id; allow_patterns regex matches skip
flagging entirely; min_words gates the insufficient_context vague rule below
its floor. Penalties flow from Policy through calculate_score(findings,
policy=None).

Closes review concern C5 (art_hC18m78C): one severity vocabulary
(policy.SEVERITIES) shared by registration and a new single choke point,
scorer.require_known_severity, which the analyzer applies to every emitted
finding; registry.KNOWN_SEVERITIES is cross-pinned to it by test.

Zero new dependencies; stdlib only.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* feat(two-layer-banding): policy-driven status bands and a borderline near-miss signal

get_status(score, mode, policy=None) now reads its bands from Policy
(ready_at 85, usable_at 60, strict_clarify_at 65 — the v0.2 constants, pinned
byte-exactly in tests/test_banding.py), replacing the four hard-coded scorer
thresholds. Severity still decides what fires; bands decide what happens.

AnalysisResult gains an additive borderline: bool field (default False, in
to_dict) set when the score lands in [policy.borderline_at, ready_at) — the
distinct near-miss "worth one more pass" signal from the spec's result states.
interpretation_note is untouched, so existing serialized results are
byte-identical under default policy.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* feat(input-cap): enforce Policy.max_chars with visible truncation

analyze() now slices input to policy.max_chars before intent detection and
rule execution — analysis cost is bounded regardless of input size (v0.2 took
~3.8 s on a 1.6 MB input; capped analysis of the same input runs in
milliseconds). Truncation is never silent: AnalysisResult gains an additive
truncated: bool field (default False, in to_dict), and the allowlist is
evaluated against the same capped input so no scan escapes the budget.

Default cap 10_000 matches the Policy default, pinned by test.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* feat(score-breakdown): additive score_breakdown audit field on the result

calculate_score_with_breakdown(findings, policy=None) scores findings and
returns the audit trail — {"base": 100, "penalties": [{"code", "severity",
"points"(negative)}...], "final"} — one entry per distinct gap (or code when
gap is None) in first-occurrence order, matching the result's gaps order.
calculate_score keeps its v0.2 signature and delegates to it, so both stay in
lockstep by construction.

AnalysisResult gains an additive score_breakdown field (default None, deep-
copied in to_dict so a serialized snapshot can never be rewritten by later
mutation). With default policy the penalties mirror the v0.2 deductions
exactly, so serialized defaults differ only by the new keys.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* test(pin-v02-defaults): pin every Policy default byte-exactly to v0.2

Dedicated pinning tests guard the calibration-drift risk the spec flags:
each Policy default (ready_at 85, usable_at 60, strict_clarify_at 65,
penalties 5/15/25, min_words 3, max_chars 10000, borderline_at 74) is
asserted three ways — dataclass field defaults, constructed instance values,
and the literal source line (underscore separators normalized) — plus a
no-unpinned-fields guard so future fields cannot ship uncalibrated.

Behavioral pins reproduce the observed v0.2 reference outputs under default
policy (75/usable_with_warnings, 35/needs_clarification, 50, strict
blocked/needs_clarification) and the v0.2 result and recommendation key
shapes.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* test: integrate policy keys into sibling additive-contract tests

The followups parity test and the language ordered-key-list test pin the
exhaustive to_dict key set as of their own merge; the policy-calibration
fields (borderline, truncated, score_breakdown) are additive per the compat
contract, so the expected sets extend rather than the fields disappearing.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

---------

Co-authored-by: Obvious <obvious@obvious.ai>
Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* feat: first-party writing domain — six rules, compose intent, complete advice (#10)

* feat(writing-domain): first-party writing domain — six rules, compose intent, complete advice

Register the writing domain through the same registry path user domains
take: one globally unique fallback intent ("compose") and six rules
covering the eval vocabulary's six gaps — audience, purpose,
structure/format, source material, context, completeness — at the
vocabulary-pinned severities (high/high/medium/high/medium/low).

All term matching goes through the shared word-boundary matcher
(inputguard.matching, PR #6). Rules follow the trigger-and-satisfy shape;
the source-material rule fires only when existing text is referenced but
not provided. Recommendations and follow-up entries ship for all six gaps
in the same change, keeping the registry completeness invariant whole.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* test(writing-domain): per-rule, completeness, integration, no-leakage, and eval-label coverage

Five layers: per-rule fire/silent behavior; the writing-domain
completeness invariant (every writing gap has a four-key recommendation
and one or two follow-up questions); analyze() integration end to end
(underspecified -> needs_clarification at score 30, fully specified ->
ready, partially specified -> usable_with_warnings); no-leakage in both
directions at registry and analyze level; and the 14 labeled English
writing rows from eval/cases.csv pinned as a calibration net.

The exact-set registry expectations extend from 19 coding rules to 25
rules across two domains — the guards are strengthened, not weakened.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

---------

Co-authored-by: Obvious <obvious@obvious.ai>
Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* feat: first-party data-analysis domain — six rules, analysis intent, complete advice (#11)

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* fix: degradation coverage for DG-011..014 (detector) (#12)

Function-word layer (fr/es/pt stop-word lexicons with English margin) plus uncovered-run layer (3+ consecutive uncovered-script letters) route accented-Latin and mixed English+Han input to the explicit degraded path; DG-012's silent ready and DG-011/013/014's unnoted rule findings are gone. Degradation eval FP 3/14 -> 0/14, note honesty 10/14 -> 14/14, all 121 probe verdicts otherwise byte-identical.

* fix: complete v0.3.0 release sweep — version, cap, degradation, docs, CI

- Bump pyproject.toml and inputguard.__version__ to 0.3.0; write the full
  0.3.0 CHANGELOG section from the merged release-wave history.
- Report the literal "degraded" status on the degraded path in both modes;
  add the mixed-script (>=5% uncovered remainder) and non-English-Latin
  (<10% English evidence, filler-resistant denominator) gates so French/
  Spanish/Portuguese and mixed English+Han prompts degrade with zero gaps
  instead of spurious ones; preserve the truncated flag on that path.
- Linear dataset-filename extraction (follow-ups + data-analysis rule)
  with 1K/10K timing tests; benchmark doc records the before/after.
- README: rewritten Non-English section, full result-object field table
  (follow_ups, probe fields, borderline, truncated, score_breakdown),
  writing and data-analysis rule tables.
- eval: docs/false-positive-benchmark.md from measured results.
- CI: minimal GitHub Actions workflow running pytest on push/PR (3.9+3.13).

Eval: 116/121 matching (all 14 degradation rows, notes on 14/14); the 5
residual mismatches are pre-existing coding boundary/intent rows.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
Human author: Kalisetti Nihanth Naidu (nihanthnaidu007@gmail.com)

* feat: CI matrix, latency budgets, CLI, and packaging metadata for v0.3 (#13)

* test(registry-contract-guards): CI-enforced guards for the three review findings

Dedicated suite for the adversarial-API-review gaps (art_hC18m78C), as the
named acceptance artifact:

- C3/P2: duplicate rule ids inside one register_domain call raise
  ValueError naming the id; a rejected batch mutates nothing.
- B1/N3/P1: a rule whose check cannot be called as check(self, text)
  (extra param, zero param, keyword-only) is rejected at registration with
  an actionable message; defaulted and spec signatures accepted and fired.
- C2/P4: an exception inside check() aborts analyze() as a RuntimeError
  carrying the rule id, its registration origin, and the original
  exception chained -- via the register_domain path (register_rule is
  covered in test_registry.py).

The registry-isolation fixture moves to tests/conftest.py so the guard
suite and the policy suite share one snapshot/restore implementation
instead of importing across test modules.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* build(typing): annotate for mypy --strict without behavior change

Pre-work for the CI strict-typing gate: missing Iterable[str] on the five
rule modules' _contains_any helpers, Dict[str, Any] / Dict[str, str] generics
on types.py and recommender.py, a pinned Rule-typed local in
registry._instantiate, and the writing signals Tuple[str, ...] type argument.
Also removes a literal duplicate "filter by"/"sort by" set entry in the
feature scope signals (ruff B033); set semantics make this a no-op at
runtime, and rules/__init__ marks its v0.2 signal re-exports as intentional
via __all__ (ruff F401).

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* ci(actions-matrix): GitHub Actions CI with lint, typed, tested, benchmark, wheel gates

The v0.2 failure mode was "every change ships on trust" — no CI, lint,
typing, or coverage tooling at all. This adds the production gate set from
the spec (§8) as three jobs:

- test: Python 3.9-3.13 matrix running ruff, pytest with a 90%
  branch-coverage floor (--cov-branch --cov-fail-under=90 --strict), and
  mypy --strict (finally checking the shipped py.typed).
- benchmark: latency budgets as merge gates, not notes. Budgets are derived
  from measured numbers (build-intent p50 21.94 ms at 10k chars including
  the C1 InsufficientContextRule double-run, debug 6.59 ms, ready 5.62 ms;
  1.6 MB input bounded by the Policy cap at 17.1 ms vs the v0.2 unbounded
  3.8 s scan). Hard gate is p50 with CI-noise headroom; p50/p99 print to
  the job log per spec. Budgets and derivation documented in
  tests/test_benchmarks.py.
- wheel-zero-dep: the zero-dependency promise verified against the built
  artifact — wheel METADATA carries no Requires-Dist, a bare venv install
  pulls no third-party package, and the public API works from the
  installed wheel. Smoke steps run from a scratch directory so the repo's
  local inputguard/ package can never shadow the installed wheel.

pyproject gains the matching tool config ([tool.ruff] E/F/W/B with E501
ignored, [tool.mypy] strict, explicit pytest testpaths) and the expanded
[dev] extras the CI jobs install; all dev tooling stays out of runtime
dependencies.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* feat(cli): inputguard analyze with argparse, to_dict JSON, CI-friendly exit codes

The developer surface the adoption research flagged (spec §5): a
zero-dependency argparse CLI shipped as a console script.

- "inputguard analyze TEXT [--stdin] [--domain D] [--mode M]
  [--format text|json] [--min-score N]"; JSON output is the result's own
  to_dict() contract, byte-equal to the library's (asserted by test).
- Exit codes documented in the module docstring and --help: 0 = analysis
  ok and at/above the --min-score floor when given; 1 = below the floor
  (the commit-hook / CI gate); 2 = usage error (missing text, unknown
  domain) reported cleanly on stderr, never a traceback.
- Text rendering follows the spec §5 example shape: status, clarity,
  intent, domain, missing gaps, ask lines, and the degradation note when
  the language probe marks the result degraded (probe P2 input shows the
  note, never a silent ready).

Entry point added in pyproject [project.scripts]; stdlib argparse and json
only — the zero-dependency promise is untouched.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* chore(packaging-metadata): 0.3.0, Beta status, 3.13 classifier, project URLs

The packaging drift the survey flagged is closed: classifiers extend to
Python 3.13 (the runtime the project is developed and tested on),
Development Status moves to 4 - Beta per the release plan (1.0 deliberately
deferred until the extension API sees real use), [project.urls] points at
the repository, changelog, and issue tracker, and inputguard.__version__
syncs with the pyproject version. The changelog's Unreleased v0.3 section
is promoted to [0.3.0] with the CI and CLI entries from this PR.

Also amends the wheel-gate assertion to what the zero-dependency promise
actually means: no UNCONDITIONAL Requires-Dist entries (setuptools emits
extra-gated `Requires-Dist: ...; extra == "dev"` lines for the optional
tooling — verified against the built wheel, and the bare-venv install
proves none of them install at runtime).

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* fix(benchmark-budgets): gate latency only in the untraced benchmark job, budgets from measured CI data

The first CI run (35277175054) failed the latency asserts inside the
coverage pass: coverage tracing roughly triples per-call cost (debug
p50 21.14 ms vs 6.59 ms untraced; ready 20.35 ms vs 5.62 ms), and runner
hardware is slower than the dev sandbox. Gating the same budgets under
coverage in the test job AND untraced in the benchmark job was a design
error — two different measurement conditions cannot share one threshold.

- The coverage pass now ignores tests/test_benchmarks.py (with the
  observed numbers in a comment); latency gates belong exclusively to
  the dedicated benchmark job, which runs untraced and reflects
  real-world analysis cost.
- Budgets re-derived from measured environments documented in the test
  module: sandbox untraced (build 21.94 / debug 6.59 / ready 5.62 ms
  p50) and the CI-traced observation above. New p50 budgets: build 65
  ms (~3x), debug 25 ms (~3.8x), ready 25 ms (~4.4x); p95 sustained
  guards at 2.5x unchanged. Still order-of-magnitude gates: they catch
  algorithmic regressions (accidental O(n^2), a lost cap) that blow
  past by 10x+, while absorbing CI hardware variance.

Local verification of the exact split: coverage pass 441 passed, 96.29%
branch (floor 90); benchmark job untraced 4 passed; ruff + mypy --strict
clean.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* fix: give the extraction linearization bounds CI headroom

test (3.10) flaked on the rebased head (run 35277834661): the dotted-filler
linearization test's absolute bound asserted a single wall-clock sample
under 50 ms, and the coverage-traced 3.10 runner measured 50.61 ms. The
meaningful guard is the 40x growth-rate assertion — the regression it
protects against measured 1.3 s at 10 K, ~26x the bound. Both timing tests
now use a 250 ms absolute bound (5x+ detection margin, no CI-hardware
flake); the growth-rate assertion is unchanged.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

---------

Co-authored-by: Obvious <obvious@obvious.ai>
Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* docs(readme-tested-examples): rewrite README for v0.3 with every example executed by tests

The v0.2 README drifted from recommender.py within two releases. The
rewritten README documents the v0.3 surface (extension API, Policy,
follow-ups, degradation, CLI) and every example now runs: a new
tests/test_readme_examples.py executes all python blocks in document
order in a subprocess, diffs the embedded to_dict() JSON against live
analyzer output, and checks exports, rule tables, gap vocabulary, and
CLI output against the live registry. Example values carry real
assertions sourced from runtime probes.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* docs(integration-recipes): LangChain and LiteLLM recipes consuming to_dict()

Copy-paste integration patterns over the to_dict() serialization
boundary — preflight-and-gate and clarify-then-complete. The recipes
are documentation only: framework packages stay out of core and out of
the dependency tree. tests/test_recipes.py executes every recipe block
in document order, stubbing the single framework idiom each wiring
block uses, so the inputguard-side contract cannot drift silently.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* docs(fp-benchmark): refresh measured rates and state the boundary target honestly

Re-ran eval/measure_fp.py on the release branch: overall 116/121 match,
FP 3.3% / FN 0.8%, zero false positives on true negatives and
degradation rows, notes on 14/14 degradation rows, PF rows 19-20 ms.
Adds the labeling-guide target contrast the doc lacked: fixture-style
boundary FP measured 4/6 (v0.2 baseline 6/6, target 0/6) — driven by
the matcher's deliberate inflection tolerance, documented as an
eval-driven trade-off for a future release, not a label edit.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* chore(release-0.3.0): complete the 0.3.0 changelog for the docs wave

Version 0.3.0 and the Development Status :: 4 - Beta classifier were
already in pyproject.toml and inputguard.__version__ — verified, no
bump needed. This commit completes the 0.3.0 entry with the docs wave:
tested README rewrite, integration recipes, and the measured
false-positive benchmark with the boundary-target contrast.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

* fix(readme-examples): 3.9-compatible annotation in the custom-rule example

The custom-rule example annotated check() with PEP 604 union syntax
(-> RuleFinding | None), which is evaluated at definition time and
raises TypeError on Python 3.9 — caught by the CI matrix's 3.9 job
running the README anti-drift gate. Use typing.Optional instead; an
AST scan of every executed README block confirms no other
runtime-evaluated 3.9-incompatible construct remains.

Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>

---------

Co-authored-by: Obvious <obvious@obvious.ai>
Co-authored-by: obvious-autobuild[bot] <262744130+obvious-autobuild[bot]@users.noreply.github.com>
Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants