Repository navigation
feat: CI matrix, latency budgets, CLI, and packaging metadata for v0.3 - #13
Merged
Merged
Conversation
…ew findings Dedicated suite for the adversarial-API-review gaps (art_hC18m78C), as the named acceptance artifact: - C3/P2: duplicate rule ids inside one register_domain call raise ValueError naming the id; a rejected batch mutates nothing. - B1/N3/P1: a rule whose check cannot be called as check(self, text) (extra param, zero param, keyword-only) is rejected at registration with an actionable message; defaulted and spec signatures accepted and fired. - C2/P4: an exception inside check() aborts analyze() as a RuntimeError carrying the rule id, its registration origin, and the original exception chained -- via the register_domain path (register_rule is covered in test_registry.py). The registry-isolation fixture moves to tests/conftest.py so the guard suite and the policy suite share one snapshot/restore implementation instead of importing across test modules. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
Pre-work for the CI strict-typing gate: missing Iterable[str] on the five rule modules' _contains_any helpers, Dict[str, Any] / Dict[str, str] generics on types.py and recommender.py, a pinned Rule-typed local in registry._instantiate, and the writing signals Tuple[str, ...] type argument. Also removes a literal duplicate "filter by"/"sort by" set entry in the feature scope signals (ruff B033); set semantics make this a no-op at runtime, and rules/__init__ marks its v0.2 signal re-exports as intentional via __all__ (ruff F401). Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
…mark, wheel gates The v0.2 failure mode was "every change ships on trust" — no CI, lint, typing, or coverage tooling at all. This adds the production gate set from the spec (§8) as three jobs: - test: Python 3.9-3.13 matrix running ruff, pytest with a 90% branch-coverage floor (--cov-branch --cov-fail-under=90 --strict), and mypy --strict (finally checking the shipped py.typed). - benchmark: latency budgets as merge gates, not notes. Budgets are derived from measured numbers (build-intent p50 21.94 ms at 10k chars including the C1 InsufficientContextRule double-run, debug 6.59 ms, ready 5.62 ms; 1.6 MB input bounded by the Policy cap at 17.1 ms vs the v0.2 unbounded 3.8 s scan). Hard gate is p50 with CI-noise headroom; p50/p99 print to the job log per spec. Budgets and derivation documented in tests/test_benchmarks.py. - wheel-zero-dep: the zero-dependency promise verified against the built artifact — wheel METADATA carries no Requires-Dist, a bare venv install pulls no third-party package, and the public API works from the installed wheel. Smoke steps run from a scratch directory so the repo's local inputguard/ package can never shadow the installed wheel. pyproject gains the matching tool config ([tool.ruff] E/F/W/B with E501 ignored, [tool.mypy] strict, explicit pytest testpaths) and the expanded [dev] extras the CI jobs install; all dev tooling stays out of runtime dependencies. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
…y exit codes The developer surface the adoption research flagged (spec §5): a zero-dependency argparse CLI shipped as a console script. - "inputguard analyze TEXT [--stdin] [--domain D] [--mode M] [--format text|json] [--min-score N]"; JSON output is the result's own to_dict() contract, byte-equal to the library's (asserted by test). - Exit codes documented in the module docstring and --help: 0 = analysis ok and at/above the --min-score floor when given; 1 = below the floor (the commit-hook / CI gate); 2 = usage error (missing text, unknown domain) reported cleanly on stderr, never a traceback. - Text rendering follows the spec §5 example shape: status, clarity, intent, domain, missing gaps, ask lines, and the degradation note when the language probe marks the result degraded (probe P2 input shows the note, never a silent ready). Entry point added in pyproject [project.scripts]; stdlib argparse and json only — the zero-dependency promise is untouched. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
…ct URLs The packaging drift the survey flagged is closed: classifiers extend to Python 3.13 (the runtime the project is developed and tested on), Development Status moves to 4 - Beta per the release plan (1.0 deliberately deferred until the extension API sees real use), [project.urls] points at the repository, changelog, and issue tracker, and inputguard.__version__ syncs with the pyproject version. The changelog's Unreleased v0.3 section is promoted to [0.3.0] with the CI and CLI entries from this PR. Also amends the wheel-gate assertion to what the zero-dependency promise actually means: no UNCONDITIONAL Requires-Dist entries (setuptools emits extra-gated `Requires-Dist: ...; extra == "dev"` lines for the optional tooling — verified against the built wheel, and the bare-venv install proves none of them install at runtime). Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
…ob, budgets from measured CI data The first CI run (35277175054) failed the latency asserts inside the coverage pass: coverage tracing roughly triples per-call cost (debug p50 21.14 ms vs 6.59 ms untraced; ready 20.35 ms vs 5.62 ms), and runner hardware is slower than the dev sandbox. Gating the same budgets under coverage in the test job AND untraced in the benchmark job was a design error — two different measurement conditions cannot share one threshold. - The coverage pass now ignores tests/test_benchmarks.py (with the observed numbers in a comment); latency gates belong exclusively to the dedicated benchmark job, which runs untraced and reflects real-world analysis cost. - Budgets re-derived from measured environments documented in the test module: sandbox untraced (build 21.94 / debug 6.59 / ready 5.62 ms p50) and the CI-traced observation above. New p50 budgets: build 65 ms (~3x), debug 25 ms (~3.8x), ready 25 ms (~4.4x); p95 sustained guards at 2.5x unchanged. Still order-of-magnitude gates: they catch algorithmic regressions (accidental O(n^2), a lost cap) that blow past by 10x+, while absorbing CI hardware variance. Local verification of the exact split: coverage pass 441 passed, 96.29% branch (floor 90); benchmark job untraced 4 passed; ruff + mypy --strict clean. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
obvious-autobuild
Bot
force-pushed
the
feat/ci-packaging
branch
from
September 17, 2026 21:38
96581e8 to
d5f97af
Compare
test (3.10) flaked on the rebased head (run 35277834661): the dotted-filler linearization test's absolute bound asserted a single wall-clock sample under 50 ms, and the coverage-traced 3.10 runner measured 50.61 ms. The meaningful guard is the 40x growth-rate assertion — the regression it protects against measured 1.3 s at 10 K, ~26x the bound. Both timing tests now use a 250 ms absolute bound (5x+ detection margin, no CI-hardware flake); the growth-rate assertion is unchanged. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
This was referenced Sep 17, 2026
obvious-autobuild Bot
added a commit
that referenced
this pull request
Sep 18, 2026
…measured FP benchmark, changelog (#14) * chore: release branch for InputGuard v0.3 production upgrade * feat: typed Rule protocol, in-process registry, and loud scorer failures (#3) * feat(rule-protocol-registry): typed Rule protocol and in-process registry Open the engine (spec art_bTvdPdJS §1): rules and domains become registered data instead of code paths. - inputguard/registry.py: runtime-checkable Rule protocol (id, domain, severity, gap + check(text, intent)) and RuleRegistry with register_rule (decorator and imperative forms) and register_domain (priority-ordered intent signals, single empty-terms fallback intent). Duplicate rule ids, unknown severities, and — once a domain exists — undeclared rule domains raise at registration time. - detector.py: the coding priority chain becomes INTENT_SIGNALS data; detect_intent gains an optional signals parameter so registered domains run the same chain; normalize() becomes public so analyze() can honor the "rules receive normalized text" contract. Nothing consumes the registry yet; behavior is unchanged (89 tests green). Zero new dependencies. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * refactor(built-ins-via-registry): dispatch analyze() through the registry Replace the hardcoded domain whitelist and the if/elif intent dispatch with registry lookups (spec art_bTvdPdJS §1): - analyzer.py: analyze() resolves domain signals via REGISTRY.get_domain_signals (unknown domain still raises ValueError, now naming every registered domain) and runs rules via REGISTRY.rules_for_intent — the deduped, deterministic-finding-order contract of the v0.2 runners is preserved by the registry's registration-order iteration. - rules/*: each of the 19 built-in rules gains a thin registry adapter class exposing the v0.2 check function through the Rule protocol, registered with @register_rule — the built-ins dogfood the exact extension path users get. The coding catch-all adapter re-runs the built-in build rules internally to preserve its fires-only-when-empty semantics under the frozen check(text, intent) contract. - rules/__init__.py: registers the coding domain (INTENT_SIGNALS chain plus all 19 rule adapters) through register_domain; the v0.2 runner exports are unchanged. Behavior-identical: the 89 existing tests pass unmodified. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * fix(scorer-loud-failures): unknown severities raise; unknown gaps keep advice Close the two silent-failure edges the survey pinned (art_CnghyDxp §2.3, §2.4) per spec art_bTvdPdJS §6: - scorer.py: calculate_score validates every finding's severity up front and raises ValueError on an unknown one — the silent zero penalty from .get(severity, 0) is gone (rank comparisons and penalty lookups go through validated paths). - recommender.py: get_recommendations no longer drops gaps without a curated entry; unknown gaps get a documented four-key generic fallback (_fallback_recommendation), keeping advice complete for rule authors who ship a new gap. Existing behavior for known severities and gaps is unchanged: the 89 existing tests pass unmodified. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * test(registry-invariants): registry contract and parity tests New tests/test_registry.py pins the foundation's acceptance points: - dogfooding: all 19 built-in rules registered through the registry, coding domain signals in priority order with a single fallback intent, deterministic build-intent rule order; - validation: duplicate rule id, unknown severity, and unknown domain raise ValueError at registration; decorator form registers and satisfies the Rule protocol; - loud failures: unknown severity raises in calculate_score (including alongside known severities); unknown gaps get the documented four-key fallback, never a silent drop; - end to end: a custom registered rule changes a result (finding, gap, score -15, fallback recommendation); a registered custom domain analyzes; duplicate domain registration raises; - invariants: the registry-walking completeness check (every built-in gap has a complete recommendation entry), dispatch parity with the v0.2 runners across nine inputs, and 65 parallel analyze() calls across 8 threads matching sequential results. Full suite: 114 passed (89 pre-existing, unmodified). Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> --------- Co-authored-by: Obvious <obvious@obvious.ai> Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * feat: per-gap follow-up questions engine with additive follow_ups result field (#4) Per-gap templated clarifying questions (one or two per gap, deduped, gap-ordered), slot fills for function/dataset names, a documented fallback for unknown gaps, the additive AnalysisResult.follow_ups field serialized in to_dict(), and the registry-walking completeness invariant extended to require follow-ups for every built-in gap. 139/139 tests green; existing 114 unmodified. - feat(follow-up-templates): per-gap clarifying question engine - feat(result-field): additive follow_ups on AnalysisResult - test(follow-up-completeness): extended invariant + engine behavior tests Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * feat: multilingual degradation — script probe and explicit degraded path for uncovered scripts (#5) * feat(script-probe): add unicodedata script-histogram language probe Pure, thread-safe script classification built on unicodedata only — each letter's Unicode name mentions its script, so token scanning classifies scripts with no hardcoded range tables. Classifies the input's dominant script, a coarse script-derived language guess, and the English-heuristic coverage band (full / partial / none / unknown). A deterministic stride sample bounds probe cost on arbitrarily long input. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * feat(degraded-path): explicit degradation for scripts without heuristic coverage Wire the script probe into analyze(): uncovered scripts skip the English-only rules outright, take a 20-point confidence penalty (score 80 — never 'ready' in either mode), and report detected_intent 'undetermined' instead of an unearned fallback. Results carry the three additive fields detected_language, heuristic_coverage, and degradation_note; to_dict() includes them additively. Covered scripts (full/partial coverage) run the unchanged pipeline — partial adds a note without a penalty. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * test(non-english-honesty): probe-P2 regression, Unicode samples, thread parity The probe-P2 input (Chinese) must return a degraded result with a degradation_note — never ready/100. Also covers: degraded shape and mode behavior, per-script degradation, English parity (probe fields on covered input, byte-parity scores: 'fix my code' -> 35, repeated build -> 50), partial-coverage rules-run-with-note, validation order, additive to_dict() keys, systematic no-crash Unicode samples with the note-consistency invariant, bounded long-input probing, and 64-worker parallel analyze() consistency across English/Chinese inputs. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * fix(tests): adapt to_dict key-set expectations to the merged additive contract PR #4's parity test asserted the exact to_dict key set with follow_ups as the only additive key; the three language-probe fields are additive per the v0.3 spec, so the expectation now covers the merged additive set. The key-order test documents the merged 11-key contract. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> --------- Co-authored-by: Obvious <obvious@obvious.ai> Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * feat(eval-prompt-set): version the 121-case clarity-evaluation set as eval/ (#7) * feat(eval-prompt-set): version the 121-case clarity-evaluation set as eval/ Checkpoint the labeled evaluation set (workbook wb_q30ippCW + labeling guide art_XPvHhPeZ) into the repo: cases.csv (121 rows, QUOTE_ALL/CRLF), README with the labeling criteria, per-case_type rubric, gap-vocabulary calibration contract, and a stdlib-only measure_fp.py that runs the analyzer over every row and reports per-case_type and overall FP/FN rates. The tool is a measurement, not a gate: it always exits 0 and treats label/behavior disagreement as the data. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * test(eval-loader): smoke-test the eval dataset mechanics Three smoke tests: the CSV parses with all nine required columns and non-empty texts, case ids are unique, and measure_fp.py runs end to end on a 10-case sample. Deliberately no label-vs-analyzer assertions — that agreement is what the measurement tool reports over time. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> --------- Co-authored-by: Obvious <obvious@obvious.ai> Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * fix: enforce registry contract — check(text) signature, unique intents, registration guards (#8) * fix(registry-contract-spec-b1): revert Rule.check to the spec's check(self, text) signature The merged foundation (PR #3) dispatches rule.check(text, intent), but the pinned spec (art_bTvdPdJS section 1) defines check(self, text). The intent parameter is redundant by construction: analyze() dispatches REGISTRY.rules_for_intent(detected_intent), which filters on rule.domain == intent, so every rule invoked already knows the intent as its own domain member — and all 19 built-in adapters ignore it. The mismatch is not cosmetic: registration only checks callable(rule.check), so a rule written per the spec registers cleanly and then crashes analyze() mid-run with "TypeError: check() takes 2 positional arguments but 3 were given" (review probe P1). Spec and code must agree before wave-2/3 rule authors copy either version. - inputguard/registry.py: Rule protocol back to check(self, text); module and member docstrings updated to match. - inputguard/analyzer.py: dispatch rule.check(normalized). - inputguard/rules/: all 19 built-in adapters drop the unused intent parameter; InsufficientContextRule docstring no longer cites (text, intent). - tests/test_registry.py: helper and decorator-form test rules updated. Part of release PR #2. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * fix(registry-intent-scoping-b2): globally-unique intents, registration guards, loud rule attribution Closes the cross-domain leak and the registration holes the adversarial review proved against the merged foundation (art_hC18m78C). All guards fire at registration time, before any mutation. - B2: register_domain raises ValueError when an intent name is already declared by another registered domain, naming both domains. Rules dispatch on intent name alone (rules_for_intent filters on rule.domain), so a shared intent name ran one domain's rules inside the other's analysis (review probe P6b: a legal rule fired inside a coding analysis). Duplicate intent names within one signals argument raise too — the ownership map must be unambiguous. - C4: the self-contradictory error ("declares domain 'coding', which no registered domain declares as an intent. Registered domains: 'coding'") now states the actual constraint and enumerates the valid intent names; the Rule protocol documents that domain holds the intent name the rule is registered under, never a domain name. - C3: duplicate rule ids within one register_domain call raise — the old code silently dropped the second rule (review probe P2). - N3: _validate checks the check() signature via inspect.signature binding; a rule whose check cannot accept the single positional text argument is a registration error, not a mid-analyze TypeError (review probe P1's crash shape). - C2-lite: analyze() wraps rule dispatch and re-raises RuntimeError with the rule id and its registration origin ("at <file>:<line>", captured at _add time and exposed via RuleRegistry.rule_origin), chaining the original exception. Blast-radius contract documented on the Rule protocol and RuleRegistry: a rule exception aborts analyze() by design, registration is permanent, REGISTRY is process-global. - C6: the extension API (register_rule, register_domain, REGISTRY, Rule) is exported at package top level, additive on the untouched v0.2 four-export surface. Part of release PR #2. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * test(registry-contract-guards): pin every remediated registry behavior in the suite Permanent tests for the review's findings so the gates enforce the chosen contract (the review found none of these covered): - spec-compliant check(text) rule registers and fires end to end (P1) - wrong-arity and zero-argument check rejected at registration, nothing enters the registry (N3) - cross-domain intent collision raises naming both domains (B2/P6b), leakage probe cannot recur - duplicate intent within one signals argument raises (N1 registry half) - same-call duplicate rule ids raise and leave no partial registration (C3/P2) - rule exception aborts analyze() with rule id + registration origin and the original cause chained (C2-lite/P4) - extension API importable at top level; v0.2 four-export surface intact (C6) Suite: 123 passed (114 baseline + 9 new). Part of release PR #2. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * test: adapt PR #4 follow-ups test rule to the spec check(text) signature Rebase adaptation: PR #4's unknown-gap test registered a rule with the retired check(text, intent) signature; the N3 registration guard now correctly rejects it. The test's subject (unknown-gap fallback follow-up) is orthogonal to the signature — the rule moves to the spec's check(self, text). Suite: 202 passed. Part of release PR #2. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> --------- Co-authored-by: Obvious <obvious@obvious.ai> * fix: match terms at word boundaries across detector and all rule modules (#6) * fix(word-boundary-generalization): shared word-boundary matcher across detector and all rule modules Raw substring containment (detector.py:57-58) read "fixture" as the debug signal "fix" (probe P1) and flagged innocent input at 35/needs_clarification. All term lookups now route through inputguard/matching.py, which matches terms as standalone words using the boundary class the coding matcher already used, extended with inflectional endings so suffix-shaped hits (bugs, errors, debugged, refactoring, profiling, stack traces) keep firing. Coding rules keep their v0.2 multiword token fallback via token_fallback=True; other callers get strict adjacent-phrase matching. No-alnum terms (/, =>, ->) keep substring semantics. Finders compile once per term-set (lru_cache). Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * test(fp-regressions): pin word-boundary matcher contract and probe-P1 fix Regression coverage for the shared matcher: the P1 "fixture" input no longer reads as debug and is not hard-flagged; embedded words (fixture, refix, prefix, praised) stop matching; inflected forms (bugs, errors, debugged, refactoring, profiling, stack traces) keep firing so the boundary tightening adds no false negatives; phrases match adjacently with final-word inflections; the coding token fallback stays opt-in; symbol terms keep substring semantics. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> --------- Co-authored-by: Obvious <obvious@obvious.ai> Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * feat(policy): frozen validated Policy with pinned v0.2 defaults, banding, input cap, and score breakdown (#9) * feat(policy-object): frozen validated Policy with rule filters and word floor Every scoring tunable moves into a frozen, validated Policy dataclass whose defaults are the v0.2 constants: ready_at 85, usable_at 60, strict_clarify_at 65, penalties 5/15/25, min_words 3, max_chars 10000, borderline_at 74. Validation rejects mis-ordered bands, out-of-range values, uncompilable allow_patterns, and disabled rule ids that are not registered. InputGuard and analyze() take an optional policy (a per-call policy overrides the guard's). disabled_rules skips registered rules by id; allow_patterns regex matches skip flagging entirely; min_words gates the insufficient_context vague rule below its floor. Penalties flow from Policy through calculate_score(findings, policy=None). Closes review concern C5 (art_hC18m78C): one severity vocabulary (policy.SEVERITIES) shared by registration and a new single choke point, scorer.require_known_severity, which the analyzer applies to every emitted finding; registry.KNOWN_SEVERITIES is cross-pinned to it by test. Zero new dependencies; stdlib only. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * feat(two-layer-banding): policy-driven status bands and a borderline near-miss signal get_status(score, mode, policy=None) now reads its bands from Policy (ready_at 85, usable_at 60, strict_clarify_at 65 — the v0.2 constants, pinned byte-exactly in tests/test_banding.py), replacing the four hard-coded scorer thresholds. Severity still decides what fires; bands decide what happens. AnalysisResult gains an additive borderline: bool field (default False, in to_dict) set when the score lands in [policy.borderline_at, ready_at) — the distinct near-miss "worth one more pass" signal from the spec's result states. interpretation_note is untouched, so existing serialized results are byte-identical under default policy. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * feat(input-cap): enforce Policy.max_chars with visible truncation analyze() now slices input to policy.max_chars before intent detection and rule execution — analysis cost is bounded regardless of input size (v0.2 took ~3.8 s on a 1.6 MB input; capped analysis of the same input runs in milliseconds). Truncation is never silent: AnalysisResult gains an additive truncated: bool field (default False, in to_dict), and the allowlist is evaluated against the same capped input so no scan escapes the budget. Default cap 10_000 matches the Policy default, pinned by test. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * feat(score-breakdown): additive score_breakdown audit field on the result calculate_score_with_breakdown(findings, policy=None) scores findings and returns the audit trail — {"base": 100, "penalties": [{"code", "severity", "points"(negative)}...], "final"} — one entry per distinct gap (or code when gap is None) in first-occurrence order, matching the result's gaps order. calculate_score keeps its v0.2 signature and delegates to it, so both stay in lockstep by construction. AnalysisResult gains an additive score_breakdown field (default None, deep- copied in to_dict so a serialized snapshot can never be rewritten by later mutation). With default policy the penalties mirror the v0.2 deductions exactly, so serialized defaults differ only by the new keys. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * test(pin-v02-defaults): pin every Policy default byte-exactly to v0.2 Dedicated pinning tests guard the calibration-drift risk the spec flags: each Policy default (ready_at 85, usable_at 60, strict_clarify_at 65, penalties 5/15/25, min_words 3, max_chars 10000, borderline_at 74) is asserted three ways — dataclass field defaults, constructed instance values, and the literal source line (underscore separators normalized) — plus a no-unpinned-fields guard so future fields cannot ship uncalibrated. Behavioral pins reproduce the observed v0.2 reference outputs under default policy (75/usable_with_warnings, 35/needs_clarification, 50, strict blocked/needs_clarification) and the v0.2 result and recommendation key shapes. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * test: integrate policy keys into sibling additive-contract tests The followups parity test and the language ordered-key-list test pin the exhaustive to_dict key set as of their own merge; the policy-calibration fields (borderline, truncated, score_breakdown) are additive per the compat contract, so the expected sets extend rather than the fields disappearing. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> --------- Co-authored-by: Obvious <obvious@obvious.ai> Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * feat: first-party writing domain — six rules, compose intent, complete advice (#10) * feat(writing-domain): first-party writing domain — six rules, compose intent, complete advice Register the writing domain through the same registry path user domains take: one globally unique fallback intent ("compose") and six rules covering the eval vocabulary's six gaps — audience, purpose, structure/format, source material, context, completeness — at the vocabulary-pinned severities (high/high/medium/high/medium/low). All term matching goes through the shared word-boundary matcher (inputguard.matching, PR #6). Rules follow the trigger-and-satisfy shape; the source-material rule fires only when existing text is referenced but not provided. Recommendations and follow-up entries ship for all six gaps in the same change, keeping the registry completeness invariant whole. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * test(writing-domain): per-rule, completeness, integration, no-leakage, and eval-label coverage Five layers: per-rule fire/silent behavior; the writing-domain completeness invariant (every writing gap has a four-key recommendation and one or two follow-up questions); analyze() integration end to end (underspecified -> needs_clarification at score 30, fully specified -> ready, partially specified -> usable_with_warnings); no-leakage in both directions at registry and analyze level; and the 14 labeled English writing rows from eval/cases.csv pinned as a calibration net. The exact-set registry expectations extend from 19 coding rules to 25 rules across two domains — the guards are strengthened, not weakened. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> --------- Co-authored-by: Obvious <obvious@obvious.ai> Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * feat: first-party data-analysis domain — six rules, analysis intent, complete advice (#11) Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * fix: degradation coverage for DG-011..014 (detector) (#12) Function-word layer (fr/es/pt stop-word lexicons with English margin) plus uncovered-run layer (3+ consecutive uncovered-script letters) route accented-Latin and mixed English+Han input to the explicit degraded path; DG-012's silent ready and DG-011/013/014's unnoted rule findings are gone. Degradation eval FP 3/14 -> 0/14, note honesty 10/14 -> 14/14, all 121 probe verdicts otherwise byte-identical. * fix: complete v0.3.0 release sweep — version, cap, degradation, docs, CI - Bump pyproject.toml and inputguard.__version__ to 0.3.0; write the full 0.3.0 CHANGELOG section from the merged release-wave history. - Report the literal "degraded" status on the degraded path in both modes; add the mixed-script (>=5% uncovered remainder) and non-English-Latin (<10% English evidence, filler-resistant denominator) gates so French/ Spanish/Portuguese and mixed English+Han prompts degrade with zero gaps instead of spurious ones; preserve the truncated flag on that path. - Linear dataset-filename extraction (follow-ups + data-analysis rule) with 1K/10K timing tests; benchmark doc records the before/after. - README: rewritten Non-English section, full result-object field table (follow_ups, probe fields, borderline, truncated, score_breakdown), writing and data-analysis rule tables. - eval: docs/false-positive-benchmark.md from measured results. - CI: minimal GitHub Actions workflow running pytest on push/PR (3.9+3.13). Eval: 116/121 matching (all 14 degradation rows, notes on 14/14); the 5 residual mismatches are pre-existing coding boundary/intent rows. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> Human author: Kalisetti Nihanth Naidu (nihanthnaidu007@gmail.com) * feat: CI matrix, latency budgets, CLI, and packaging metadata for v0.3 (#13) * test(registry-contract-guards): CI-enforced guards for the three review findings Dedicated suite for the adversarial-API-review gaps (art_hC18m78C), as the named acceptance artifact: - C3/P2: duplicate rule ids inside one register_domain call raise ValueError naming the id; a rejected batch mutates nothing. - B1/N3/P1: a rule whose check cannot be called as check(self, text) (extra param, zero param, keyword-only) is rejected at registration with an actionable message; defaulted and spec signatures accepted and fired. - C2/P4: an exception inside check() aborts analyze() as a RuntimeError carrying the rule id, its registration origin, and the original exception chained -- via the register_domain path (register_rule is covered in test_registry.py). The registry-isolation fixture moves to tests/conftest.py so the guard suite and the policy suite share one snapshot/restore implementation instead of importing across test modules. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * build(typing): annotate for mypy --strict without behavior change Pre-work for the CI strict-typing gate: missing Iterable[str] on the five rule modules' _contains_any helpers, Dict[str, Any] / Dict[str, str] generics on types.py and recommender.py, a pinned Rule-typed local in registry._instantiate, and the writing signals Tuple[str, ...] type argument. Also removes a literal duplicate "filter by"/"sort by" set entry in the feature scope signals (ruff B033); set semantics make this a no-op at runtime, and rules/__init__ marks its v0.2 signal re-exports as intentional via __all__ (ruff F401). Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * ci(actions-matrix): GitHub Actions CI with lint, typed, tested, benchmark, wheel gates The v0.2 failure mode was "every change ships on trust" — no CI, lint, typing, or coverage tooling at all. This adds the production gate set from the spec (§8) as three jobs: - test: Python 3.9-3.13 matrix running ruff, pytest with a 90% branch-coverage floor (--cov-branch --cov-fail-under=90 --strict), and mypy --strict (finally checking the shipped py.typed). - benchmark: latency budgets as merge gates, not notes. Budgets are derived from measured numbers (build-intent p50 21.94 ms at 10k chars including the C1 InsufficientContextRule double-run, debug 6.59 ms, ready 5.62 ms; 1.6 MB input bounded by the Policy cap at 17.1 ms vs the v0.2 unbounded 3.8 s scan). Hard gate is p50 with CI-noise headroom; p50/p99 print to the job log per spec. Budgets and derivation documented in tests/test_benchmarks.py. - wheel-zero-dep: the zero-dependency promise verified against the built artifact — wheel METADATA carries no Requires-Dist, a bare venv install pulls no third-party package, and the public API works from the installed wheel. Smoke steps run from a scratch directory so the repo's local inputguard/ package can never shadow the installed wheel. pyproject gains the matching tool config ([tool.ruff] E/F/W/B with E501 ignored, [tool.mypy] strict, explicit pytest testpaths) and the expanded [dev] extras the CI jobs install; all dev tooling stays out of runtime dependencies. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * feat(cli): inputguard analyze with argparse, to_dict JSON, CI-friendly exit codes The developer surface the adoption research flagged (spec §5): a zero-dependency argparse CLI shipped as a console script. - "inputguard analyze TEXT [--stdin] [--domain D] [--mode M] [--format text|json] [--min-score N]"; JSON output is the result's own to_dict() contract, byte-equal to the library's (asserted by test). - Exit codes documented in the module docstring and --help: 0 = analysis ok and at/above the --min-score floor when given; 1 = below the floor (the commit-hook / CI gate); 2 = usage error (missing text, unknown domain) reported cleanly on stderr, never a traceback. - Text rendering follows the spec §5 example shape: status, clarity, intent, domain, missing gaps, ask lines, and the degradation note when the language probe marks the result degraded (probe P2 input shows the note, never a silent ready). Entry point added in pyproject [project.scripts]; stdlib argparse and json only — the zero-dependency promise is untouched. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * chore(packaging-metadata): 0.3.0, Beta status, 3.13 classifier, project URLs The packaging drift the survey flagged is closed: classifiers extend to Python 3.13 (the runtime the project is developed and tested on), Development Status moves to 4 - Beta per the release plan (1.0 deliberately deferred until the extension API sees real use), [project.urls] points at the repository, changelog, and issue tracker, and inputguard.__version__ syncs with the pyproject version. The changelog's Unreleased v0.3 section is promoted to [0.3.0] with the CI and CLI entries from this PR. Also amends the wheel-gate assertion to what the zero-dependency promise actually means: no UNCONDITIONAL Requires-Dist entries (setuptools emits extra-gated `Requires-Dist: ...; extra == "dev"` lines for the optional tooling — verified against the built wheel, and the bare-venv install proves none of them install at runtime). Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * fix(benchmark-budgets): gate latency only in the untraced benchmark job, budgets from measured CI data The first CI run (35277175054) failed the latency asserts inside the coverage pass: coverage tracing roughly triples per-call cost (debug p50 21.14 ms vs 6.59 ms untraced; ready 20.35 ms vs 5.62 ms), and runner hardware is slower than the dev sandbox. Gating the same budgets under coverage in the test job AND untraced in the benchmark job was a design error — two different measurement conditions cannot share one threshold. - The coverage pass now ignores tests/test_benchmarks.py (with the observed numbers in a comment); latency gates belong exclusively to the dedicated benchmark job, which runs untraced and reflects real-world analysis cost. - Budgets re-derived from measured environments documented in the test module: sandbox untraced (build 21.94 / debug 6.59 / ready 5.62 ms p50) and the CI-traced observation above. New p50 budgets: build 65 ms (~3x), debug 25 ms (~3.8x), ready 25 ms (~4.4x); p95 sustained guards at 2.5x unchanged. Still order-of-magnitude gates: they catch algorithmic regressions (accidental O(n^2), a lost cap) that blow past by 10x+, while absorbing CI hardware variance. Local verification of the exact split: coverage pass 441 passed, 96.29% branch (floor 90); benchmark job untraced 4 passed; ruff + mypy --strict clean. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * fix: give the extraction linearization bounds CI headroom test (3.10) flaked on the rebased head (run 35277834661): the dotted-filler linearization test's absolute bound asserted a single wall-clock sample under 50 ms, and the coverage-traced 3.10 runner measured 50.61 ms. The meaningful guard is the 40x growth-rate assertion — the regression it protects against measured 1.3 s at 10 K, ~26x the bound. Both timing tests now use a 250 ms absolute bound (5x+ detection margin, no CI-hardware flake); the growth-rate assertion is unchanged. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> --------- Co-authored-by: Obvious <obvious@obvious.ai> Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * docs(readme-tested-examples): rewrite README for v0.3 with every example executed by tests The v0.2 README drifted from recommender.py within two releases. The rewritten README documents the v0.3 surface (extension API, Policy, follow-ups, degradation, CLI) and every example now runs: a new tests/test_readme_examples.py executes all python blocks in document order in a subprocess, diffs the embedded to_dict() JSON against live analyzer output, and checks exports, rule tables, gap vocabulary, and CLI output against the live registry. Example values carry real assertions sourced from runtime probes. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * docs(integration-recipes): LangChain and LiteLLM recipes consuming to_dict() Copy-paste integration patterns over the to_dict() serialization boundary — preflight-and-gate and clarify-then-complete. The recipes are documentation only: framework packages stay out of core and out of the dependency tree. tests/test_recipes.py executes every recipe block in document order, stubbing the single framework idiom each wiring block uses, so the inputguard-side contract cannot drift silently. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * docs(fp-benchmark): refresh measured rates and state the boundary target honestly Re-ran eval/measure_fp.py on the release branch: overall 116/121 match, FP 3.3% / FN 0.8%, zero false positives on true negatives and degradation rows, notes on 14/14 degradation rows, PF rows 19-20 ms. Adds the labeling-guide target contrast the doc lacked: fixture-style boundary FP measured 4/6 (v0.2 baseline 6/6, target 0/6) — driven by the matcher's deliberate inflection tolerance, documented as an eval-driven trade-off for a future release, not a label edit. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * chore(release-0.3.0): complete the 0.3.0 changelog for the docs wave Version 0.3.0 and the Development Status :: 4 - Beta classifier were already in pyproject.toml and inputguard.__version__ — verified, no bump needed. This commit completes the 0.3.0 entry with the docs wave: tested README rewrite, integration recipes, and the measured false-positive benchmark with the boundary-target contrast. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> * fix(readme-examples): 3.9-compatible annotation in the custom-rule example The custom-rule example annotated check() with PEP 604 union syntax (-> RuleFinding | None), which is evaluated at definition time and raises TypeError on Python 3.9 — caught by the CI matrix's 3.9 job running the README anti-drift gate. Use typing.Optional instead; an AST scan of every executed README block confirms no other runtime-evaluated 3.9-incompatible construct remains. Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com> --------- Co-authored-by: Obvious <obvious@obvious.ai> Co-authored-by: obvious-autobuild[bot] <262744130+obvious-autobuild[bot]@users.noreply.github.com> Co-authored-by: Kalisetti Nihanth Naidu <nihanthnaidu007@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Closes the v0.3 spec's §8 engineering baseline: until now the package shipped on trust — no CI, lint, typing, or coverage tooling at all. This PR adds the production gate set as three GitHub Actions jobs, the
inputguard analyzeCLI, and the release packaging metadata. Rebased ontorelease/v0.3.0@70f41ab(the release-sweep commit; its minimal 21-line CI workflow is superseded by this PR's full matrix, whose jobs are a strict superset).CI (
.github/workflows/ci.yml), three jobs:test— Python 3.9–3.13 matrix:ruff check,pytest -q --cov=inputguard --cov-branch --cov-fail-under=90 --strict(latency asserts excluded from the traced pass — they gate only in the dedicated benchmark job; coverage tracing roughly triples per-call cost, observed on run 35277175054), andmypy --strict inputguard(the shippedpy.typedis finally checked).benchmark— untraced latency budgets as merge gates; p50/p99 print to the job log.wheel-zero-dep— the zero-dependency promise verified against the built artifact: wheel METADATA has zero unconditionalRequires-Distentries, a bare-venv install pulls no third-party package, and the public API + CLI work from the installed wheel alone (smoke steps run from a scratch directory so the repo's localinputguard/can never shadow the wheel).Registry contract guards (
tests/test_registry_contract_guards.py) — the three tests the adversarial API review (art_hC18m78C) found missing, CI-enforced: same-call duplicate rule ids inregister_domainraise with nothing mutating; achecksignature that cannot be called ascheck(self, text)is rejected at registration with an actionable message; dispatch exceptions surface with rule id, registration origin, and the original exception chained — via theregister_domainpath (register_rulewas already covered).CLI (
inputguard/cli.py) —inputguard analyzeconsole script: stdlib argparse,--format jsonemitting byte-equalto_dict()output (asserted by test), exit codes 0 ok / 1 below--min-score/ 2 usage error, degradation note surfaced in text output.Packaging — version 0.3.0, Development Status → 4 - Beta, 3.13 classifier,
[project.urls], dev tooling confined to[dev]extras (zero runtime dependencies preserved), explicit[tool.ruff]/[tool.mypy]/[tool.pytest.ini_options]config. The release sweep already bumped the version; this PR contributes the Beta status, classifier, URLs, and changelog entries (woven into the sweep's[0.3.0]section). Two post-rebase CI flakes in sweep-authored timing tests were fixed by giving single-sample absolute bounds CI headroom (250 ms; the guarded regression is 1.3 s).Suite + coverage evidence
455 tests pass locally on the rebased head; the 89 v0.2 baseline tests remain untouched and green.
Coverage 96.28% branch (
--cov-branch --cov-fail-under=90 --strict— the exact CI command, run locally).Latency, measured on the rebased head (untraced, Python 3.13, 10,000-char inputs = the default cap, 60 samples after warmup):
InsufficientContextRuledouble-run)truncated=TrueBudget derivation (documented in the test module): untraced sandbox measurements multiplied ~3–4× to absorb CI hardware variance, calibrated against the first CI run's traced observations (debug p50 21.14 ms, ready 20.35 ms under coverage tracing — the reason latency asserts were moved out of the coverage pass). The budgets still gate by an order of magnitude: algorithmic regressions (accidental O(n²), a lost cap) blow past them ~10×, as the v0.2 unbounded 3.8 s scan would. Hard gate is p50; p95 sustained-regression guards sit at 2.5×.
Wheel gate verified locally end to end:
inputguard-0.3.0-py3-none-any.whlbuilds, installs into a bare venv (pip list→inputguard+ pip only), imports and runsanalyze()from outside the source tree, CLI JSON + exit-code smoke passes, and METADATA shows 0 unconditional / 6 extra-gated dev-onlyRequires-Distentries.Part of release PR #2.
Backend-only infrastructure change; visual evidence not applicable.
Human author: Kalisetti Nihanth Naidu (nihanthnaidu007@gmail.com)
🔗 Obvious Project · 🧵 Obvious Thread