From 94508cbe075d96896d251aeb13873500bc91dff7 Mon Sep 17 00:00:00 2001 From: Obvious Date: Thu, 17 Sep 2026 18:53:31 +0000 Subject: [PATCH 01/18] chore: release branch for InputGuard v0.3 production upgrade From efa8c63d0387169976eeebd6399c2d545c74e058 Mon Sep 17 00:00:00 2001 From: "obvious-autobuild[bot]" <262744130+obvious-autobuild[bot]@users.noreply.github.com> Date: Thu, 17 Sep 2026 19:26:08 +0000 Subject: [PATCH 02/18] feat: typed Rule protocol, in-process registry, and loud scorer failures (#3) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * feat(rule-protocol-registry): typed Rule protocol and in-process registry Open the engine (spec art_bTvdPdJS §1): rules and domains become registered data instead of code paths. - inputguard/registry.py: runtime-checkable Rule protocol (id, domain, severity, gap + check(text, intent)) and RuleRegistry with register_rule (decorator and imperative forms) and register_domain (priority-ordered intent signals, single empty-terms fallback intent). Duplicate rule ids, unknown severities, and — once a domain exists — undeclared rule domains raise at registration time. - detector.py: the coding priority chain becomes INTENT_SIGNALS data; detect_intent gains an optional signals parameter so registered domains run the same chain; normalize() becomes public so analyze() can honor the "rules receive normalized text" contract. Nothing consumes the registry yet; behavior is unchanged (89 tests green). Zero new dependencies. Co-authored-by: Kalisetti Nihanth Naidu * refactor(built-ins-via-registry): dispatch analyze() through the registry Replace the hardcoded domain whitelist and the if/elif intent dispatch with registry lookups (spec art_bTvdPdJS §1): - analyzer.py: analyze() resolves domain signals via REGISTRY.get_domain_signals (unknown domain still raises ValueError, now naming every registered domain) and runs rules via REGISTRY.rules_for_intent — the deduped, deterministic-finding-order contract of the v0.2 runners is preserved by the registry's registration-order iteration. - rules/*: each of the 19 built-in rules gains a thin registry adapter class exposing the v0.2 check function through the Rule protocol, registered with @register_rule — the built-ins dogfood the exact extension path users get. The coding catch-all adapter re-runs the built-in build rules internally to preserve its fires-only-when-empty semantics under the frozen check(text, intent) contract. - rules/__init__.py: registers the coding domain (INTENT_SIGNALS chain plus all 19 rule adapters) through register_domain; the v0.2 runner exports are unchanged. Behavior-identical: the 89 existing tests pass unmodified. Co-authored-by: Kalisetti Nihanth Naidu * fix(scorer-loud-failures): unknown severities raise; unknown gaps keep advice Close the two silent-failure edges the survey pinned (art_CnghyDxp §2.3, §2.4) per spec art_bTvdPdJS §6: - scorer.py: calculate_score validates every finding's severity up front and raises ValueError on an unknown one — the silent zero penalty from .get(severity, 0) is gone (rank comparisons and penalty lookups go through validated paths). - recommender.py: get_recommendations no longer drops gaps without a curated entry; unknown gaps get a documented four-key generic fallback (_fallback_recommendation), keeping advice complete for rule authors who ship a new gap. Existing behavior for known severities and gaps is unchanged: the 89 existing tests pass unmodified. Co-authored-by: Kalisetti Nihanth Naidu * test(registry-invariants): registry contract and parity tests New tests/test_registry.py pins the foundation's acceptance points: - dogfooding: all 19 built-in rules registered through the registry, coding domain signals in priority order with a single fallback intent, deterministic build-intent rule order; - validation: duplicate rule id, unknown severity, and unknown domain raise ValueError at registration; decorator form registers and satisfies the Rule protocol; - loud failures: unknown severity raises in calculate_score (including alongside known severities); unknown gaps get the documented four-key fallback, never a silent drop; - end to end: a custom registered rule changes a result (finding, gap, score -15, fallback recommendation); a registered custom domain analyzes; duplicate domain registration raises; - invariants: the registry-walking completeness check (every built-in gap has a complete recommendation entry), dispatch parity with the v0.2 runners across nine inputs, and 65 parallel analyze() calls across 8 threads matching sequential results. Full suite: 114 passed (89 pre-existing, unmodified). Co-authored-by: Kalisetti Nihanth Naidu --------- Co-authored-by: Obvious Co-authored-by: Kalisetti Nihanth Naidu --- inputguard/analyzer.py | 40 ++-- inputguard/detector.py | 56 ++++-- inputguard/recommender.py | 24 ++- inputguard/registry.py | 327 +++++++++++++++++++++++++++++++ inputguard/rules/__init__.py | 44 ++++- inputguard/rules/coding.py | 142 ++++++++++++++ inputguard/rules/debug.py | 52 +++++ inputguard/rules/explanation.py | 38 ++++ inputguard/rules/feature.py | 52 +++++ inputguard/rules/optimization.py | 52 +++++ inputguard/scorer.py | 25 ++- tests/test_registry.py | 306 +++++++++++++++++++++++++++++ 12 files changed, 1113 insertions(+), 45 deletions(-) create mode 100644 inputguard/registry.py create mode 100644 tests/test_registry.py diff --git a/inputguard/analyzer.py b/inputguard/analyzer.py index 03c3844..9eb3b2e 100644 --- a/inputguard/analyzer.py +++ b/inputguard/analyzer.py @@ -1,14 +1,11 @@ from __future__ import annotations -from typing import List +from typing import List, Set -from inputguard.detector import detect_intent +import inputguard.rules # noqa: F401 — importing registers the coding domain and the 19 built-in rules +from inputguard.detector import detect_intent, normalize from inputguard.recommender import get_recommendations -from inputguard.rules.coding import run_coding_rules -from inputguard.rules.debug import run_debug_rules -from inputguard.rules.optimization import run_optimization_rules -from inputguard.rules.explanation import run_explanation_rules -from inputguard.rules.feature import run_feature_rules +from inputguard.registry import REGISTRY from inputguard.scorer import calculate_score, get_status from inputguard.types import AnalysisResult, RuleFinding @@ -36,29 +33,26 @@ def analyze(self, user_input: str, domain: str = "coding") -> AnalysisResult: ) if not user_input.strip(): raise ValueError("user_input must be a non-empty, non-whitespace string.") - if domain != "coding": - raise ValueError( - f"Unsupported domain: {domain!r}. Phase 1 only supports 'coding'." - ) - detected_intent = detect_intent(user_input) + domain_signals = REGISTRY.get_domain_signals(domain) + + detected_intent = detect_intent(user_input, domain_signals) + normalized = normalize(user_input) - if detected_intent == "debug": - findings: List[RuleFinding] = run_debug_rules(user_input) - elif detected_intent == "optimization": - findings = run_optimization_rules(user_input) - elif detected_intent == "explanation": - findings = run_explanation_rules(user_input) - elif detected_intent == "feature": - findings = run_feature_rules(user_input) - else: - findings = run_coding_rules(user_input) + findings: List[RuleFinding] = [] + seen_codes: Set[str] = set() + for rule in REGISTRY.rules_for_intent(detected_intent): + finding = rule.check(normalized, detected_intent) + if finding is None or finding.code in seen_codes: + continue + seen_codes.add(finding.code) + findings.append(finding) score = calculate_score(findings) status = get_status(score, self.mode) gaps: List[str] = [] - seen = set() + seen: Set[str] = set() for f in findings: if f.gap is not None and f.gap not in seen: gaps.append(f.gap) diff --git a/inputguard/detector.py b/inputguard/detector.py index 35a17eb..31db70e 100644 --- a/inputguard/detector.py +++ b/inputguard/detector.py @@ -2,6 +2,7 @@ import re +from typing import Iterable, Mapping, Optional, Tuple DEBUG_SIGNALS = { "error", "exception", "traceback", "not working", "isn't working", @@ -50,7 +51,20 @@ } -def _normalize(text: str) -> str: +# The coding intent chain: priority-ordered (intent, terms) pairs. The single +# intent with empty terms ("build") is the fallback. This is the single source +# of truth for detection — the rule modules key their gates off the same sets, +# and the coding domain registers this chain in the rule registry. +INTENT_SIGNALS = ( + ("debug", DEBUG_SIGNALS), + ("optimization", OPTIMIZATION_SIGNALS), + ("explanation", EXPLANATION_SIGNALS), + ("feature", FEATURE_SIGNALS), + ("build", ()), +) + + +def normalize(text: str) -> str: return re.sub(r"\s+", " ", text.strip().lower()) @@ -58,7 +72,10 @@ def _contains_any(text: str, terms) -> bool: return any(term in text for term in terms) -def detect_intent(text: str) -> str: +def detect_intent( + text: str, + signals: Optional[Iterable[Tuple[str, Iterable[str]]]] = None, +) -> str: """ Detect the intent type of a coding input. @@ -70,15 +87,28 @@ def detect_intent(text: str) -> str: feature beats build. Build is the fallback. Ambiguous inputs always resolve to the highest-priority match. + + Pass ``signals`` to run the same chain over a domain's own + priority-ordered ``(intent, terms)`` pairs — the single intent with + empty terms is the fallback. Defaults to :data:`INTENT_SIGNALS`, the + built-in coding chain. """ - normalized = _normalize(text) - - if _contains_any(normalized, DEBUG_SIGNALS): - return "debug" - if _contains_any(normalized, OPTIMIZATION_SIGNALS): - return "optimization" - if _contains_any(normalized, EXPLANATION_SIGNALS): - return "explanation" - if _contains_any(normalized, FEATURE_SIGNALS): - return "feature" - return "build" + normalized = normalize(text) + + chain = INTENT_SIGNALS if signals is None else tuple( + signals.items() if isinstance(signals, Mapping) else signals + ) + + fallback: Optional[str] = None + for intent, terms in chain: + if not terms: + fallback = intent + elif _contains_any(normalized, terms): + return intent + + if fallback is not None: + return fallback + raise ValueError( + "detect_intent requires exactly one fallback intent (an entry with " + "empty signal terms); none was provided." + ) diff --git a/inputguard/recommender.py b/inputguard/recommender.py index f37275c..56c42bf 100644 --- a/inputguard/recommender.py +++ b/inputguard/recommender.py @@ -115,9 +115,29 @@ } +def _fallback_recommendation(gap: str) -> dict: + """Generic four-key advice for a gap with no curated entry. + + The documented fallback (spec art_bTvdPdJS §6): unknown gaps keep their + advice complete instead of vanishing silently from the recommendations + list — the v0.2 behavior that let rule authors ship empty advice. + """ + return { + "gap": gap, + "what_is_missing": f"You haven't provided the {gap} needed to act on this request.", + "what_to_provide": f"Describe the {gap} explicitly — one or two concrete sentences is enough.", + "why_it_matters": f"Without the {gap}, the request can be read several ways and the answer may miss what you actually need.", + } + + def get_recommendations(gaps: List[str]) -> List[dict]: + """Build one four-key recommendation per gap, in input order. + + A gap without a curated entry gets the documented fallback + (see :func:`_fallback_recommendation`) — never a silent drop. + """ out: List[dict] = [] for gap in gaps: - if gap in _RECOMMENDATIONS: - out.append(dict(_RECOMMENDATIONS[gap])) + entry = _RECOMMENDATIONS.get(gap) + out.append(dict(entry) if entry is not None else _fallback_recommendation(gap)) return out diff --git a/inputguard/registry.py b/inputguard/registry.py new file mode 100644 index 0000000..3762bcc --- /dev/null +++ b/inputguard/registry.py @@ -0,0 +1,327 @@ +"""Typed Rule protocol and the in-process registries that open the engine. + +v0.3 turns InputGuard's closed, coding-only checker into a pluggable clarity +engine: rules and domains are registered data, not code paths. + +- A rule is any object with the four members (``id``, ``domain``, + ``severity``, ``gap``) and a ``check(text, intent)`` method — the + :class:`Rule` protocol. +- A domain is a named analysis scope (``"coding"`` ships built in) that + declares its intent signals in priority order via + :meth:`RuleRegistry.register_domain`. + +Registration happens at import/startup time; ``analyze()`` only reads the +registry, preserving the thread-safe, dependency-free pipeline v0.2 +established. +""" + +from __future__ import annotations + +from typing import ( + Dict, + Iterable, + List, + Mapping, + Optional, + Protocol, + Set, + Tuple, + Union, + runtime_checkable, +) + +from inputguard.types import RuleFinding + +__all__ = [ + "KNOWN_SEVERITIES", + "REGISTRY", + "Rule", + "RuleRegistry", + "register_domain", + "register_rule", +] + +KNOWN_SEVERITIES: Tuple[str, ...] = ("low", "medium", "high") + +# Normalized domain signals: priority-ordered (intent, terms) pairs. The one +# intent with empty terms is the fallback for inputs no other intent matches. +SignalSpec = Tuple[Tuple[str, Tuple[str, ...]], ...] + +_REQUIRED_MEMBERS = ("id", "domain", "severity", "gap", "check") + + +@runtime_checkable +class Rule(Protocol): + """The v0.3 extension contract: four members and one method. + + ``id`` must be unique across the registry, ``severity`` must be one of + ``'low' | 'medium' | 'high'`` (validated at registration), and ``gap`` + groups the rule's findings for scoring dedup (``None`` dedupes by code + instead). ``check`` receives lowercased, whitespace-collapsed text plus + the intent detected for this ``analyze()`` call, and returns at most one + :class:`~inputguard.types.RuleFinding` — ``None`` when the rule does not + fire. + """ + + id: str + domain: str + severity: str + gap: Optional[str] + + def check(self, text: str, intent: str) -> Optional[RuleFinding]: + ... # pragma: no cover — protocol body + + +def _instantiate(rule_cls: type) -> Rule: + try: + return rule_cls() + except TypeError as exc: + raise TypeError( + f"Cannot register rule class {rule_cls.__name__!r}: it must be " + f"constructible with no arguments. Register an instance instead: " + f"register_rule({rule_cls.__name__}())." + ) from exc + + +def _normalize_signals( + name: str, + signals: Union[Mapping[str, Iterable[str]], Iterable[Tuple[str, Iterable[str]]]], +) -> SignalSpec: + items = tuple(signals.items() if isinstance(signals, Mapping) else signals) + if not items: + raise ValueError(f"Domain {name!r}: signals must declare at least one intent.") + + normalized: List[Tuple[str, Tuple[str, ...]]] = [] + fallbacks = 0 + for intent, terms in items: + if not isinstance(intent, str) or not intent: + raise ValueError( + f"Domain {name!r}: intent names must be non-empty strings, got {intent!r}." + ) + if isinstance(terms, str) or not isinstance(terms, Iterable): + raise TypeError( + f"Domain {name!r}: terms for intent {intent!r} must be an iterable " + f"of strings, got {terms!r}." + ) + term_tuple = tuple(terms) + for term in term_tuple: + if not isinstance(term, str) or not term: + raise ValueError( + f"Domain {name!r}: terms for intent {intent!r} must be " + f"non-empty strings, got {term!r}." + ) + if not term_tuple: + fallbacks += 1 + normalized.append((intent, term_tuple)) + + if fallbacks != 1: + raise ValueError( + f"Domain {name!r}: exactly one intent must have empty signal terms " + f"(the fallback intent); found {fallbacks}." + ) + return tuple(normalized) + + +class RuleRegistry: + """In-process registry of rules and analysis domains. + + Writes happen at import/startup; ``analyze()`` reads are plain dict + lookups and iteration over registered rules, so parallel ``analyze()`` + calls stay consistent (the v0.2 thread-safety contract). Registering + while another thread is mid-``analyze()`` is not supported. + """ + + def __init__(self) -> None: + self._rules: Dict[str, Rule] = {} + self._domains: Dict[str, SignalSpec] = {} + + # -- registration ---------------------------------------------------- + + def register_rule(self, rule: Union[type, Rule]) -> Union[type, Rule]: + """Register a rule — decorator or imperative form. + + ``@register_rule`` above a rule class registers a zero-argument + instance and returns the class unchanged; ``register_rule(instance)`` + registers the instance and returns it. + + Raises ``ValueError`` on a duplicate rule id, an unknown severity, or + — once any domain is registered — a domain no registered domain + declares as an intent. Raises ``TypeError`` when the object is not + shaped like a :class:`Rule` or the class cannot be constructed with + no arguments. + """ + if isinstance(rule, type): + instance = _instantiate(rule) + self._validate(instance) + self._add(instance) + return rule + self._validate(rule) + self._add(rule) + return rule + + def register_domain( + self, + name: str, + signals: Union[Mapping[str, Iterable[str]], Iterable[Tuple[str, Iterable[str]]]], + rules: Iterable[Union[type, Rule]] = (), + ) -> None: + """Register an analysis domain: intent signals plus its rules. + + ``signals`` maps intent names to trigger terms; insertion order is the + priority chain, and the single intent with empty terms is the fallback + for inputs no other intent matches. Every rule passed in must declare + one of those intents as its ``domain``; rules already registered (for + example via the ``@register_rule`` decorator) are reused, and the rest + are registered through the same path. + + Raises ``ValueError`` on a duplicate domain name, malformed signals + (no fallback intent), or a rule whose domain the domain does not + declare. + """ + if not isinstance(name, str) or not name: + raise ValueError(f"Domain name must be a non-empty string, got {name!r}.") + if name in self._domains: + raise ValueError( + f"Duplicate domain: {name!r}. Domains must be registered once." + ) + normalized_signals = _normalize_signals(name, signals) + declared = {intent for intent, _ in normalized_signals} + + coerced = [self._coerce(entry) for entry in rules] + for rule in coerced: + self._validate(rule) + for rule in coerced: + if rule.domain not in declared: + raise ValueError( + f"Rule {rule.id!r} declares domain {rule.domain!r}, which is not " + f"an intent of domain {name!r}. Declared intents: {sorted(declared)}." + ) + for rule in coerced: + existing = self._rules.get(rule.id) + if existing is not None and existing is not rule: + raise ValueError( + f"Duplicate rule id: {rule.id!r}. Rule ids must be unique." + ) + + # All checks passed — mutate. The domain is stored before its rules so + # the known-domain check in _add sees the intents it declares. + self._domains[name] = normalized_signals + for rule in coerced: + if rule.id not in self._rules: + self._add(rule) + + # -- reads (thread-safe: no mutation happens during analyze) ---------- + + def get_domain_signals(self, name: str) -> SignalSpec: + """Return the domain's priority-ordered ``(intent, terms)`` pairs. + + Raises ``ValueError`` for an unregistered domain — the registry-owned + replacement for v0.2's hardcoded domain whitelist. + """ + try: + return self._domains[name] + except KeyError: + registered = ", ".join(repr(d) for d in self._domains) or "none" + raise ValueError( + f"Unsupported domain: {name!r}. Registered domains: {registered}." + ) from None + + def domain_names(self) -> Tuple[str, ...]: + """Registered domain names, in registration order.""" + return tuple(self._domains) + + def rules_for_intent(self, intent: str) -> List[Rule]: + """Rules whose domain is this intent scope, in registration order.""" + return [rule for rule in self._rules.values() if rule.domain == intent] + + def rule_ids(self) -> Tuple[str, ...]: + """Registered rule ids, in registration order.""" + return tuple(self._rules) + + def rules(self) -> Tuple[Rule, ...]: + """All registered rules, in registration order.""" + return tuple(self._rules.values()) + + def get_rule(self, rule_id: str) -> Optional[Rule]: + """Return the rule registered under ``rule_id``, or ``None``.""" + return self._rules.get(rule_id) + + # -- internals --------------------------------------------------------- + + def _coerce(self, entry: Union[type, Rule]) -> Rule: + """Accept a rule class or instance. + + A class already registered (via the decorator, which instantiates) + resolves to its registered instance; a fresh class is instantiated. + """ + if not isinstance(entry, type): + return entry + probe_id = getattr(entry, "id", None) + existing = self._rules.get(probe_id) if isinstance(probe_id, str) else None + return existing if existing is not None else _instantiate(entry) + + def _validate(self, rule: Rule) -> None: + missing = [m for m in _REQUIRED_MEMBERS if not hasattr(rule, m)] + if missing: + raise TypeError( + f"Rule {rule!r} is missing required member(s): " + f"{', '.join(repr(m) for m in missing)}. A rule needs id, domain, " + "severity, gap, and a check(text, intent) method." + ) + if not isinstance(rule.id, str) or not rule.id: + raise ValueError(f"Rule id must be a non-empty string, got {rule.id!r}.") + if not isinstance(rule.domain, str) or not rule.domain: + raise ValueError( + f"Rule {rule.id!r}: domain must be a non-empty string, got {rule.domain!r}." + ) + if rule.severity not in KNOWN_SEVERITIES: + raise ValueError( + f"Rule {rule.id!r}: unknown severity {rule.severity!r}. " + f"Expected one of: 'low', 'medium', 'high'." + ) + if rule.gap is not None and (not isinstance(rule.gap, str) or not rule.gap): + raise ValueError( + f"Rule {rule.id!r}: gap must be None or a non-empty string, got {rule.gap!r}." + ) + if not callable(rule.check): + raise TypeError( + f"Rule {rule.id!r}: check must be callable — " + "check(text, intent) -> Optional[RuleFinding]." + ) + + def _add(self, rule: Rule) -> None: + if rule.id in self._rules: + raise ValueError(f"Duplicate rule id: {rule.id!r}. Rule ids must be unique.") + # Known-domain check: vacuous until the first domain registers (built-in + # rules register during package import, before any domain exists), then + # enforced for everything registered afterwards. + if self._domains and rule.domain not in self._declared_intents(): + raise ValueError( + f"Rule {rule.id!r} declares domain {rule.domain!r}, which no " + f"registered domain declares as an intent. Registered domains: " + f"{', '.join(repr(d) for d in self._domains)}." + ) + self._rules[rule.id] = rule + + def _declared_intents(self) -> Set[str]: + declared: Set[str] = set() + for signal_spec in self._domains.values(): + declared.update(intent for intent, _ in signal_spec) + return declared + + +REGISTRY = RuleRegistry() + + +def register_rule(rule: Union[type, Rule]) -> Union[type, Rule]: + """Module-level form of :meth:`RuleRegistry.register_rule` — usable as a decorator.""" + return REGISTRY.register_rule(rule) + + +def register_domain( + name: str, + signals: Union[Mapping[str, Iterable[str]], Iterable[Tuple[str, Iterable[str]]]], + rules: Iterable[Union[type, Rule]] = (), +) -> None: + """Module-level form of :meth:`RuleRegistry.register_domain`.""" + REGISTRY.register_domain(name, signals, rules) diff --git a/inputguard/rules/__init__.py b/inputguard/rules/__init__.py index f7c6066..fc32218 100644 --- a/inputguard/rules/__init__.py +++ b/inputguard/rules/__init__.py @@ -1,8 +1,27 @@ -from inputguard.rules.coding import run_coding_rules -from inputguard.rules.debug import run_debug_rules -from inputguard.rules.optimization import run_optimization_rules -from inputguard.rules.explanation import run_explanation_rules -from inputguard.rules.feature import run_feature_rules +"""Built-in rule modules and the registry wiring for the coding domain. + +Importing this package registers the coding domain — its intent signals and +all 19 built-in rules — through the exact same registry path a user rule +takes. The v0.2 ``run_*_rules`` functions stay exported for backward +compatibility, but the analyzer dispatches through the registry now. +""" + +from inputguard.detector import ( + DEBUG_SIGNALS, + EXPLANATION_SIGNALS, + FEATURE_SIGNALS, + INTENT_SIGNALS, + OPTIMIZATION_SIGNALS, +) +from inputguard.registry import REGISTRY +from inputguard.rules.coding import CODING_RULES, run_coding_rules +from inputguard.rules.debug import DEBUG_RULES, run_debug_rules +from inputguard.rules.explanation import EXPLANATION_RULES, run_explanation_rules +from inputguard.rules.feature import FEATURE_RULES, run_feature_rules +from inputguard.rules.optimization import ( + OPTIMIZATION_RULES, + run_optimization_rules, +) __all__ = [ "run_coding_rules", @@ -11,3 +30,18 @@ "run_explanation_rules", "run_feature_rules", ] + +# The coding domain: intent signals in strict priority order (the single +# empty-terms entry, "build", is the fallback) and all 19 built-in rules, +# registered through the same register_domain path a user domain takes. +REGISTRY.register_domain( + "coding", + INTENT_SIGNALS, + rules=( + *CODING_RULES, + *DEBUG_RULES, + *OPTIMIZATION_RULES, + *EXPLANATION_RULES, + *FEATURE_RULES, + ), +) diff --git a/inputguard/rules/coding.py b/inputguard/rules/coding.py index eeb1025..60ca560 100644 --- a/inputguard/rules/coding.py +++ b/inputguard/rules/coding.py @@ -3,6 +3,7 @@ import re from typing import List, Optional, Set +from inputguard.registry import register_rule from inputguard.types import RuleFinding @@ -299,3 +300,144 @@ def run_coding_rules(text: str) -> List[RuleFinding]: seen_codes.add(catchall.code) return findings + + +# Registry adapters: the v0.2 check functions above stay the single home of +# the rule logic; these classes expose it through the v0.3 Rule protocol and +# register it through the same path a user rule takes. + + +@register_rule +class MissingLanguageRule: + """Registry adapter for check_missing_language.""" + + id = "missing_language" + domain = "build" + severity = "high" + gap = "programming language" + + def check(self, text: str, intent: str) -> Optional[RuleFinding]: + return check_missing_language(text) + + +@register_rule +class MissingApiStructureRule: + """Registry adapter for check_missing_api_structure.""" + + id = "missing_api_structure" + domain = "build" + severity = "high" + gap = "api structure" + + def check(self, text: str, intent: str) -> Optional[RuleFinding]: + return check_missing_api_structure(text) + + +@register_rule +class MissingDataModelRule: + """Registry adapter for check_missing_data_model.""" + + id = "missing_data_model" + domain = "build" + severity = "high" + gap = "data model" + + def check(self, text: str, intent: str) -> Optional[RuleFinding]: + return check_missing_data_model(text) + + +@register_rule +class MissingIntegrationSpecificsRule: + """Registry adapter for check_missing_integration_specifics.""" + + id = "missing_integration_specifics" + domain = "build" + severity = "medium" + gap = "integration specifics" + + def check(self, text: str, intent: str) -> Optional[RuleFinding]: + return check_missing_integration_specifics(text) + + +@register_rule +class MissingAuthTypeRule: + """Registry adapter for check_missing_auth_type.""" + + id = "missing_auth_type" + domain = "build" + severity = "high" + gap = "authentication type" + + def check(self, text: str, intent: str) -> Optional[RuleFinding]: + return check_missing_auth_type(text) + + +@register_rule +class MissingOutputFormatRule: + """Registry adapter for check_missing_output_format.""" + + id = "missing_output_format" + domain = "build" + severity = "medium" + gap = "output format" + + def check(self, text: str, intent: str) -> Optional[RuleFinding]: + return check_missing_output_format(text) + + +@register_rule +class IntentWithoutDetailRule: + """Registry adapter for _check_intent_without_detail.""" + + id = "intent_without_language" + domain = "build" + severity = "high" + gap = "programming language" + + def check(self, text: str, intent: str) -> Optional[RuleFinding]: + return _check_intent_without_detail(text) + + +@register_rule +class InsufficientContextRule: + """Catch-all safety net for build-intent input (registry adapter). + + v0.2's run_coding_rules passed this rule the findings collected so far and + it fired only when that list was empty. A Rule sees only (text, intent), + so the adapter re-runs the other built-in build rules — they are pure + functions, so the verdict is identical. Findings from user-registered + rules are not visible here: the catch-all suppresses on the built-in + build rules only. + """ + + id = "insufficient_context" + domain = "build" + severity = "high" + gap = "task context" + + def check(self, text: str, intent: str) -> Optional[RuleFinding]: + normalized = _normalize(text) + seen_codes: Set[str] = set() + prior: List[RuleFinding] = [] + for check_fn in _CHECKS: + result = check_fn(normalized) + if result is not None and result.code not in seen_codes: + prior.append(result) + seen_codes.add(result.code) + intent_finding = _check_intent_without_detail(normalized) + if intent_finding is not None and intent_finding.code not in seen_codes: + prior.append(intent_finding) + seen_codes.add(intent_finding.code) + return _check_insufficient_context(normalized, prior) + + +CODING_RULES = ( + MissingLanguageRule, + MissingApiStructureRule, + MissingDataModelRule, + MissingIntegrationSpecificsRule, + MissingAuthTypeRule, + MissingOutputFormatRule, + IntentWithoutDetailRule, + InsufficientContextRule, +) diff --git a/inputguard/rules/debug.py b/inputguard/rules/debug.py index 9369619..3067ae7 100644 --- a/inputguard/rules/debug.py +++ b/inputguard/rules/debug.py @@ -4,6 +4,7 @@ from typing import List, Optional from inputguard.detector import DEBUG_SIGNALS +from inputguard.registry import register_rule from inputguard.types import RuleFinding @@ -103,3 +104,54 @@ def run_debug_rules(text: str) -> List[RuleFinding]: if result: findings.append(result) return _dedupe(findings) + + +# Registry adapters: the v0.2 check functions above stay the single home of +# the rule logic; these classes expose it through the v0.3 Rule protocol and +# register it through the same path a user rule takes. + + +@register_rule +class MissingErrorMessageRule: + """Registry adapter for check_missing_error_message.""" + + id = "missing_error_message" + domain = "debug" + severity = "high" + gap = "error description" + + def check(self, text: str, intent: str) -> Optional[RuleFinding]: + return check_missing_error_message(text) + + +@register_rule +class MissingExpectedVsActualRule: + """Registry adapter for check_missing_expected_vs_actual.""" + + id = "missing_expected_vs_actual" + domain = "debug" + severity = "high" + gap = "expected vs actual behavior" + + def check(self, text: str, intent: str) -> Optional[RuleFinding]: + return check_missing_expected_vs_actual(text) + + +@register_rule +class MissingDebugCodeContextRule: + """Registry adapter for check_missing_debug_code_context.""" + + id = "missing_debug_code_context" + domain = "debug" + severity = "medium" + gap = "code context" + + def check(self, text: str, intent: str) -> Optional[RuleFinding]: + return check_missing_debug_code_context(text) + + +DEBUG_RULES = ( + MissingErrorMessageRule, + MissingExpectedVsActualRule, + MissingDebugCodeContextRule, +) diff --git a/inputguard/rules/explanation.py b/inputguard/rules/explanation.py index ef356f4..8a5a5cf 100644 --- a/inputguard/rules/explanation.py +++ b/inputguard/rules/explanation.py @@ -4,6 +4,7 @@ from typing import List, Optional from inputguard.detector import EXPLANATION_SIGNALS +from inputguard.registry import register_rule from inputguard.types import RuleFinding @@ -84,3 +85,40 @@ def run_explanation_rules(text: str) -> List[RuleFinding]: if result: findings.append(result) return _dedupe(findings) + + +# Registry adapters: the v0.2 check functions above stay the single home of +# the rule logic; these classes expose it through the v0.3 Rule protocol and +# register it through the same path a user rule takes. + + +@register_rule +class MissingCodeReferenceRule: + """Registry adapter for check_missing_code_reference.""" + + id = "missing_code_reference" + domain = "explanation" + severity = "high" + gap = "code reference" + + def check(self, text: str, intent: str) -> Optional[RuleFinding]: + return check_missing_code_reference(text) + + +@register_rule +class MissingExplanationDepthRule: + """Registry adapter for check_missing_explanation_depth.""" + + id = "missing_explanation_depth" + domain = "explanation" + severity = "low" + gap = "explanation depth" + + def check(self, text: str, intent: str) -> Optional[RuleFinding]: + return check_missing_explanation_depth(text) + + +EXPLANATION_RULES = ( + MissingCodeReferenceRule, + MissingExplanationDepthRule, +) diff --git a/inputguard/rules/feature.py b/inputguard/rules/feature.py index de629fb..7afa551 100644 --- a/inputguard/rules/feature.py +++ b/inputguard/rules/feature.py @@ -4,6 +4,7 @@ from typing import List, Optional from inputguard.detector import FEATURE_SIGNALS +from inputguard.registry import register_rule from inputguard.types import RuleFinding @@ -104,3 +105,54 @@ def run_feature_rules(text: str) -> List[RuleFinding]: if result: findings.append(result) return _dedupe(findings) + + +# Registry adapters: the v0.2 check functions above stay the single home of +# the rule logic; these classes expose it through the v0.3 Rule protocol and +# register it through the same path a user rule takes. + + +@register_rule +class MissingExistingStackRule: + """Registry adapter for check_missing_existing_stack.""" + + id = "missing_existing_stack" + domain = "feature" + severity = "high" + gap = "existing stack" + + def check(self, text: str, intent: str) -> Optional[RuleFinding]: + return check_missing_existing_stack(text) + + +@register_rule +class MissingFeatureScopeRule: + """Registry adapter for check_missing_feature_scope.""" + + id = "missing_feature_scope" + domain = "feature" + severity = "high" + gap = "feature scope" + + def check(self, text: str, intent: str) -> Optional[RuleFinding]: + return check_missing_feature_scope(text) + + +@register_rule +class MissingCompletionCriteriaRule: + """Registry adapter for check_missing_completion_criteria.""" + + id = "missing_completion_criteria" + domain = "feature" + severity = "low" + gap = "completion criteria" + + def check(self, text: str, intent: str) -> Optional[RuleFinding]: + return check_missing_completion_criteria(text) + + +FEATURE_RULES = ( + MissingExistingStackRule, + MissingFeatureScopeRule, + MissingCompletionCriteriaRule, +) diff --git a/inputguard/rules/optimization.py b/inputguard/rules/optimization.py index d89d34b..031346c 100644 --- a/inputguard/rules/optimization.py +++ b/inputguard/rules/optimization.py @@ -4,6 +4,7 @@ from typing import List, Optional from inputguard.detector import OPTIMIZATION_SIGNALS +from inputguard.registry import register_rule from inputguard.types import RuleFinding @@ -103,3 +104,54 @@ def run_optimization_rules(text: str) -> List[RuleFinding]: if result: findings.append(result) return _dedupe(findings) + + +# Registry adapters: the v0.2 check functions above stay the single home of +# the rule logic; these classes expose it through the v0.3 Rule protocol and +# register it through the same path a user rule takes. + + +@register_rule +class MissingOptimizationTargetRule: + """Registry adapter for check_missing_optimization_target.""" + + id = "missing_optimization_target" + domain = "optimization" + severity = "high" + gap = "optimization target" + + def check(self, text: str, intent: str) -> Optional[RuleFinding]: + return check_missing_optimization_target(text) + + +@register_rule +class MissingPerformanceBaselineRule: + """Registry adapter for check_missing_performance_baseline.""" + + id = "missing_performance_baseline" + domain = "optimization" + severity = "medium" + gap = "performance baseline" + + def check(self, text: str, intent: str) -> Optional[RuleFinding]: + return check_missing_performance_baseline(text) + + +@register_rule +class MissingOptimizationConstraintRule: + """Registry adapter for check_missing_optimization_constraint.""" + + id = "missing_optimization_constraint" + domain = "optimization" + severity = "low" + gap = "optimization constraint" + + def check(self, text: str, intent: str) -> Optional[RuleFinding]: + return check_missing_optimization_constraint(text) + + +OPTIMIZATION_RULES = ( + MissingOptimizationTargetRule, + MissingPerformanceBaselineRule, + MissingOptimizationConstraintRule, +) diff --git a/inputguard/scorer.py b/inputguard/scorer.py index 113fba7..d6894ad 100644 --- a/inputguard/scorer.py +++ b/inputguard/scorer.py @@ -18,17 +18,38 @@ _STRICT_NEEDS = 65 +def _severity_rank(severity: str) -> int: + """Rank lookup that fails loudly: an unknown severity must never silently score zero.""" + try: + return _SEVERITY_RANK[severity] + except KeyError: + raise ValueError( + f"Unknown severity: {severity!r}. Expected one of: 'low', 'medium', 'high'." + ) from None + + def calculate_score(findings: List[RuleFinding]) -> int: + """100 minus one penalty per distinct gap (or code, when gap is None), clamped to [0, 100]. + + Unknown severities raise ``ValueError`` — a typo'd severity would + otherwise distort every score silently (the v0.2 ``.get(severity, 0)`` + behavior). + """ + # Validate every severity before scoring so a bad finding fails loudly + # even when a later finding would otherwise mask it in the dedup loop. + for f in findings: + _severity_rank(f.severity) + highest_by_gap: Dict[str, str] = {} for f in findings: key = f.gap if f.gap is not None else f.code current = highest_by_gap.get(key) - if current is None or _SEVERITY_RANK.get(f.severity, 0) > _SEVERITY_RANK.get(current, 0): + if current is None or _severity_rank(f.severity) > _severity_rank(current): highest_by_gap[key] = f.severity score = 100 for severity in highest_by_gap.values(): - score -= SEVERITY_PENALTIES.get(severity, 0) + score -= SEVERITY_PENALTIES[severity] return max(0, min(100, score)) diff --git a/tests/test_registry.py b/tests/test_registry.py new file mode 100644 index 0000000..7e0da55 --- /dev/null +++ b/tests/test_registry.py @@ -0,0 +1,306 @@ +"""Registry contract tests for the v0.3 foundation. + +Covers the acceptance points of the rule-registry PR: registration +validation (duplicate id, unknown severity, unknown domain), the +unknown-gap fallback, a custom rule changing a result end to end, a +registered custom domain, the registry-walking completeness invariant, +dispatch parity with the v0.2 runners, and thread-safe parallel +analyze() consistency. +""" + +from __future__ import annotations + +from concurrent.futures import ThreadPoolExecutor + +import pytest + +from inputguard import InputGuard, RuleFinding +from inputguard.recommender import _RECOMMENDATIONS, get_recommendations +from inputguard.registry import REGISTRY, Rule, register_domain, register_rule +from inputguard.rules.coding import run_coding_rules +from inputguard.rules.debug import run_debug_rules +from inputguard.rules.explanation import run_explanation_rules +from inputguard.rules.feature import run_feature_rules +from inputguard.rules.optimization import run_optimization_rules +from inputguard.scorer import calculate_score + +# The 19 built-in rule ids, as documented in the README table. +EXPECTED_BUILTIN_IDS = { + "missing_language", + "missing_api_structure", + "missing_data_model", + "missing_integration_specifics", + "missing_auth_type", + "missing_output_format", + "intent_without_language", + "insufficient_context", + "missing_error_message", + "missing_expected_vs_actual", + "missing_debug_code_context", + "missing_optimization_target", + "missing_performance_baseline", + "missing_optimization_constraint", + "missing_code_reference", + "missing_explanation_depth", + "missing_existing_stack", + "missing_feature_scope", + "missing_completion_criteria", +} + +_V02_INTENT_RUNNERS = { + "build": run_coding_rules, + "debug": run_debug_rules, + "optimization": run_optimization_rules, + "explanation": run_explanation_rules, + "feature": run_feature_rules, +} + + +def _test_rule(rule_id="test_rule", domain="debug", severity="low", gap=None, finding=None): + """Build an instance of a minimal rule class with configurable members.""" + + class TestRule: + pass + + TestRule.id = rule_id + TestRule.domain = domain + TestRule.severity = severity + TestRule.gap = gap + TestRule.check = lambda self, text, intent: finding + return TestRule() + + +@pytest.fixture +def registry_isolation(): + """Snapshot the registry around tests that mutate it.""" + rules_before = dict(REGISTRY._rules) + domains_before = dict(REGISTRY._domains) + yield REGISTRY + REGISTRY._rules.clear() + REGISTRY._rules.update(rules_before) + REGISTRY._domains.clear() + REGISTRY._domains.update(domains_before) + + +# --- the 19 built-ins dogfood the registry path --------------------------- + + +def test_all_builtin_rules_registered_through_registry(): + assert len(REGISTRY.rule_ids()) == 19 + assert set(REGISTRY.rule_ids()) == EXPECTED_BUILTIN_IDS + for rule in REGISTRY.rules(): + assert isinstance(rule, Rule) + assert rule.severity in ("low", "medium", "high") + + +def test_coding_domain_registered_with_priority_signals(): + assert REGISTRY.domain_names() == ("coding",) + chain = REGISTRY.get_domain_signals("coding") + assert [intent for intent, _ in chain] == [ + "debug", + "optimization", + "explanation", + "feature", + "build", + ] + # The single fallback intent carries no terms. + assert [terms for _, terms in chain if not terms] == [()] + + +def test_build_intent_rule_order_is_deterministic(): + ids = [r.id for r in REGISTRY.rules_for_intent("build")] + assert ids == [ + "missing_language", + "missing_api_structure", + "missing_data_model", + "missing_integration_specifics", + "missing_auth_type", + "missing_output_format", + "intent_without_language", + "insufficient_context", + ] + + +# --- registration validation ---------------------------------------------- + + +def test_duplicate_rule_id_raises_value_error(registry_isolation): + with pytest.raises(ValueError, match="Duplicate rule id"): + register_rule(_test_rule(rule_id="missing_error_message")) + + +def test_unknown_severity_raises_value_error_at_registration(registry_isolation): + with pytest.raises(ValueError, match="unknown severity"): + register_rule(_test_rule(severity="critical")) + + +def test_unknown_domain_raises_value_error_at_registration(registry_isolation): + with pytest.raises(ValueError, match="domain"): + register_rule(_test_rule(domain="nosuchintent")) + + +def test_register_rule_decorator_form(registry_isolation): + @register_rule + class AlwaysAskForDeadlineRule: + id = "test_always_ask_deadline" + domain = "debug" + severity = "medium" + gap = "deadline" + + def check(self, text, intent): + return RuleFinding( + code=self.id, + message="Test rule fired.", + severity=self.severity, + gap=self.gap, + ) + + registered = REGISTRY.get_rule("test_always_ask_deadline") + assert registered is not None + assert isinstance(registered, AlwaysAskForDeadlineRule) + assert isinstance(registered, Rule) + + +# --- scorer / recommender loud failures ------------------------------------ + + +def test_scorer_unknown_severity_raises_value_error(): + finding = RuleFinding(code="x", message="m", severity="catastrophic") + with pytest.raises(ValueError, match="Unknown severity"): + calculate_score([finding]) + + +def test_scorer_unknown_severity_raises_even_alongside_known_severities(): + known = RuleFinding(code="y", message="m", severity="high") + unknown = RuleFinding(code="x", message="m", severity="typo") + with pytest.raises(ValueError, match="Unknown severity"): + calculate_score([known, unknown]) + + +def test_unknown_gap_gets_documented_fallback(): + recs = get_recommendations(["totally new gap"]) + assert len(recs) == 1 + entry = recs[0] + assert set(entry) == {"gap", "what_is_missing", "what_to_provide", "why_it_matters"} + assert entry["gap"] == "totally new gap" + assert all(isinstance(v, str) and v for v in entry.values()) + + +def test_unknown_gap_fallback_preserves_order_and_known_entries(): + recs = get_recommendations(["programming language", "totally new gap"]) + assert [r["gap"] for r in recs] == ["programming language", "totally new gap"] + # The known gap still gets its curated entry (copied, not shared). + assert recs[0]["why_it_matters"] == _RECOMMENDATIONS["programming language"]["why_it_matters"] + + +# --- custom rules and domains end to end ----------------------------------- + + +def test_custom_rule_changes_result(registry_isolation): + text = "fix the bug in my app" + before = InputGuard().analyze(text) + assert not any(f.code == "test_always_ask_deadline" for f in before.findings) + + finding = RuleFinding( + code="test_always_ask_deadline", + message="Test rule fired.", + severity="medium", + gap="deadline", + ) + register_rule(_test_rule("test_always_ask_deadline", finding=finding)) + + result = InputGuard().analyze(text) + codes = [f.code for f in result.findings] + assert "test_always_ask_deadline" in codes + assert "deadline" in result.gaps + assert result.clarity_score == before.clarity_score - 15 + deadline_rec = next(r for r in result.recommendations if r["gap"] == "deadline") + assert all(str(v) for v in deadline_rec.values()) + + +def test_custom_domain_analyzes_end_to_end(registry_isolation): + register_domain( + "testdom", + {"thing": ("widget",), "fallback": ()}, + rules=[], + ) + result = InputGuard().analyze("widget please", domain="testdom") + assert result.detected_intent == "thing" + assert result.findings == [] + assert result.clarity_score == 100 + assert result.status == "ready" + + +def test_duplicate_domain_registration_raises(registry_isolation): + with pytest.raises(ValueError, match="Duplicate domain"): + register_domain("coding", {"x": ("term",), "fallback": ()}) + + +# --- completeness invariant ------------------------------------------------- + + +def test_every_builtin_gap_has_recommendation_and_complete_entry(): + gaps = {rule.gap for rule in REGISTRY.rules() if rule.gap is not None} + assert gaps, "built-in rules must declare gaps" + for gap in gaps: + assert gap in _RECOMMENDATIONS, f"gap {gap!r} has no recommendation entry" + entry = _RECOMMENDATIONS[gap] + assert set(entry) == {"gap", "what_is_missing", "what_to_provide", "why_it_matters"} + assert all(isinstance(v, str) and v for v in entry.values()) + + +# --- dispatch parity with the v0.2 runners ---------------------------------- + + +def _ordered_unique_gaps(findings): + seen = set() + out = [] + for f in findings: + if f.gap is not None and f.gap not in seen: + out.append(f.gap) + seen.add(f.gap) + return out + + +@pytest.mark.parametrize( + "text", + [ + "Build a REST API using FastAPI. Store users in PostgreSQL with fields for name and email. Add email and password login.", + "make me something good", + "I need a mobile app", + "fix the bug in my app", + "the deploy fails with a TypeError: cannot read property of undefined, it should return the cached value instead", + "make my database query faster", + "explain recursion in this function", + "add search to my app", + "what does this code do", + ], +) +def test_registry_dispatch_matches_v02_runner_output(text): + result = InputGuard().analyze(text) + legacy = _V02_INTENT_RUNNERS[result.detected_intent](text) + assert [f.code for f in result.findings] == [f.code for f in legacy] + assert result.gaps == _ordered_unique_gaps(legacy) + assert result.clarity_score == calculate_score(legacy) + + +# --- thread safety ----------------------------------------------------------- + + +def test_parallel_analyze_stays_consistent(): + guard = InputGuard() + texts = [ + "fix the bug in my app", + "make my database query faster", + "what does this code do", + "add search to my app", + "Build a REST API using FastAPI. Store users in PostgreSQL. Add login.", + ] + sequential = [guard.analyze(t).to_dict() for t in texts] + + with ThreadPoolExecutor(max_workers=8) as pool: + parallel = list(pool.map(lambda t: guard.analyze(t).to_dict(), texts * 13)) + + assert len(parallel) == 65 + for offset in range(0, len(parallel), len(texts)): + assert parallel[offset : offset + len(texts)] == sequential From 38a9e56ef51c73c285c4f783789c0265fe489bb9 Mon Sep 17 00:00:00 2001 From: "obvious-autobuild[bot]" <262744130+obvious-autobuild[bot]@users.noreply.github.com> Date: Thu, 17 Sep 2026 19:41:31 +0000 Subject: [PATCH 03/18] feat: per-gap follow-up questions engine with additive follow_ups result field (#4) Per-gap templated clarifying questions (one or two per gap, deduped, gap-ordered), slot fills for function/dataset names, a documented fallback for unknown gaps, the additive AnalysisResult.follow_ups field serialized in to_dict(), and the registry-walking completeness invariant extended to require follow-ups for every built-in gap. 139/139 tests green; existing 114 unmodified. - feat(follow-up-templates): per-gap clarifying question engine - feat(result-field): additive follow_ups on AnalysisResult - test(follow-up-completeness): extended invariant + engine behavior tests Co-authored-by: Kalisetti Nihanth Naidu --- inputguard/analyzer.py | 5 + inputguard/followups.py | 203 +++++++++++++++++++++++++ inputguard/types.py | 5 + tests/test_followups.py | 327 ++++++++++++++++++++++++++++++++++++++++ 4 files changed, 540 insertions(+) create mode 100644 inputguard/followups.py create mode 100644 tests/test_followups.py diff --git a/inputguard/analyzer.py b/inputguard/analyzer.py index 9eb3b2e..23900f9 100644 --- a/inputguard/analyzer.py +++ b/inputguard/analyzer.py @@ -4,6 +4,7 @@ import inputguard.rules # noqa: F401 — importing registers the coding domain and the 19 built-in rules from inputguard.detector import detect_intent, normalize +from inputguard.followups import get_follow_ups from inputguard.recommender import get_recommendations from inputguard.registry import REGISTRY from inputguard.scorer import calculate_score, get_status @@ -59,6 +60,9 @@ def analyze(self, user_input: str, domain: str = "coding") -> AnalysisResult: seen.add(f.gap) recommendations = get_recommendations(gaps) + # Slot fills read the original input (case preserved); rules ran on + # the normalized text, question extraction does not need to. + follow_ups = get_follow_ups(gaps, user_input) high_count = sum(1 for f in findings if f.severity == "high") interpretation_note = None @@ -73,4 +77,5 @@ def analyze(self, user_input: str, domain: str = "coding") -> AnalysisResult: recommendations=recommendations, findings=findings, interpretation_note=interpretation_note, + follow_ups=follow_ups, ) diff --git a/inputguard/followups.py b/inputguard/followups.py new file mode 100644 index 0000000..ff8e590 --- /dev/null +++ b/inputguard/followups.py @@ -0,0 +1,203 @@ +"""Per-gap templated clarifying questions — the v0.3 differentiator. + +InputGuard already knows exactly what is missing from a prompt; this module +turns each gap finding into one or two clarifying questions the user can +answer verbatim. That converts "your input is vague" into "here is the +sentence to send next" — no adjacent tool generates structured follow-up +questions from gap findings (domain brief art_bzqrA3c7 §2). + +Design, mirroring the recommender: + +- ``_FOLLOW_UP_QUESTIONS`` is the gap -> questions table, one or two + templates per built-in gap, in stable order. +- A template may reference a slot (``{function}``, ``{dataset}``); the slot + is filled from the original input when something recognizable is there and + the template is skipped otherwise. Every gap keeps at least one slot-free + template, so a gap never yields zero questions. +- A gap with no table entry gets the documented fallback question — never + silence (the same contract the recommendations fallback keeps). + +The module is pure (table lookups plus regex over the input, no shared +mutation), so it preserves the thread-safe pipeline ``analyze()`` runs. +""" + +from __future__ import annotations + +import re +from typing import Dict, List, Mapping, Optional, Set, Tuple + +__all__ = ["get_follow_ups"] + + +# One or two question templates per built-in gap, in stable order. Templates +# are phrased for the person typing the prompt; {function} / {dataset} slots +# are filled from the input when extractable. The spec-pinned examples +# ("Which function or module should get faster?", "How slow is it today, and +# what latency would be acceptable?") are kept byte-exact for the gaps the +# spec's example result shows them for. +_FOLLOW_UP_QUESTIONS: Dict[str, Tuple[str, ...]] = { + "programming language": ( + "Which programming language or framework should this use?", + "Where does it need to run — a web browser, a server, your terminal, or a phone?", + ), + "api structure": ( + "What actions should the API support — for example, list, create, or delete a resource?", + ), + "data model": ( + "Which fields does {dataset} contain, and which of them matter for this task?", + "What data needs to be stored, and which fields matter for each item?", + ), + "integration specifics": ( + "Which specific feature of the integration do you need — for example, one-time payments, monthly subscriptions, or SMS notifications?", + ), + "authentication type": ( + "Which login method should it use — email and password, Google or GitHub sign-in, an API key, or a magic link?", + ), + "output format": ( + "What should this be when it's done — a web app, a command-line tool, a REST API, a script, or a mobile app?", + ), + "task context": ( + "What are you trying to build or accomplish, in a sentence or two?", + ), + "error description": ( + "What is the exact error message or exception you're seeing (copy it verbatim if you can)?", + ), + "expected vs actual behavior": ( + "What did you expect to happen, and what actually happens instead?", + ), + "code context": ( + "Where is {function} defined, and what does it currently do?", + "Which file, function, or part of your code does the problem live in?", + ), + "optimization target": ( + "Which function or module should get faster?", + "What makes {function} slow today, and how fast should it be?", + ), + "performance baseline": ( + "How slow is it today, and what latency would be acceptable?", + ), + "optimization constraint": ( + "What must not change while it gets faster — an interface, readability, behavior others depend on?", + ), + "code reference": ( + "Which function, class, or concept should the explanation focus on?", + ), + "explanation depth": ( + "How deep should the explanation go — a quick overview, step by step, or all the way down to internals?", + ), + "existing stack": ( + "What is your existing app built with — languages, frameworks, and database?", + ), + "feature scope": ( + "What exactly should the new feature do, from the user's point of view?", + ), + "completion criteria": ( + "What does 'done' look like — what should you be able to do when the feature works?", + ), +} + +# Slot names a template may reference. A template naming anything else is a +# table bug and raises at render time instead of silently never firing (the +# same loud-failure contract the scorer applies to unknown severities). +_KNOWN_SLOTS: Tuple[str, ...] = ("function", "dataset") + +_SLOT_RE = re.compile(r"\{([a-z_]+)\}") + +# A code call site: an identifier immediately followed by "(" — no space, so +# natural-language parentheticals ("fix this (urgently)") never match, while +# process_orders() and def process_orders( both do. Control-flow keywords are +# excluded; the first match in the input wins (deterministic). +_FUNCTION_RE = re.compile(r"(? Optional[str]: + for match in _FUNCTION_RE.finditer(text): + name = match.group(1) + if name not in _FUNCTION_KEYWORDS: + return name + return None + + +def _extract_dataset(text: str) -> Optional[str]: + file_match = _DATASET_FILE_RE.search(text) + if file_match is not None: + return file_match.group(1) + named_match = _DATASET_NAMED_RE.search(text) + if named_match is not None: + return named_match.group(1) + return None + + +def _extract_slots(text: str) -> Dict[str, str]: + """Pull fillable slots from the original input (case preserved).""" + slots: Dict[str, str] = {} + function_name = _extract_function(text) + if function_name is not None: + slots["function"] = function_name + dataset_name = _extract_dataset(text) + if dataset_name is not None: + slots["dataset"] = dataset_name + return slots + + +def _render(template: str, slots: Mapping[str, str]) -> Optional[str]: + """Fill a template's slots, or return None when one is not fillable. + + A slot name outside ``_KNOWN_SLOTS`` is a table bug and raises instead of + quietly producing a template that can never fire. + """ + names = _SLOT_RE.findall(template) + unknown = [name for name in names if name not in _KNOWN_SLOTS] + if unknown: + raise ValueError( + f"Unknown follow-up slot(s) {unknown} in template {template!r}. " + f"Known slots: {', '.join(_KNOWN_SLOTS)}." + ) + if any(name not in slots for name in names): + return None + return template.format(**slots) + + +def _fallback_question(gap: str) -> str: + """Generic question for a gap with no curated entry. + + The follow-up mirror of the recommendations fallback: a user-registered + rule may declare a gap this table has never seen, and its findings still + earn a question instead of silently empty advice. + """ + return f"Can you add the {gap} this request is missing?" + + +def get_follow_ups(gaps: List[str], text: str) -> List[str]: + """Build deduped clarifying questions for ``gaps``, in gap order. + + Each gap contributes its one or two templates, rendered against the + slots extractable from ``text``; a question already produced by an + earlier gap is dropped. A gap without a curated entry gets the + documented fallback question. + """ + slots = _extract_slots(text) + out: List[str] = [] + seen: Set[str] = set() + for gap in gaps: + templates = _FOLLOW_UP_QUESTIONS.get(gap, (_fallback_question(gap),)) + for template in templates: + rendered = _render(template, slots) + if rendered is None or rendered in seen: + continue + seen.add(rendered) + out.append(rendered) + return out diff --git a/inputguard/types.py b/inputguard/types.py index e389c32..8dcdd5c 100644 --- a/inputguard/types.py +++ b/inputguard/types.py @@ -21,6 +21,10 @@ class AnalysisResult: recommendations: List[dict] = field(default_factory=list) findings: List[RuleFinding] = field(default_factory=list) interpretation_note: Optional[str] = None + # v0.3, additive: templated clarifying questions, one or two per gap, + # deduped and ordered with `gaps`. Appended after the v0.2 fields so any + # positional construction keeps its meaning. + follow_ups: List[str] = field(default_factory=list) def to_dict(self) -> dict: return { @@ -29,6 +33,7 @@ def to_dict(self) -> dict: "detected_intent": self.detected_intent, "gaps": list(self.gaps), "recommendations": [dict(r) for r in self.recommendations], + "follow_ups": list(self.follow_ups), "findings": [asdict(f) for f in self.findings], "interpretation_note": self.interpretation_note, } diff --git a/tests/test_followups.py b/tests/test_followups.py new file mode 100644 index 0000000..fd74b0e --- /dev/null +++ b/tests/test_followups.py @@ -0,0 +1,327 @@ +"""Follow-up question engine tests. + +Three layers, matching the acceptance criteria: + +- the registry-walking completeness invariant extended to follow-ups: + every built-in gap has one or two curated questions, and the question + table covers exactly the built-in gaps (drift in either direction fails); +- behavior on results: follow_ups present and non-empty wherever findings + fire, deduped, ordered with gaps, empty when the input is clean; +- slot fills (function name, dataset name), the unknown-gap fallback, + English score parity with v0.2, and thread-safe parallel consistency. + +All existing tests are untouched — this file only adds. +""" + +from __future__ import annotations + +from concurrent.futures import ThreadPoolExecutor + +import pytest + +from inputguard import AnalysisResult, InputGuard +from inputguard.followups import ( + _FOLLOW_UP_QUESTIONS, + _extract_slots, + _render, + get_follow_ups, +) +from inputguard.recommender import _RECOMMENDATIONS +from inputguard.registry import REGISTRY, register_rule +from inputguard.rules.coding import run_coding_rules +from inputguard.rules.debug import run_debug_rules +from inputguard.rules.explanation import run_explanation_rules +from inputguard.rules.feature import run_feature_rules +from inputguard.rules.optimization import run_optimization_rules +from inputguard.scorer import calculate_score +from inputguard.types import RuleFinding + + +# --- the extended registry-walking completeness invariant -------------------- + + +def _builtin_gaps(): + gaps = {rule.gap for rule in REGISTRY.rules() if rule.gap is not None} + assert gaps, "built-in rules must declare gaps" + return gaps + + +def test_every_builtin_gap_has_at_least_one_follow_up_question(): + # The extended invariant: a rule can ship a recommendation entry and + # still leave the user with nothing to answer — every built-in gap must + # also carry one or two curated clarifying questions. + for gap in _builtin_gaps(): + assert gap in _FOLLOW_UP_QUESTIONS, f"gap {gap!r} has no follow-up question" + questions = _FOLLOW_UP_QUESTIONS[gap] + assert 1 <= len(questions) <= 2, f"gap {gap!r} must have one or two questions" + for question in questions: + assert isinstance(question, str) and question.endswith("?"), ( + f"gap {gap!r} has a malformed follow-up question: {question!r}" + ) + + +def test_follow_up_table_covers_exactly_the_builtin_gaps(): + # Two-way check: a question entry for a gap no rule declares is dead + # config, so the table keys must equal the built-in gaps exactly. + assert set(_FOLLOW_UP_QUESTIONS) == _builtin_gaps() + + +def test_every_template_renders_with_unknown_slot_failing_loudly(): + # Rendering every template against empty slots exercises the render path: + # a typo'd slot name raises (loud table bug); a known-but-unfilled slot + # skips the template (documented behavior); plain templates pass through. + for gap, questions in _FOLLOW_UP_QUESTIONS.items(): + for question in questions: + rendered = _render(question, slots={}) + if "{" in question: + assert rendered is None + else: + assert rendered == question + + +def test_every_builtin_gap_has_recommendation_and_follow_up_pair(): + # The completeness invariant read whole: advice AND a question per gap. + for gap in _builtin_gaps(): + assert gap in _RECOMMENDATIONS, f"gap {gap!r} has no recommendation entry" + assert gap in _FOLLOW_UP_QUESTIONS, f"gap {gap!r} has no follow-up question" + + +# --- follow_ups on results ---------------------------------------------------- + + +@pytest.mark.parametrize( + "text", + [ + "build a REST API", # build + "fix the login crash", # debug + "make this faster", # optimization + "explain async", # explanation + "add auth to my app", # feature + ], +) +def test_follow_ups_present_and_non_empty_on_findings_bearing_results(text): + result = InputGuard().analyze(text) + assert result.findings, "test input must be findings-bearing" + assert result.gaps + assert isinstance(result.follow_ups, list) + assert result.follow_ups, "findings-bearing results must ask at least one question" + assert all( + isinstance(q, str) and q for q in result.follow_ups + ), "follow-ups must be non-empty strings" + # One or two per gap, deduped. + assert len(result.gaps) <= len(result.follow_ups) <= 2 * len(result.gaps) + assert len(result.follow_ups) == len(set(result.follow_ups)), "follow-ups must be deduped" + + +def test_ready_results_have_empty_follow_ups(): + text = ( + "Write a Python function using FastAPI and PostgreSQL that returns " + "a JSON list of users" + ) + result = InputGuard().analyze(text) + assert result.clarity_score == 100 + assert result.follow_ups == [] + + +def test_follow_ups_order_matches_gap_order_and_spec_examples(): + # "make this faster" is the spec's example input: its first two + # follow-ups are pinned byte-exact to the spec's example output, which + # also pins follow-up order to gap order. + result = InputGuard().analyze("make this faster") + assert result.gaps == [ + "optimization target", + "performance baseline", + "optimization constraint", + ] + assert result.follow_ups[0] == "Which function or module should get faster?" + assert result.follow_ups[1] == "How slow is it today, and what latency would be acceptable?" + assert result.follow_ups[2] in { + "What must not change while it gets faster — an interface, readability, " + "behavior others depend on?", + } + + +# --- dedup, ordering, and the unknown-gap fallback ---------------------------- + + +def test_follow_ups_are_deduped_across_gaps(): + questions = get_follow_ups( + ["task context", "task context", "error description"], + "fix the login crash", + ) + assert len(questions) == len(set(questions)) + + +def test_follow_ups_order_follows_the_gap_list(): + error_first = get_follow_ups(["error description", "task context"], "text") + task_first = get_follow_ups(["task context", "error description"], "text") + assert error_first[0].startswith("What is the exact error") + assert task_first[0].startswith("What are you trying to build") + + +def test_unknown_gap_gets_the_documented_fallback_question(): + assert get_follow_ups(["deadline"], "ship the report by Friday") == [ + "Can you add the deadline this request is missing?" + ] + + +def test_user_rule_with_unknown_gap_still_gets_a_follow_up(): + finding = RuleFinding( + code="missing_deadline", + message="No deadline given.", + severity="medium", + gap="deadline", + ) + + class DeadlineRule: + id = "test_followups_missing_deadline" + domain = "build" + severity = "medium" + gap = "deadline" + + def check(self, text: str, intent: str): + return finding + + rules_before = {rule.id: rule for rule in REGISTRY.rules()} + try: + register_rule(DeadlineRule()) + result = InputGuard().analyze("build a REST API") + assert result.findings + # "build a REST API" also fires the built-in missing-language rule; + # the point here is the unknown gap: it joins the result and its + # fallback question lands in follow_ups, in gap order. + assert result.gaps[-1] == "deadline" + assert result.follow_ups[-1] == "Can you add the deadline this request is missing?" + assert result.follow_ups[-1] in result.to_dict()["follow_ups"] + finally: + REGISTRY._rules.clear() + REGISTRY._rules.update(rules_before) + + +# --- slot fills --------------------------------------------------------------- + + +def test_function_slot_fill_from_call_site(): + result = InputGuard().analyze( + "speed up my code — the render loop calls update_positions() on every frame" + ) + assert "optimization target" in result.gaps + assert ( + "What makes update_positions slow today, and how fast should it be?" + in result.follow_ups + ) + + +def test_dataset_slot_fill_from_named_file(): + result = InputGuard().analyze( + "analyze the sales_data.csv table and find the top customers" + ) + assert "data model" in result.gaps + assert ( + "Which fields does sales_data.csv contain, and which of them matter for this task?" + in result.follow_ups + ) + + +def test_prose_parentheticals_never_fill_the_function_slot(): + # Natural-language parentheses and control-flow keywords are not call + # sites; without a fillable slot the slotted template is skipped and no + # unfilled "{...}" ever leaks into a question. + assert "function" not in _extract_slots("fix this (urgently) broken build") + assert "function" not in _extract_slots("if(x > 3) the loop is wrong") + for question in get_follow_ups(["code context"], "the build is broken"): + assert "{" not in question + + +def test_slot_extraction_preserves_the_original_case(): + slots = _extract_slots("speed up ProcessOrders(") + assert slots["function"] == "ProcessOrders" + + +# --- English score parity with v0.2 ------------------------------------------- + +_V02_INTENT_RUNNERS = { + "build": run_coding_rules, + "debug": run_debug_rules, + "optimization": run_optimization_rules, + "explanation": run_explanation_rules, + "feature": run_feature_rules, +} + +_V02_TO_DICT_KEYS = { + "status", + "clarity_score", + "detected_intent", + "gaps", + "recommendations", + "findings", + "interpretation_note", +} + + +def _ordered_unique_gaps(findings): + seen = set() + out = [] + for f in findings: + if f.gap is not None and f.gap not in seen: + out.append(f.gap) + seen.add(f.gap) + return out + + +@pytest.mark.parametrize( + "text", + [ + "build a REST API", + "fix the login crash", + "make this faster", + "explain async", + "add auth to my app", + ], +) +def test_english_score_parity_with_v02(text): + # follow_ups are derived from gaps after scoring; the v0.2 numbers must + # not move: same findings, same gap order, same score as the legacy + # runners, and to_dict() gains exactly one additive key. + result = InputGuard().analyze(text) + legacy = _V02_INTENT_RUNNERS[result.detected_intent](text) + + assert [f.code for f in result.findings] == [f.code for f in legacy] + assert result.gaps == _ordered_unique_gaps(legacy) + assert result.clarity_score == calculate_score(legacy) + + d = result.to_dict() + assert set(d) == _V02_TO_DICT_KEYS | {"follow_ups"} + v02_shaped = AnalysisResult( + status=result.status, + clarity_score=result.clarity_score, + detected_intent=result.detected_intent, + gaps=result.gaps, + recommendations=result.recommendations, + findings=result.findings, + interpretation_note=result.interpretation_note, + ) + assert {k: d[k] for k in _V02_TO_DICT_KEYS} == { + k: v02_shaped.to_dict()[k] for k in _V02_TO_DICT_KEYS + } + + +# --- thread safety ------------------------------------------------------------- + + +def test_parallel_analyze_follow_ups_stay_consistent(): + texts = [ + "build a REST API", + "fix the login crash", + "make this faster", + "explain async", + "add auth to my app", + ] + guard = InputGuard() + sequential = [guard.analyze(t).to_dict() for t in texts] + with ThreadPoolExecutor(max_workers=13) as pool: + parallel = list(pool.map(lambda t: guard.analyze(t).to_dict(), texts * 13)) + for i in range(len(sequential)): + expected = sequential[i] + for j in range(i, len(parallel), len(sequential)): + assert parallel[j] == expected From 7f15ed02cb4514fac54ed0241bdd5511610271eb Mon Sep 17 00:00:00 2001 From: "obvious-autobuild[bot]" <262744130+obvious-autobuild[bot]@users.noreply.github.com> Date: Thu, 17 Sep 2026 19:52:27 +0000 Subject: [PATCH 04/18] =?UTF-8?q?feat:=20multilingual=20degradation=20?= =?UTF-8?q?=E2=80=94=20script=20probe=20and=20explicit=20degraded=20path?= =?UTF-8?q?=20for=20uncovered=20scripts=20(#5)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * feat(script-probe): add unicodedata script-histogram language probe Pure, thread-safe script classification built on unicodedata only — each letter's Unicode name mentions its script, so token scanning classifies scripts with no hardcoded range tables. Classifies the input's dominant script, a coarse script-derived language guess, and the English-heuristic coverage band (full / partial / none / unknown). A deterministic stride sample bounds probe cost on arbitrarily long input. Co-authored-by: Kalisetti Nihanth Naidu * feat(degraded-path): explicit degradation for scripts without heuristic coverage Wire the script probe into analyze(): uncovered scripts skip the English-only rules outright, take a 20-point confidence penalty (score 80 — never 'ready' in either mode), and report detected_intent 'undetermined' instead of an unearned fallback. Results carry the three additive fields detected_language, heuristic_coverage, and degradation_note; to_dict() includes them additively. Covered scripts (full/partial coverage) run the unchanged pipeline — partial adds a note without a penalty. Co-authored-by: Kalisetti Nihanth Naidu * test(non-english-honesty): probe-P2 regression, Unicode samples, thread parity The probe-P2 input (Chinese) must return a degraded result with a degradation_note — never ready/100. Also covers: degraded shape and mode behavior, per-script degradation, English parity (probe fields on covered input, byte-parity scores: 'fix my code' -> 35, repeated build -> 50), partial-coverage rules-run-with-note, validation order, additive to_dict() keys, systematic no-crash Unicode samples with the note-consistency invariant, bounded long-input probing, and 64-worker parallel analyze() consistency across English/Chinese inputs. Co-authored-by: Kalisetti Nihanth Naidu * fix(tests): adapt to_dict key-set expectations to the merged additive contract PR #4's parity test asserted the exact to_dict key set with follow_ups as the only additive key; the three language-probe fields are additive per the v0.3 spec, so the expectation now covers the merged additive set. The key-order test documents the merged 11-key contract. Co-authored-by: Kalisetti Nihanth Naidu --------- Co-authored-by: Obvious Co-authored-by: Kalisetti Nihanth Naidu --- CHANGELOG.md | 14 ++ README.md | 29 +++ inputguard/analyzer.py | 41 +++- inputguard/language.py | 288 +++++++++++++++++++++++++++ inputguard/types.py | 10 + tests/test_followups.py | 11 +- tests/test_language.py | 425 ++++++++++++++++++++++++++++++++++++++++ 7 files changed, 815 insertions(+), 3 deletions(-) create mode 100644 inputguard/language.py create mode 100644 tests/test_language.py diff --git a/CHANGELOG.md b/CHANGELOG.md index 9910fbd..4879789 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,5 +1,19 @@ # Changelog +## [Unreleased] +### Added +- Multilingual degradation: a zero-dependency script probe + (`unicodedata`-based histogram, `inputguard/language.py`) classifies + each input's script before rules run. Scripts without English + heuristic coverage take an explicit degraded path — rules are skipped, + a 20-point confidence penalty applies, and the result carries + `detected_language`, `heuristic_coverage`, and `degradation_note` + instead of silently scoring 100/ready. +- Additive `AnalysisResult` fields: `detected_language`, + `heuristic_coverage`, `degradation_note` (included in `to_dict()`). +- Mixed-script input: rules run whenever the covered share of letters is + at least 50%, with a `partial` note between 50–70% coverage. + ## [0.2.0] — 2026-05-29 ### Added - Auto intent detection. `.analyze()` now detects whether the input diff --git a/README.md b/README.md index 7bc578b..8c4d206 100644 --- a/README.md +++ b/README.md @@ -95,6 +95,35 @@ Use `warning` when you want to surface gaps to the user without blocking. Use `s --- +## Non-English input + +InputGuard's rules are English-language heuristics. Before any rule runs, a zero-dependency script probe (stdlib `unicodedata` only) classifies the input's script. When the script has no heuristic coverage, InputGuard says so instead of pretending: + +```python +result = guard.analyze("建造一个用户登录应用") + +result.status # 'usable_with_warnings' — never 'ready' +result.clarity_score # 80 (100 minus the degradation penalty) +result.detected_intent # 'undetermined' +result.detected_language # 'zh' (coarse, script-derived guess) +result.heuristic_coverage # 'none' +result.degradation_note # explains that rules were skipped and why +``` + +The rules are **skipped explicitly** on uncovered scripts — running English keyword rules on text they cannot read would produce a silent, unearned verdict. A degraded result is never `ready` in either mode (strict mode returns `needs_clarification`). In v0.2 this input silently scored 100/ready; v0.3 refuses to assert a confidence it does not have. + +Three additive fields on the result carry the probe's verdict: + +| Field | Values | +|---|---| +| `detected_language` | coarse script-derived guess (`'en'`, `'zh'`, `'ja'`, `'ko'`, `'ru'`, `'ar'`, ...; `'und'` when unclassifiable) | +| `heuristic_coverage` | `'full'` (≥ 70% of letters covered — rules run exactly as before), `'partial'` (50–70% — rules run, note flags the uncovered remainder), `'none'` (degraded path), `'unknown'` (no letters to classify) | +| `degradation_note` | `None`, or an explanation of what was skipped and why | + +Mixed input is handled by share, not by exclusion: `"make it faster 这个"` is still fully analyzed (English dominates and the rules run); input whose letters fall 50–70% inside coverage gets a `partial` note without a penalty. + +--- + ## How intent detection works InputGuard automatically detects what kind of coding input it is receiving. No extra parameters needed. The same `.analyze()` call handles all five intent types. diff --git a/inputguard/analyzer.py b/inputguard/analyzer.py index 23900f9..94074ca 100644 --- a/inputguard/analyzer.py +++ b/inputguard/analyzer.py @@ -5,6 +5,15 @@ import inputguard.rules # noqa: F401 — importing registers the coding domain and the 19 built-in rules from inputguard.detector import detect_intent, normalize from inputguard.followups import get_follow_ups +from inputguard.language import ( + COVERAGE_NONE, + COVERAGE_PARTIAL, + DEGRADATION_PENALTY, + DEGRADED_INTENT, + degradation_note_for, + partial_coverage_note, + probe_script, +) from inputguard.recommender import get_recommendations from inputguard.registry import REGISTRY from inputguard.scorer import calculate_score, get_status @@ -36,9 +45,32 @@ def analyze(self, user_input: str, domain: str = "coding") -> AnalysisResult: raise ValueError("user_input must be a non-empty, non-whitespace string.") domain_signals = REGISTRY.get_domain_signals(domain) + normalized = normalize(user_input) + + # Language probe (pipeline step 2): classify the input's script before + # any rule runs, so scripts outside heuristic coverage take the + # explicit degradation path instead of silently passing as ready. + probe = probe_script(normalized) + if probe.heuristic_coverage == COVERAGE_NONE: + # English-only rules are skipped outright — running them on a + # script they cannot read would produce a silent, unearned + # ready. The penalty keeps the result out of "ready" in both + # modes; the note says honestly what the tool does not know. + score = max(0, 100 - DEGRADATION_PENALTY) + return AnalysisResult( + status=get_status(score, self.mode), + clarity_score=score, + detected_intent=DEGRADED_INTENT, + gaps=[], + recommendations=[], + findings=[], + interpretation_note=None, + detected_language=probe.detected_language, + heuristic_coverage=probe.heuristic_coverage, + degradation_note=degradation_note_for(probe), + ) detected_intent = detect_intent(user_input, domain_signals) - normalized = normalize(user_input) findings: List[RuleFinding] = [] seen_codes: Set[str] = set() @@ -78,4 +110,11 @@ def analyze(self, user_input: str, domain: str = "coding") -> AnalysisResult: findings=findings, interpretation_note=interpretation_note, follow_ups=follow_ups, + detected_language=probe.detected_language, + heuristic_coverage=probe.heuristic_coverage, + degradation_note=( + partial_coverage_note(probe) + if probe.heuristic_coverage == COVERAGE_PARTIAL + else None + ), ) diff --git a/inputguard/language.py b/inputguard/language.py new file mode 100644 index 0000000..13be44e --- /dev/null +++ b/inputguard/language.py @@ -0,0 +1,288 @@ +"""Language probe: stdlib-only script detection for honest degradation. + +Every clarity rule in InputGuard is an English-keyword heuristic. On input in +another script (Chinese, Cyrillic, Arabic, ...) none of those keywords match, +so v0.2 silently returned ``ready`` at score 100 for input it had not assessed +at all — the tool asserted a confidence it did not have. + +This module is the v0.3 fix (spec: "Honest degradation for non-English +input"): a pure script histogram built exclusively on :mod:`unicodedata` — +no langdetect, no fastText, no network, honoring the zero-dependency +constraint. Each letter's Unicode *name* mentions its script (``"CJK +UNIFIED IDEOGRAPH-..."``, ``"CYRILLIC SMALL LETTER BE"``, ...), so scanning +name tokens classifies scripts with no hardcoded range tables to rot. + +When the dominant script has no heuristic coverage, ``analyze()`` takes the +explicit degraded path: English-only rules are skipped, a confidence penalty +is applied, and the result carries a ``degradation_note`` — never a silent +``ready``. Scripts with coverage (Latin) behave exactly as they did before +the probe existed. + +The probe is pure and thread-safe: no module-level mutation, only stdlib +lookups, matching the ``analyze()`` thread-safety contract. +""" + +from __future__ import annotations + +import re +import unicodedata +from dataclasses import dataclass +from typing import Dict, Optional + +__all__ = [ + "COVERAGE_FULL", + "COVERAGE_NONE", + "COVERAGE_PARTIAL", + "COVERAGE_UNKNOWN", + "COVERED_SCRIPTS", + "DEGRADATION_PENALTY", + "DEGRADED_INTENT", + "ScriptProbe", + "degradation_note_for", + "partial_coverage_note", + "probe_script", +] + +# Scripts the English keyword heuristics can actually assess. Every rule fires +# on English terms, so Latin script is the only fully covered script today. +COVERED_SCRIPTS = frozenset({"latin"}) + +# Coverage vocabulary for ``AnalysisResult.heuristic_coverage``: +# - "full": >= 70% of classified letters are in covered scripts; rules run +# exactly as before, no note. +# - "partial": 50-70% covered; rules run (the covered majority still drives +# detection), a note flags the uncovered remainder, no penalty. +# - "none": < 50% covered; degraded path — rules skipped, penalty, note. +# - "unknown": no letters at all (digits, punctuation, emoji) — nothing to +# classify; rules run as-is, no note. +COVERAGE_FULL = "full" +COVERAGE_PARTIAL = "partial" +COVERAGE_NONE = "none" +COVERAGE_UNKNOWN = "unknown" + +_FULL_COVERAGE_AT = 0.7 +_NONE_COVERAGE_BELOW = 0.5 + +# Confidence penalty applied on the degraded path (rules skipped). 20 points +# puts a degraded result at score 80: never ``ready`` (85) in either mode — +# ``usable_with_warnings`` in warning mode, ``needs_clarification`` in strict +# mode — so a degraded result can never masquerade as a clean bill of health. +# Policy-tunable in a later release; this constant is the default. +DEGRADATION_PENALTY = 20 + +# Intent reported for degraded results. Intent detection is English-keyword +# based, so on uncovered scripts it has no evidence; "undetermined" is an +# additive intent value that says so instead of guessing the "build" fallback. +DEGRADED_INTENT = "undetermined" + +# The probe examines a deterministic stride sample of at most this many +# characters, bounding probe cost on arbitrarily long input. +_SAMPLE_LIMIT = 4096 + +# Unicode name token -> script label. Name tokens are matched left-to-right, +# so the first script mention in the name wins ("FULLWIDTH LATIN CAPITAL +# LETTER A" -> "latin"; "KATAKANA-HIRAGANA PROLONGED SOUND MARK" -> +# "katakana"). +_SCRIPT_KEYWORDS = { + "latin": "latin", + "cjk": "han", + "ideograph": "han", + "han": "han", + "hiragana": "hiragana", + "katakana": "katakana", + "hangul": "hangul", + "cyrillic": "cyrillic", + "greek": "greek", + "arabic": "arabic", + "hebrew": "hebrew", + "syriac": "syriac", + "thaana": "thaana", + "devanagari": "devanagari", + "bengali": "bengali", + "gurmukhi": "gurmukhi", + "gujarati": "gujarati", + "oriya": "oriya", + "tamil": "tamil", + "telugu": "telugu", + "kannada": "kannada", + "malayalam": "malayalam", + "sinhala": "sinhala", + "thai": "thai", + "lao": "lao", + "tibetan": "tibetan", + "myanmar": "myanmar", + "georgian": "georgian", + "armenian": "armenian", + "ethiopic": "ethiopic", + "cherokee": "cherokee", + "mongolian": "mongolian", + "khmer": "khmer", +} + +# Script label -> coarse language guess. Script-level detection cannot pick +# between languages sharing a script (Han is written as Chinese *and* Japanese +# kanji); the ``ja`` refinement below handles the kana case, and the guess is +# documented as script-derived, never certain. +_SCRIPT_LANGUAGE = { + "latin": "en", + "han": "zh", + "hiragana": "ja", + "katakana": "ja", + "hangul": "ko", + "cyrillic": "ru", + "greek": "el", + "arabic": "ar", + "hebrew": "he", + "syriac": "syc", + "thaana": "dv", + "devanagari": "hi", + "bengali": "bn", + "gurmukhi": "pa", + "gujarati": "gu", + "oriya": "or", + "tamil": "ta", + "telugu": "te", + "kannada": "kn", + "malayalam": "ml", + "sinhala": "si", + "thai": "th", + "lao": "lo", + "tibetan": "bo", + "myanmar": "my", + "georgian": "ka", + "armenian": "hy", + "ethiopic": "am", + "cherokee": "chr", + "mongolian": "mn", + "khmer": "km", +} + + +@dataclass(frozen=True) +class ScriptProbe: + """Result of the script histogram over one input. + + ``heuristic_coverage`` is the field ``analyze()`` branches on; the other + fields feed the degradation note and the result's ``detected_language``. + """ + + detected_language: str + dominant_script: Optional[str] + covered_share: float + heuristic_coverage: str + + +def _sample(text: str, limit: int) -> str: + """Deterministic stride sample of at most ``limit`` characters. + + Striding (not truncating) keeps the histogram representative of the whole + input — a script buried at the end of a long paste is still seen. + """ + if len(text) <= limit: + return text + step = (len(text) + limit - 1) // limit # ceil division + return text[::step] + + +def _script_of(char: str) -> Optional[str]: + """Best-effort script label for one character, or ``None``. + + Non-letters (digits, punctuation, emoji, marks) and unnamed code points + carry no script evidence and return ``None``. + """ + if not unicodedata.category(char).startswith("L"): + return None + try: + name = unicodedata.name(char) + except ValueError: # unassigned / unnamed code point + return None + for token in re.findall(r"[a-z]+", name.lower()): + script = _SCRIPT_KEYWORDS.get(token) + if script is not None: + return script + return None + + +def probe_script(text: str) -> ScriptProbe: + """Classify an input's script coverage with a :mod:`unicodedata` histogram. + + Pure function: same input, same probe, no shared state — safe to call + from parallel ``analyze()`` threads. The classified population is the + letters whose script the probe recognizes; scripts outside + :data:`_SCRIPT_KEYWORDS` (Runic, Deseret, ...) leave the covered-share + denominator, which is the honest signal that heuristics do not apply. + """ + letter_count = 0 + script_counts: Dict[str, int] = {} + for char in _sample(text, _SAMPLE_LIMIT): + if not unicodedata.category(char).startswith("L"): + continue + letter_count += 1 + script = _script_of(char) + if script is not None: + script_counts[script] = script_counts.get(script, 0) + 1 + + classified = sum(script_counts.values()) + if classified == 0: + # Letters with no recognizable script, or no letters at all. + coverage = COVERAGE_NONE if letter_count > 0 else COVERAGE_UNKNOWN + return ScriptProbe( + detected_language="und", + dominant_script=None, + covered_share=0.0, + heuristic_coverage=coverage, + ) + + covered_share = script_counts.get("latin", 0) / classified + if covered_share >= _FULL_COVERAGE_AT: + coverage = COVERAGE_FULL + elif covered_share >= _NONE_COVERAGE_BELOW: + coverage = COVERAGE_PARTIAL + else: + coverage = COVERAGE_NONE + + # Deterministic dominant script: highest count, alphabetical tie-break. + dominant = min(script_counts, key=lambda s: (-script_counts[s], s)) + detected_language = _detected_language(dominant, script_counts) + + return ScriptProbe( + detected_language=detected_language, + dominant_script=dominant, + covered_share=covered_share, + heuristic_coverage=coverage, + ) + + +def _detected_language(dominant: str, script_counts: Dict[str, int]) -> str: + """Coarse script-derived language guess; "und" when nothing maps.""" + kana_present = script_counts.get("hiragana", 0) + script_counts.get("katakana", 0) + if dominant in ("han", "hiragana", "katakana") and kana_present: + return "ja" # kana alongside han (or alone) reads as Japanese + return _SCRIPT_LANGUAGE.get(dominant, "und") + + +def degradation_note_for(probe: ScriptProbe) -> str: + """The note carried by results on the degraded (uncovered) path.""" + if probe.dominant_script is None: + script_desc = "a script outside the probe's coverage" + else: + script_desc = f"the {probe.dominant_script} script" + return ( + f"InputGuard's clarity rules are English-language heuristics; this " + f"input appears to use {script_desc} (guessed language: " + f"{probe.detected_language}), which they do not cover. Rule analysis " + "was skipped rather than run with false confidence — this result " + "reports a language limitation of the tool, not a judgment of the " + "input's clarity. Rephrasing the key details in English enables a " + "full analysis." + ) + + +def partial_coverage_note(probe: ScriptProbe) -> str: + """The note carried when rules ran but a minority script is uncovered.""" + covered_pct = round(probe.covered_share * 100) + return ( + f"Input mixes scripts — about {covered_pct}% of its letters fall " + "within the English-heuristic coverage. Rules ran on the input as a " + "whole, so findings may miss content in the uncovered script(s)." + ) diff --git a/inputguard/types.py b/inputguard/types.py index 8dcdd5c..573046b 100644 --- a/inputguard/types.py +++ b/inputguard/types.py @@ -25,6 +25,13 @@ class AnalysisResult: # deduped and ordered with `gaps`. Appended after the v0.2 fields so any # positional construction keeps its meaning. follow_ups: List[str] = field(default_factory=list) + # v0.3 multilingual degradation (additive, see inputguard/language.py): + # the script probe's verdict. ``heuristic_coverage`` is "full", "partial", + # "none", or "unknown"; a non-None ``degradation_note`` marks results the + # rules could not fully assess. + detected_language: str = "en" + heuristic_coverage: str = "full" + degradation_note: Optional[str] = None def to_dict(self) -> dict: return { @@ -36,6 +43,9 @@ def to_dict(self) -> dict: "follow_ups": list(self.follow_ups), "findings": [asdict(f) for f in self.findings], "interpretation_note": self.interpretation_note, + "detected_language": self.detected_language, + "heuristic_coverage": self.heuristic_coverage, + "degradation_note": self.degradation_note, } def is_clear(self) -> bool: diff --git a/tests/test_followups.py b/tests/test_followups.py index fd74b0e..6e21a5d 100644 --- a/tests/test_followups.py +++ b/tests/test_followups.py @@ -282,7 +282,9 @@ def _ordered_unique_gaps(findings): def test_english_score_parity_with_v02(text): # follow_ups are derived from gaps after scoring; the v0.2 numbers must # not move: same findings, same gap order, same score as the legacy - # runners, and to_dict() gains exactly one additive key. + # runners, and to_dict() gains exactly the merged additive keys + # (follow_ups from the questions engine; the language-probe fields from + # the multilingual degradation work). result = InputGuard().analyze(text) legacy = _V02_INTENT_RUNNERS[result.detected_intent](text) @@ -291,7 +293,12 @@ def test_english_score_parity_with_v02(text): assert result.clarity_score == calculate_score(legacy) d = result.to_dict() - assert set(d) == _V02_TO_DICT_KEYS | {"follow_ups"} + assert set(d) == _V02_TO_DICT_KEYS | { + "follow_ups", + "detected_language", + "heuristic_coverage", + "degradation_note", + } v02_shaped = AnalysisResult( status=result.status, clarity_score=result.clarity_score, diff --git a/tests/test_language.py b/tests/test_language.py new file mode 100644 index 0000000..d235d0f --- /dev/null +++ b/tests/test_language.py @@ -0,0 +1,425 @@ +"""Tests for the language probe and the multilingual degradation path. + +Probe units cover script classification and coverage bands +(``inputguard.language``). The integration section covers the degraded +``analyze()`` path end to end: the probe-P2 regression (Chinese input +must never silently return ready/100), mode behavior, additive result +fields, English parity, systematic Unicode no-crash samples, and +thread-safe parallel analysis. +""" + +from __future__ import annotations + +import json +import unicodedata +from concurrent.futures import ThreadPoolExecutor + +import pytest + +from inputguard import InputGuard +from inputguard.language import ( + COVERAGE_FULL, + COVERAGE_NONE, + COVERAGE_PARTIAL, + COVERAGE_UNKNOWN, + DEGRADATION_PENALTY, + DEGRADED_INTENT, + _sample, + degradation_note_for, + partial_coverage_note, + probe_script, +) + +# Well-known probe-P2 input: pure Chinese, zero English heuristic coverage. +CHINESE = "建造一个用户登录应用" + + +def test_pure_english_is_full_coverage_en(): + probe = probe_script("build a rest api with users in postgresql") + assert probe.dominant_script == "latin" + assert probe.detected_language == "en" + assert probe.heuristic_coverage == COVERAGE_FULL + assert probe.covered_share == 1.0 + + +def test_pure_chinese_is_uncovered_han(): + probe = probe_script(CHINESE) + assert probe.dominant_script == "han" + assert probe.detected_language == "zh" + assert probe.heuristic_coverage == COVERAGE_NONE + assert probe.covered_share == 0.0 + + +def test_cyrillic_is_uncovered(): + probe = probe_script("почему моя программа не работает") + assert probe.dominant_script == "cyrillic" + assert probe.detected_language == "ru" + assert probe.heuristic_coverage == COVERAGE_NONE + + +def test_arabic_is_uncovered(): + probe = probe_script("لماذا لا يعمل هذا الكود") + assert probe.dominant_script == "arabic" + assert probe.detected_language == "ar" + assert probe.heuristic_coverage == COVERAGE_NONE + + +def test_hebrew_is_uncovered(): + probe = probe_script("למה הקוד הזה לא עובד") + assert probe.dominant_script == "hebrew" + assert probe.detected_language == "he" + assert probe.heuristic_coverage == COVERAGE_NONE + + +def test_greek_is_uncovered(): + probe = probe_script("γιατί δεν λειτουργεί") + assert probe.dominant_script == "greek" + assert probe.detected_language == "el" + assert probe.heuristic_coverage == COVERAGE_NONE + + +def test_korean_is_uncovered(): + probe = probe_script("이 코드가 왜 작동하지 않는 건가요") + assert probe.dominant_script == "hangul" + assert probe.detected_language == "ko" + assert probe.heuristic_coverage == COVERAGE_NONE + + +def test_japanese_kana_reads_as_ja(): + probe = probe_script("このコードが動作しないのはなぜですか") + assert probe.dominant_script in ("hiragana", "katakana", "han") + assert probe.detected_language == "ja" + assert probe.heuristic_coverage == COVERAGE_NONE + + +def test_fullwidth_latin_reads_as_latin(): + # Fullwidth Latin: the Unicode name's first token ("FULLWIDTH") is not + # the script; the token scan must find "latin". + probe = probe_script("API") + assert probe.dominant_script == "latin" + assert probe.heuristic_coverage == COVERAGE_FULL + + +def test_fullwidth_latin_mixed_with_japanese_is_partial(): + # 3 fullwidth latin letters vs 2 japanese letters -> exactly 0.5 share, + # the partial band boundary: rules still run, a note flags the rest. + probe = probe_script("APIを作る") + assert probe.dominant_script == "latin" + assert probe.heuristic_coverage == COVERAGE_PARTIAL + + +def test_digits_and_punctuation_have_no_coverage_signal(): + probe = probe_script("12345 !!! ??? ...") + assert probe.dominant_script is None + assert probe.detected_language == "und" + assert probe.heuristic_coverage == COVERAGE_UNKNOWN + + +def test_majority_english_with_minority_han_is_full(): + # 12 latin letters vs 2 han letters -> covered share well above 0.7. + probe = probe_script("make it faster 这个") + assert probe.dominant_script == "latin" + assert probe.heuristic_coverage == COVERAGE_FULL + + +def test_partial_band_input_is_between_full_and_none(): + # 5 latin letters vs 4 han letters -> covered share 5/9 (~0.56), inside + # the partial band (>= 0.5, < 0.7). + probe = probe_script("abcde 这是测试") + assert probe.heuristic_coverage == COVERAGE_PARTIAL + + +def test_unknown_script_letters_take_none_path(): + # Runic letters: real letters whose Unicode names the keyword table does + # not map — the honest classification is "no coverage", not a silent pass. + probe = probe_script("ᚠᚢᚦᚨᚱᚲ") + assert probe.dominant_script is None + assert probe.heuristic_coverage == COVERAGE_NONE + + +def test_probe_is_deterministic(): + text = "mixed 混合 input テキスト with 数字 123 and émojis 🏳️‍🌈" + assert probe_script(text) == probe_script(text) + + +def test_sample_bounds_length_and_is_deterministic(): + text = "a" * 100_000 + sample = _sample(text, 4096) + assert len(sample) <= 4096 + assert sample == _sample(text, 4096) + assert _sample("short", 4096) == "short" + + +def test_sample_lets_tail_scripts_surface(): + # A buried script at the very end of a huge paste must still be seen by + # the stride sample (stride 3 over 10 000 chars hits index 9999). + text = "a" * 9999 + "中" + probe = probe_script(text) + assert probe.dominant_script == "latin" + assert probe.covered_share < 1.0 + + +def test_degradation_penalty_never_allows_ready(): + # 100 - penalty must stay below the 85 ready floor in both modes. + assert 100 - DEGRADATION_PENALTY < 85 + + +def test_degraded_intent_value_is_undetermined(): + assert DEGRADED_INTENT == "undetermined" + + +def test_note_for_dominant_script_mentions_script_and_language(): + probe = probe_script(CHINESE) + note = degradation_note_for(probe) + assert "han" in note + assert "zh" in note + assert "English" in note + + +def test_note_for_unrecognized_script_names_the_limitation(): + probe = probe_script("ᚠᚢᚦᚨᚱᚲ") + note = degradation_note_for(probe) + assert "outside the probe's coverage" in note + + +def test_partial_note_reports_coverage_percentage(): + probe = probe_script("abcde 这是测试") + note = partial_coverage_note(probe) + assert "uncovered" in note + + +def test_script_probe_is_frozen(): + probe = probe_script(CHINESE) + try: + probe.detected_language = "en" + raised = False + except Exception: + raised = True + assert raised + + +def test_non_letters_are_skipped_by_category(): + # Digits, punctuation, emoji, combining marks, and control characters + # contribute nothing to the histogram — only letters classify. + for ch in "0九!!。🏳️‍🌈́\u0000": + assert unicodedata.category(ch) # sanity: all real code points + probe = probe_script("... 123 🏳️‍🌈") + assert probe.heuristic_coverage == COVERAGE_UNKNOWN + + +# --------------------------------------------------------------------------- +# Integration: the degraded analyze() path (non-English honesty) +# --------------------------------------------------------------------------- + +# Probe P2's exact input: v0.2 returned intent=build, score 100, ready, +# zero findings — a silent pass on input the rules cannot read. +PROBE_P2_INPUT = "建造一个用户登录应用" + + +def test_probe_p2_regression_chinese_never_silent_ready(): + guard = InputGuard() + r = guard.analyze(PROBE_P2_INPUT) + assert r.heuristic_coverage == COVERAGE_NONE + assert r.degradation_note is not None + assert r.status != "ready" + assert r.clarity_score != 100 + assert not r.is_clear() + + +def test_degraded_result_shape(): + r = InputGuard().analyze(PROBE_P2_INPUT) + assert r.clarity_score == 100 - DEGRADATION_PENALTY + assert r.detected_intent == DEGRADED_INTENT + assert r.detected_language == "zh" + assert r.gaps == [] + assert r.findings == [] + assert r.recommendations == [] + assert r.interpretation_note is None + + +def test_degraded_strict_mode_returns_needs_clarification_not_blocked(): + r = InputGuard(mode="strict").analyze(PROBE_P2_INPUT) + # Score 80 lands in the strict clarify band (65-84), below the ready + # floor in both modes. + assert r.status == "needs_clarification" + + +def test_degraded_warning_mode_returns_usable_with_warnings(): + r = InputGuard().analyze(PROBE_P2_INPUT) + assert r.status == "usable_with_warnings" + + +def test_degradation_note_is_actionable(): + r = InputGuard().analyze(PROBE_P2_INPUT) + assert "English" in r.degradation_note + assert "han" in r.degradation_note + + +def test_degraded_applies_to_every_uncovered_script(): + for text in ( + "почему моя программа не работает", + "لماذا لا يعمل هذا الكود", + "이 코드가 왜 작동하지 않나요", + "このコードが動作しないのはなぜですか", + "γιατί δεν λειτουργεί", + "ᚠᚢᚦᚨᚱᚲ", + ): + r = InputGuard().analyze(text) + assert r.heuristic_coverage == COVERAGE_NONE, text + assert r.degradation_note is not None, text + assert r.status != "ready", text + + +def test_english_parity_probe_fields(): + # Covered input carries the probe's verdict but no degradation. The + # clarity verdict itself is unchanged v0.2 behavior: one distinct gap + # (api structure) -> 100 - 25 = 75. + r = InputGuard().analyze("Build a REST API using FastAPI. Store users in PostgreSQL.") + assert r.detected_language == "en" + assert r.heuristic_coverage == COVERAGE_FULL + assert r.degradation_note is None + assert r.status == "usable_with_warnings" + assert r.clarity_score == 75 + + +def test_english_parity_score_unchanged(): + # Byte-parity with the pre-probe pipeline on a vague English input: + # debug intent, three distinct gaps -> 100 - 25 - 25 - 15 = 35 (the + # score the v0.2 README documents for this input). + r = InputGuard().analyze("fix my code") + assert r.detected_intent == "debug" + assert r.clarity_score == 35 + assert r.status == "needs_clarification" + assert r.heuristic_coverage == COVERAGE_FULL + assert r.degradation_note is None + + +def test_partial_coverage_runs_rules_with_note_no_penalty(): + # ~56% covered share: rules run (findings computed as usual), the note + # flags the uncovered remainder, and no degradation penalty applies. + r = InputGuard().analyze("abcde 这是测试") + assert r.heuristic_coverage == COVERAGE_PARTIAL + assert r.degradation_note is not None + assert r.clarity_score == 100 # 5-letter input fires no rules + + +def test_validation_contracts_precede_the_probe(): + guard = InputGuard() + with pytest.raises(TypeError): + guard.analyze(12345) # type: ignore[arg-type] + with pytest.raises(ValueError): + guard.analyze(" ") + # Unknown domain raises even for input the probe would degrade. + with pytest.raises(ValueError): + guard.analyze(PROBE_P2_INPUT, domain="legal") + + +def test_to_dict_includes_additive_fields(): + d = InputGuard().analyze(PROBE_P2_INPUT).to_dict() + assert d["detected_language"] == "zh" + assert d["heuristic_coverage"] == "none" + assert isinstance(d["degradation_note"], str) + # Merged additive contract: v0.2 keys keep their names and relative + # order, follow_ups slots in after recommendations (questions engine), + # and the three probe keys trail. + assert list(d) == [ + "status", + "clarity_score", + "detected_intent", + "gaps", + "recommendations", + "follow_ups", + "findings", + "interpretation_note", + "detected_language", + "heuristic_coverage", + "degradation_note", + ] + json.dumps(d, ensure_ascii=False) # JSON-serializable as before + + +# --------------------------------------------------------------------------- +# Property-style no-crash on arbitrary Unicode (systematic samples — +# hypothesis is not a dev dependency; zero-dep constraint honored) +# --------------------------------------------------------------------------- + +_VALID_STATUSES = {"ready", "usable_with_warnings", "needs_clarification", "blocked"} + +_UNICODE_SAMPLES = [ + "emoji only 🚀🔥🏳️‍🌈", + "mixed 混合 text وبالعربية معا", + "zero width zero​width joiner", + "rtl override ‮reverse‭", + "combining marks é̈ 👨‍👩‍👧‍👦", + "unassigned \u0378 codepoint", + "cjk mixed with ascii: build api 用户", + "fullwidth letters 123", + "tab\tand\nnewline\r\nmixes", + "ʼn concatenations ʻʼʽ", + "íàéä ççñň ņņň — diacritic soup", + "文字化け mojibake", + "سيب ذلك mixed rtl ltr", + "҈ all the combining things ✈ ✈ ✈", +] + + +@pytest.mark.parametrize("text", _UNICODE_SAMPLES) +def test_no_crash_valid_result_on_arbitrary_unicode(text): + r = InputGuard().analyze(text) + assert r.status in _VALID_STATUSES + assert 0 <= r.clarity_score <= 100 + # A note exists exactly when the probe could not fully cover the input. + if r.heuristic_coverage in (COVERAGE_FULL, COVERAGE_UNKNOWN): + assert r.degradation_note is None + else: + assert r.degradation_note is not None + if r.heuristic_coverage == COVERAGE_NONE: + # Degraded path contract: rules skipped, honest intent, never ready. + assert r.findings == [] + assert r.detected_intent == DEGRADED_INTENT + assert r.status != "ready" + + +def test_long_input_bounded_probe_still_classifies(): + # ~100k chars: the stride sample keeps the probe bounded while still + # classifying correctly. The clarity verdict is unchanged v0.2 + # behavior for the repeated build request (missing language + api + # structure -> 100 - 25 - 25 = 50). + r = InputGuard().analyze("build a rest api " * 6000) + assert r.status == "needs_clarification" + assert r.clarity_score == 50 + assert r.heuristic_coverage == COVERAGE_FULL + assert r.degradation_note is None + + +def test_parallel_analyze_thread_safety_64_workers(): + # 64 workers over mixed English/Chinese inputs must agree with the + # single-threaded results (the v0.2 thread-safety contract, rerun with + # the probe in the pipeline). + guard = InputGuard() + strict = InputGuard(mode="strict") + inputs = [ + "fix my code", + PROBE_P2_INPUT, + "Build a REST API using FastAPI. Store users in PostgreSQL.", + "почему моя программа не работает", + "make it faster 这个", + ] + expected = {text: guard.analyze(text).to_dict() for text in inputs} + expected_strict = {PROBE_P2_INPUT: strict.analyze(PROBE_P2_INPUT).to_dict()} + + def run(pair): + text, mode = pair + target = strict if mode == "strict" else guard + return text, mode, target.analyze(text).to_dict() + + jobs = [(text, "warning") for text in inputs for _ in range(12)] + [ + (PROBE_P2_INPUT, "strict") for _ in range(4) + ] + assert len(jobs) == 64 + with ThreadPoolExecutor(max_workers=16) as pool: + results = list(pool.map(run, jobs)) + + for text, mode, result in results: + baseline = expected_strict if mode == "strict" else expected + assert result == baseline[text] From c52f4859383d15e0397e24180d9f516efd97774b Mon Sep 17 00:00:00 2001 From: "obvious-autobuild[bot]" <262744130+obvious-autobuild[bot]@users.noreply.github.com> Date: Thu, 17 Sep 2026 19:53:18 +0000 Subject: [PATCH 05/18] feat(eval-prompt-set): version the 121-case clarity-evaluation set as eval/ (#7) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * feat(eval-prompt-set): version the 121-case clarity-evaluation set as eval/ Checkpoint the labeled evaluation set (workbook wb_q30ippCW + labeling guide art_XPvHhPeZ) into the repo: cases.csv (121 rows, QUOTE_ALL/CRLF), README with the labeling criteria, per-case_type rubric, gap-vocabulary calibration contract, and a stdlib-only measure_fp.py that runs the analyzer over every row and reports per-case_type and overall FP/FN rates. The tool is a measurement, not a gate: it always exits 0 and treats label/behavior disagreement as the data. Co-authored-by: Kalisetti Nihanth Naidu * test(eval-loader): smoke-test the eval dataset mechanics Three smoke tests: the CSV parses with all nine required columns and non-empty texts, case ids are unique, and measure_fp.py runs end to end on a 10-case sample. Deliberately no label-vs-analyzer assertions — that agreement is what the measurement tool reports over time. Co-authored-by: Kalisetti Nihanth Naidu --------- Co-authored-by: Obvious Co-authored-by: Kalisetti Nihanth Naidu --- eval/README.md | 175 ++++++++++++++++++++ eval/cases.csv | 122 ++++++++++++++ eval/measure_fp.py | 324 ++++++++++++++++++++++++++++++++++++++ tests/test_eval_loader.py | 72 +++++++++ 4 files changed, 693 insertions(+) create mode 100644 eval/README.md create mode 100644 eval/cases.csv create mode 100644 eval/measure_fp.py create mode 100644 tests/test_eval_loader.py diff --git a/eval/README.md b/eval/README.md new file mode 100644 index 0000000..5cc1d50 --- /dev/null +++ b/eval/README.md @@ -0,0 +1,175 @@ +# eval/ — clarity evaluation set + +A versioned checkpoint of the InputGuard v0.3 labeled clarity-evaluation set: +121 prompts with expected intent, gap set, status, and per-gap severities, +plus the labeling criteria and a stdlib-only measurement script. + +Source of record: the "InputGuard v0.3 clarity-evaluation set" workbook +(sheets `cases`, `gap_vocabulary`) and its labeling guide in the Obvious +project. This directory is the git checkpoint of that dataset — it is what +the release-wave `docs/false-positive-benchmark.md` procedure consumes, and +what keeps the labels reviewable alongside the code they calibrate. + +## Files + +| File | What it is | +| --------------- | --------------------------------------------------------- | +| `cases.csv` | All 121 labeled rows (schema below). | +| `measure_fp.py` | Runs the analyzer over every row, reports FP/FN rates. | +| `README.md` | This document: labeling criteria, rubric, how to measure. | + +## `cases.csv` schema + +Nine columns, UTF-8, all fields double-quoted, CRLF line endings: + +| Column | Meaning | +| --------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `id` | Stable row id: `TP-…`, `TN-…`, `BD-…`, `DG-…`, `PF-…`. Unique across the file. | +| `text` | The prompt as a user would send it. | +| `domain` | `coding`, `writing`, or `data-analysis` (the v0.3 first-party domains). | +| `expected_intent` | Intent the analyzer must detect; `n/a` when the label asserts no intent (all degradation/performance rows, plus writing/data-analysis rows, which carry no intent machinery yet). | +| `expected_gaps` | `;`-separated gap strings from the vocabulary below; `none` when no gap may fire. | +| `expected_status` | `ready`, `usable_with_warnings`, `needs_clarification`, or `degraded`. | +| `expected_severities` | `;`-separated, positionally aligned with `expected_gaps`; `none` when there are no gaps. | +| `case_type` | `true_positive`, `true_negative`, `boundary`, `degradation`, `performance`. | +| `labeler_note` | Why the row is labeled the way it is. Read it before re-labeling or "fixing" a row. | + +Row counts: 43 true positives, 37 true negatives, 25 boundary, 14 +degradation, 2 performance. Domains: 92 coding, 14 writing, 15 +data-analysis. + +## How rows were labeled + +- **Coding rows** (`true_positive`, `true_negative`, `boundary`): + `expected_gaps` is the set of registry rules whose satisfaction signals + the text genuinely lacks, using the 19 built-in rules plus the catch-all + (`task context`). A true negative has zero such gaps by construction. + Gap strings use the coding vocabulary verbatim (see table below). +- **Writing rows** use the spec's first-party vocabulary verbatim: + `audience`, `purpose`, `structure/format`, `source material`, `context`, + `completeness`. **Data-analysis rows** use `dataset/source`, + `question/goal`, `output format`, `tooling`, `volume`, `reproducibility`. +- **Degradation rows** (`DG-001`…`DG-014`) cover Chinese, Japanese, Arabic, + Hindi, Cyrillic, accented Latin, and one mixed English+Han prompt. They + assert a non-null degradation note and a non-ready outcome — not an + ordinary status band. `expected_status` is `degraded`, + `expected_intent` is `n/a`. +- **Performance rows**: `PF-001` is 9,741 characters (under the 10,000-char + cap), `PF-002` is 10,533 (over it). Both are filler with no rule signals, + so the expected outcome isolates size handling from clarity scoring. + +## Pass rubric per case_type + +A row **passes** when the analyzer returns exactly: the expected intent +(when asserted), the expected status, and the expected gap set +(order-insensitive; no extra gaps, none missing). + +- **true_positive** — must flag: status below `ready`, correct intent, + expected gaps. The maximum severity per gap must match + `expected_severities`. +- **true_negative** — must return `ready` with zero gaps and correct + intent. Any flag on these rows is a false positive. +- **boundary** — read `labeler_note` first; it states which side the row + guards. Fixture-style rows (BD-002/003/004/006/013/014) must return + `ready` under word-boundary matching; under v0.2 substring matching all + six flag — that is the regression being killed. BD-012 is the + no-regression guard: a genuine debug request containing "fixture" must + stay `intent=debug` with one medium `code context` gap at `ready`. + BD-010 probes the catch-all: v0.2 misses it; v0.3 must report one high + `task context` gap. +- **degradation** — the analyzer must return an explicit degradation note + and NOT a `ready` status. Rule checks must not fire on Latin-only + vocabulary assumptions. +- **performance** — `PF-001` must complete normally within the latency + budget; `PF-002` must exercise the over-cap path (truncate-with-flag or + reject) with no crash and bounded time. Wall time is recorded per row. + +## Measuring + +```bash +python3 eval/measure_fp.py # full 121-row run, human-readable table +python3 eval/measure_fp.py --limit 10 # smoke run (used by tests/test_eval_loader.py) +python3 eval/measure_fp.py --json # machine-readable summary + per-case detail +``` + +Requires Python 3.9+ and nothing else — the script imports only the stdlib +and `inputguard` itself, resolved from the repo checkout (no install +needed). It always exits 0 after printing its summary: it is a +data-measurement tool, not a CI assertion. Only a broken run (missing or +malformed cases file) exits 2. + +Definitions used by `measure_fp.py`: + +- A row **matches** when actual status equals `expected_status`, actual + intent equals `expected_intent` (rows labeled `n/a` skip the intent + check), the actual gap set equals the expected set order-insensitively, + and every shared gap's maximum severity matches. +- **False positive (row-level)** — the analyzer flagged a gap the label + does not list. Includes every flagged true negative. +- **False negative (row-level)** — a labeled gap the analyzer failed to + flag. A row the analyzer cannot evaluate at all (for example a domain + whose rules are not registered yet, so `analyze()` raises) counts every + expected gap as missed and is reported as unevaluated. +- Rates are reported per `case_type` and overall as + `false-positive rows ÷ rows in group` (and likewise for false + negatives). Degradation-note presence is reported separately from the + gap-based rates, and performance rows print their wall time. + +Interpreting results: **labels may legitimately disagree with current +analyzer behavior — that disagreement is the measurement.** A rising match +rate over the release wave is the signal that the registry, word-boundary +matching, and degradation work are landing. Never edit a label to make a +measurement look better; if a row is genuinely mislabeled, fix it in a PR +with the `labeler_note` updated to say why. + +## Gap vocabulary (calibration contract) + +Per-gap default severities and the satisfied-condition examples a text must +contain to avoid the flag. The `gap_vocabulary` sheet in the source workbook +is the normative version; this table is its checkpoint. For the new-domain +gaps the examples double as the implementation contract: a text containing +any listed example must not flag that gap; a text containing none must. + +| Domain | Gap | Severity | Rule code | Satisfied when the text contains… | +| ------------- | --------------------------- | -------- | ----------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------- | +| coding | programming language | high | missing_language, intent_without_language | v0.2 gap string (compat-pinned) | +| coding | api structure | high | missing_api_structure | v0.2 gap string (compat-pinned) | +| coding | data model | high | missing_data_model | v0.2 gap string (compat-pinned) | +| coding | integration specifics | medium | missing_integration_specifics | v0.2 gap string (compat-pinned) | +| coding | authentication type | high | missing_auth_type | v0.2 gap string (compat-pinned) | +| coding | output format | medium | missing_output_format | v0.2 gap string (compat-pinned); build verbs only | +| coding | task context | high | insufficient_context | v0.2 gap string (compat-pinned); catch-all safety net | +| coding | error description | high | missing_error_message | v0.2 gap string (compat-pinned) | +| coding | expected vs actual behavior | high | missing_expected_vs_actual | v0.2 gap string (compat-pinned) | +| coding | code context | medium | missing_debug_code_context | v0.2 gap string (compat-pinned) | +| coding | optimization target | high | missing_optimization_target | v0.2 gap string (compat-pinned) | +| coding | performance baseline | medium | missing_performance_baseline | v0.2 gap string (compat-pinned) | +| coding | optimization constraint | low | missing_optimization_constraint | v0.2 gap string (compat-pinned) | +| coding | code reference | high | missing_code_reference | v0.2 gap string (compat-pinned) | +| coding | explanation depth | low | missing_explanation_depth | v0.2 gap string (compat-pinned) | +| coding | existing stack | high | missing_existing_stack | v0.2 gap string (compat-pinned) | +| coding | feature scope | high | missing_feature_scope | v0.2 gap string (compat-pinned) | +| coding | completion criteria | low | missing_completion_criteria | v0.2 gap string (compat-pinned) | +| writing | audience | high | (new v0.3 rule) | readership named: 'for beginners', 'my manager', 'the board', 'engineers', 'admissions officers', 'stakeholders' | +| writing | purpose | high | (new v0.3 rule) | goal stated: 'to convince', 'announcing', 'requesting', 'asking for', 'goal is to', 'so that', 'recap', 'lead with' | +| writing | structure/format | medium | (new v0.3 rule) | length or organization: '300 words', 'one page', 'bullets', 'sections', 'short', 'concise'; bare artifact type does NOT satisfy | +| writing | source material | high | (new v0.3 rule) | fires only when the task references existing material; satisfied when the material is actually provided | +| writing | context | medium | (new v0.3 rule) | subject or situation named: 'remote work', 'the migration', 'Q3 roadmap', 'context:', 'background:', 'the vendor' | +| writing | completeness | low | (new v0.3 rule) | required content enumerated: 'include', 'must cover', 'mention', 'must stay under 650 words', 'owners and deadlines' | +| data-analysis | dataset/source | high | (new v0.3 rule) | concrete data named: 'sales.csv', 'the attached export', 'postgres subscriptions table'; 'this dataset' alone does NOT satisfy | +| data-analysis | question/goal | high | (new v0.3 rule) | explicit question or objective: 'whether refunds spiked', 'which tier cancels most', 'I want to know', or a trailing '?' | +| data-analysis | output format | medium | (new v0.3 rule) | deliverable shape: 'bar chart', 'heatmap', 'two-slide summary', 'trend table', 'SQL to reproduce', 'themes table' | +| data-analysis | tooling | medium | (new v0.3 rule) | tool constraint: 'using pandas only', 'in python', 'dbt', 'BigQuery only', 'no external libraries' | +| data-analysis | volume | low | (new v0.3 rule) | scale stated: '80,000 rows', '400M rows', '5,000 responses', '90 days of data', '18 months of daily data' | +| data-analysis | reproducibility | low | (new v0.3 rule) | rerun/reuse concern: 'rerun weekly', 'reusable', 'rerunnable', 'reproduce it', 'monthly' | + +## Maintenance + +- Add rows via PR. Keep the id prefix and uniqueness, gap strings verbatim + from the vocabulary, and always fill `labeler_note`. +- Regenerate `cases.csv` only from the source workbook (scripted export, + QUOTE_ALL + CRLF), never by hand-editing quoted text. +- `tests/test_eval_loader.py` guards the file's mechanics (it parses, + columns exist, ids unique, the measurement runs end to end on a sample). + It deliberately does NOT assert that labels match analyzer output — that + agreement is what `measure_fp.py` measures over time. diff --git a/eval/cases.csv b/eval/cases.csv new file mode 100644 index 0000000..a4c197a --- /dev/null +++ b/eval/cases.csv @@ -0,0 +1,122 @@ +"id","text","domain","expected_intent","expected_gaps","expected_status","expected_severities","case_type","labeler_note" +"TN-DBG-01","getting a TypeError in my Python function, expected a list but got None","coding","debug","none","ready","none","true_negative","Well-specified debug request: every trigger-and-satisfy pair for this intent resolves, so no finding may fire and status must be ready at 100." +"TN-DBG-02","my express endpoint fails, the log says 'CastError: invalid ObjectId', it should return a 400 instead of a 500","coding","debug","none","ready","none","true_negative","Well-specified debug request: every trigger-and-satisfy pair for this intent resolves, so no finding may fire and status must be ready at 100." +"TN-DBG-03","the django view crashes with a ValueError when the form is empty; expected it to redirect back with an error message, code is in views.py","coding","debug","none","ready","none","true_negative","Well-specified debug request: every trigger-and-satisfy pair for this intent resolves, so no finding may fire and status must be ready at 100." +"TN-DBG-04","my python script crashes on line 42 with a KeyError 'user_id', it should skip missing keys gracefully instead of crashing","coding","debug","none","ready","none","true_negative","Well-specified debug request: every trigger-and-satisfy pair for this intent resolves, so no finding may fire and status must be ready at 100." +"TN-DBG-05","the go service hit a segfault under load, stack trace in the gist, it worked fine until the last deploy, expected zero regressions","coding","debug","none","ready","none","true_negative","Well-specified debug request: every trigger-and-satisfy pair for this intent resolves, so no finding may fire and status must be ready at 100." +"TN-OPT-01","speed up this function, it currently takes 3 seconds, must stay backward compatible","coding","optimization","none","ready","none","true_negative","Well-specified optimization request: every trigger-and-satisfy pair for this intent resolves, so no finding may fire and status must be ready at 100." +"TN-OPT-02","optimize the database query in my reporting endpoint, profiling shows it takes 2.5s today, keep the result ordering identical","coding","optimization","none","ready","none","true_negative","Well-specified optimization request: every trigger-and-satisfy pair for this intent resolves, so no finding may fire and status must be ready at 100." +"TN-OPT-03","reduce the memory usage of this module, it maxes out at 4GB right now, without breaking the streaming API contract","coding","optimization","none","ready","none","true_negative","Well-specified optimization request: every trigger-and-satisfy pair for this intent resolves, so no finding may fire and status must be ready at 100." +"TN-OPT-04","improve performance of this endpoint, benchmark before/after, acceptable tradeoff: slightly higher CPU","coding","optimization","none","ready","none","true_negative","Well-specified optimization request: every trigger-and-satisfy pair for this intent resolves, so no finding may fire and status must be ready at 100." +"TN-OPT-05","cut the cold-start latency of this function, it currently takes 1.2s, must still work offline","coding","optimization","none","ready","none","true_negative","Well-specified optimization request: every trigger-and-satisfy pair for this intent resolves, so no finding may fire and status must be ready at 100." +"TN-EXP-01","explain this function line by line, briefly","coding","explanation","none","ready","none","true_negative","Well-specified explanation request: every trigger-and-satisfy pair for this intent resolves, so no finding may fire and status must be ready at 100." +"TN-EXP-02","walk me through the following promise chain, in simple terms with an example","coding","explanation","none","ready","none","true_negative","Well-specified explanation request: every trigger-and-satisfy pair for this intent resolves, so no finding may fire and status must be ready at 100." +"TN-EXP-03","describe how this decorator works at a high level","coding","explanation","none","ready","none","true_negative","Well-specified explanation request: every trigger-and-satisfy pair for this intent resolves, so no finding may fire and status must be ready at 100." +"TN-EXP-04","break down the recursion in the above function, step by step","coding","explanation","none","ready","none","true_negative","Well-specified explanation request: every trigger-and-satisfy pair for this intent resolves, so no finding may fire and status must be ready at 100." +"TN-EXP-05","help me understand this pattern in the codebase, eli5","coding","explanation","none","ready","none","true_negative","Well-specified explanation request: every trigger-and-satisfy pair for this intent resolves, so no finding may fire and status must be ready at 100." +"TN-FEA-01","extend my express api with pagination so users can page through results, done when it works with existing page params","coding","feature","none","ready","none","true_negative","Well-specified feature request: every trigger-and-satisfy pair for this intent resolves, so no finding may fire and status must be ready at 100." +"TN-FEA-02","add full text search to my existing django app using postgres, users can search by title and author, complete means results appear as you type","coding","feature","none","ready","none","true_negative","Well-specified feature request: every trigger-and-satisfy pair for this intent resolves, so no finding may fire and status must be ready at 100." +"TN-FEA-03","implement in my vue component an upload preview, once a file is selected it should show a thumbnail","coding","feature","none","ready","none","true_negative","Well-specified feature request: every trigger-and-satisfy pair for this intent resolves, so no finding may fire and status must be ready at 100." +"TN-FEA-04","integrate into my current node codebase a role-based permission layer; the feature should support admin, editor and viewer roles, finished means all routes enforce it","coding","feature","none","ready","none","true_negative","Well-specified feature request: every trigger-and-satisfy pair for this intent resolves, so no finding may fire and status must be ready at 100." +"TN-FEA-05","extend the react dashboard with a chart export button, specifically PNG and CSV export, it should respect the current filters","coding","feature","none","ready","none","true_negative","Well-specified feature request: every trigger-and-satisfy pair for this intent resolves, so no finding may fire and status must be ready at 100." +"TN-BLD-01","build a python CLI that renames files by EXIF date, output goes to a folder tree","coding","build","none","ready","none","true_negative","Well-specified build request: every trigger-and-satisfy pair for this intent resolves, so no finding may fire and status must be ready at 100." +"TN-BLD-02","create a fastapi POST endpoint script that returns user records as a JSON body","coding","build","none","ready","none","true_negative","Well-specified build request: every trigger-and-satisfy pair for this intent resolves, so no finding may fire and status must be ready at 100." +"TN-BLD-03","write a node script that records each stripe charge into our postgres charges table, keyed by email","coding","build","none","ready","none","true_negative","Well-specified build request: every trigger-and-satisfy pair for this intent resolves, so no finding may fire and status must be ready at 100." +"TN-BLD-04","set up user login for the react web app using JWT stored in an httpOnly cookie","coding","build","none","ready","none","true_negative","Well-specified build request: every trigger-and-satisfy pair for this intent resolves, so no finding may fire and status must be ready at 100." +"TN-BLD-05","generate a python script that walks a directory tree and produces an index.json with file sizes","coding","build","none","ready","none","true_negative","Well-specified build request: every trigger-and-satisfy pair for this intent resolves, so no finding may fire and status must be ready at 100." +"TP-DBG-01","my app crashes when I click the submit button","coding","debug","error description; expected vs actual behavior; code context","needs_clarification","high; high; medium","true_positive","Gap-bearing debug request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-DBG-02","this function fails intermittently in production","coding","debug","error description; expected vs actual behavior","needs_clarification","high; high","true_positive","Gap-bearing debug request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-DBG-03","it throws an exception, here is the traceback, but I don't know why","coding","debug","expected vs actual behavior; code context","usable_with_warnings","high; medium","true_positive","Gap-bearing debug request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-DBG-04","why does my react component crash on mount","coding","debug","error description; expected vs actual behavior","needs_clarification","high; high","true_positive","Gap-bearing debug request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-DBG-05","pytest fails with a weird exception on CI only","coding","debug","error description; expected vs actual behavior; code context","needs_clarification","high; high; medium","true_positive","Gap-bearing debug request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-DBG-06","the websocket connection is broken, I pasted the log below, it should stay connected for hours","coding","debug","error description; code context","usable_with_warnings","high; medium","true_positive","Gap-bearing debug request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-OPT-01","make this faster","coding","optimization","optimization target; performance baseline; optimization constraint","needs_clarification","high; medium; low","true_positive","Gap-bearing optimization request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-OPT-02","optimize the database query in my django app","coding","optimization","performance baseline; optimization constraint","usable_with_warnings","medium; low","true_positive","Gap-bearing optimization request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-OPT-03","my api is too slow, users are complaining","coding","optimization","optimization target; optimization constraint","usable_with_warnings","high; low","true_positive","Gap-bearing optimization request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-OPT-04","reduce memory usage in the worker process","coding","optimization","optimization target; performance baseline; optimization constraint","needs_clarification","high; medium; low","true_positive","Gap-bearing optimization request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-OPT-05","refactoring my rust parser for efficiency, it parses 1MB files in 400ms right now, cannot change the output format","coding","optimization","optimization target","usable_with_warnings","high","true_positive","Gap-bearing optimization request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-OPT-06","our build pipeline is heavy, the whole thing takes 40 minutes and we want it under 10 without breaking cache correctness","coding","optimization","optimization target","usable_with_warnings","high","true_positive","Gap-bearing optimization request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-EXP-01","explain this code","coding","explanation","code reference; explanation depth","usable_with_warnings","high; low","true_positive","Gap-bearing explanation request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-EXP-02","how does the retry logic work","coding","explanation","code reference; explanation depth","usable_with_warnings","high; low","true_positive","Gap-bearing explanation request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-EXP-03","clarify the difference between these two approaches","coding","explanation","code reference; explanation depth","usable_with_warnings","high; low","true_positive","Gap-bearing explanation request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-EXP-04","explain the event loop in depth with an example","coding","explanation","code reference","usable_with_warnings","high","true_positive","Gap-bearing explanation request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-EXP-05","give me a breakdown of how the router picks a handler","coding","explanation","code reference; explanation depth","usable_with_warnings","high; low","true_positive","Gap-bearing explanation request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-EXP-06","I don't understand how our rate limiter counts requests","coding","explanation","code reference; explanation depth","usable_with_warnings","high; low","true_positive","Gap-bearing explanation request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-FEA-01","add dark mode to my app","coding","feature","existing stack; feature scope; completion criteria","needs_clarification","high; high; low","true_positive","Gap-bearing feature request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-FEA-02","implement in my react dashboard a filtering feature","coding","feature","feature scope; completion criteria","usable_with_warnings","high; low","true_positive","Gap-bearing feature request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-FEA-03","extend the cli with an export command","coding","feature","existing stack; completion criteria","usable_with_warnings","high; low","true_positive","Gap-bearing feature request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-FEA-04","add upload support to my project","coding","feature","existing stack; completion criteria","usable_with_warnings","high; low","true_positive","Gap-bearing feature request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-FEA-05","hook into my existing flask app and add notifications","coding","feature","feature scope; completion criteria","usable_with_warnings","high; low","true_positive","Gap-bearing feature request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-FEA-06","add a settings page to my app where users can toggle notifications, using vue","coding","feature","feature scope","usable_with_warnings","high","true_positive","Gap-bearing feature request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-BLD-01","build me a todo app","coding","build","programming language; output format","usable_with_warnings","high; medium","true_positive","Gap-bearing build request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-BLD-02","create a REST api for managing invoices","coding","build","programming language; api structure","needs_clarification","high; high","true_positive","Gap-bearing build request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-BLD-03","set up a postgres database to store customer orders with an order id, total and created_at","coding","build","programming language; output format","usable_with_warnings","high; medium","true_positive","Gap-bearing build request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-BLD-04","I need a script that reads a CSV and emails me a summary","coding","build","programming language","usable_with_warnings","high","true_positive","Gap-bearing build request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-BLD-05","integrate Stripe into the checkout","coding","build","programming language; integration specifics","usable_with_warnings","high; medium","true_positive","Gap-bearing build request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"TP-BLD-06","connect to the company Postgres database and pull the user records for the audit","coding","build","programming language","usable_with_warnings","high","true_positive","Gap-bearing build request; the listed gaps must fire (distinct-gap dedup applies) and warning-mode status must land in the labeled band." +"BD-001","I love this fixture in the test suite, what does it do","coding","explanation","code reference; explanation depth","usable_with_warnings","high; low","boundary","Spec verification row 'False positive eliminated' (probe P1). v0.2 reads 'fix' inside 'fixture' and answers debug/35. Expected v0.3: intent must NOT be debug; word-boundary matching leaves 'fixture' unmatched; the explanation path (score 70) is the faithful post-fix outcome." +"BD-002","our test fixtures load slowly, can we cache them?","coding","build","none","ready","none","boundary","'fixtures' must not match 'fix' and 'slowly' must not match 'slow' under boundary matching; the trailing question also skips the catch-all. v0.2 answered debug/35 (substring)." +"BD-003","the release notes praised the new fixtures; can you summarize why?","coding","build","none","ready","none","boundary","'praised' contains 'raised' (v0.2: debug/35). Post-fix both words are inert; question-shaped input skips insufficient_context, so the row must be clean." +"BD-004","I debugged the checkout flow all afternoon, can you review my patch?","coding","build","none","ready","none","boundary","'debugged' must not match 'debug' (past tense, not a request). v0.2 answered debug/35; post-fix the question shape keeps the catch-all silent." +"BD-005","After refactoring this module last sprint I finally profiled it: 400ms per request, down from 1200ms, and the public API stays backward compatible.","coding","optimization","none","ready","none","boundary","Over-suppression guard: 'refactoring' is an exact optimization signal word and must STILL be detected; all three satisfied-conditions hold so the well-specified input stays ready. Identical under v0.2." +"BD-006","did the team ship those three refactorings last quarter?","coding","build","none","ready","none","boundary","'refactorings' (plural) must not match 'refactor' or 'refactoring' under boundary matching; v0.2 answered optimization/55 (false positive). Question-shaped, so the catch-all stays silent." +"BD-007","optimize this function, it currently takes 3 seconds","coding","optimization","optimization constraint","ready","low","boundary","Threshold probe: score exactly 85 after the single low penalty; ready_at boundary must hold (ready despite a residual low-severity gap)." +"BD-008","optimize this function","coding","optimization","performance baseline; optimization constraint","usable_with_warnings","medium; low","boundary","Band probe: score 80, mid usable_with_warnings band (60-84). Minimal three-word optimization request." +"BD-009","can an HTML div be centered without flexbox?","coding","build","none","ready","none","boundary","Question-shaped non-task input: question_starters + '?' must skip insufficient_context; no intent signal may fire. Stays clean under both engines." +"BD-010","our onboarding checklist needs an owner by Friday","coding","build","task context","usable_with_warnings","high","boundary","Catch-all safety-net probe: statement form, 7 words, no specificity keyword, so insufficient_context must fire (score 75). Documents that the guard intentionally flags vague non-coding prompts routed to build." +"BD-011","help me","coding","build","none","ready","none","boundary","Word-floor probe: 2 words < min_words 3, so the catch-all must never fire on ultra-short input; no other rule has a trigger. Score 100." +"BD-012","the pytest fixture crashes with a KeyError, expected it to yield a dict, code in conftest.py","coding","debug","code context","ready","medium","boundary","No-regression guard: a GENUINE debug request containing 'fixture' must stay intent=debug and ready (85, one medium gap). The boundary fix must not over-suppress real signals." +"BD-013","should CSS classes start with a button prefix?","coding","build","none","ready","none","boundary","'prefix' contains 'fix' (v0.2: debug/35). Post-fix inert; question-shaped, so clean. 'start' avoids accidental creation/action verbs." +"BD-014","can the debugger attach to a running uvicorn worker?","coding","build","none","ready","none","boundary","'debugger' must not match 'debug' (v0.2: debug/35). Post-fix: no signal, question-shaped, clean." +"BD-015","optimise this query, it currently takes 900ms, keep the ordering stable","coding","optimization","none","ready","none","boundary","British spelling 'optimise' is an exact signal term and must survive boundary matching (no over-suppression). All three satisfied-conditions hold." +"BD-016","create a python REST api; GET /orders returns paged results","coding","build","none","ready","none","boundary","API-structure special cases: '/' and method tokens satisfy api structure; 'rest api' satisfies output format; 'python' satisfies language. Identical under v0.2." +"BD-017","fix the slow query","coding","debug","error description; expected vs actual behavior; code context","needs_clarification","high; high; medium","boundary","Priority-chain probe: debug beats optimization when both signal sets match. Documents that optimization wording under a debug trigger lands in debug rules." +"BD-018","explain how to add real-time chat to my existing app","coding","explanation","code reference; explanation depth","usable_with_warnings","high; low","boundary","Priority-chain probe: explanation beats feature when both trigger. Score 70 in warning mode." +"BD-019","make a landing page and add it to my app","coding","feature","existing stack; feature scope; completion criteria","needs_clarification","high; high; low","boundary","Priority-chain probe: feature beats the build fallback. Two high gaps + one low -> score 45." +"BD-020","I'm trying to scrape product prices","coding","build","programming language","usable_with_warnings","high","boundary","Structural-rule probe: _INTENT_PATTERN ('I'm trying to') bypasses the verb lists; intent_without_language fires without any creation verb (score 75)." +"BD-021","fix this?","coding","debug","error description; expected vs actual behavior; code context","needs_clarification","high; high; medium","boundary","Two-word input: the min_words floor guards only the catch-all; debug rules have no floor, so all three gaps fire (score 35)." +"DG-001","建造一个带用户登录和数据看板的Web应用","coding","n/a","none","degraded","none","degradation","Chinese (Han script). Spec verification row 'Non-English honesty' (probe P2): v0.2 answers score 100/ready with zero findings. Expected v0.3: degradation_note non-null, heuristic_coverage not 'full', and status never silent 'ready'. Uncovered script: English-only rules are skipped explicitly with a confidence penalty, never a silent pass." +"DG-002","帮我优化这段代码,它运行太慢了","coding","n/a","none","degraded","none","degradation","Chinese (Han script). Spec verification row 'Non-English honesty' (probe P2): v0.2 answers score 100/ready with zero findings. Expected v0.3: degradation_note non-null, heuristic_coverage not 'full', and status never silent 'ready'. Uncovered script: English-only rules are skipped explicitly with a confidence penalty, never a silent pass." +"DG-003","ユーザー認証付きのREST APIを構築してください","coding","n/a","none","degraded","none","degradation","Japanese. Spec verification row 'Non-English honesty' (probe P2): v0.2 answers score 100/ready with zero findings. Expected v0.3: degradation_note non-null, heuristic_coverage not 'full', and status never silent 'ready'. Uncovered script: English-only rules are skipped explicitly with a confidence penalty, never a silent pass." +"DG-004","この関数のパフォーマンスを改善する方法を教えてください","coding","n/a","none","degraded","none","degradation","Japanese. Spec verification row 'Non-English honesty' (probe P2): v0.2 answers score 100/ready with zero findings. Expected v0.3: degradation_note non-null, heuristic_coverage not 'full', and status never silent 'ready'. Uncovered script: English-only rules are skipped explicitly with a confidence penalty, never a silent pass." +"DG-005","أنشئ تطبيق ويب لتسجيل دخول المستخدمين مع لوحة تحكم","coding","n/a","none","degraded","none","degradation","Arabic. Spec verification row 'Non-English honesty' (probe P2): v0.2 answers score 100/ready with zero findings. Expected v0.3: degradation_note non-null, heuristic_coverage not 'full', and status never silent 'ready'. Uncovered script: English-only rules are skipped explicitly with a confidence penalty, never a silent pass." +"DG-006","لماذا يعمل هذا الكود ببطء وكيف يمكن تحسينه؟","coding","n/a","none","degraded","none","degradation","Arabic. Spec verification row 'Non-English honesty' (probe P2): v0.2 answers score 100/ready with zero findings. Expected v0.3: degradation_note non-null, heuristic_coverage not 'full', and status never silent 'ready'. Uncovered script: English-only rules are skipped explicitly with a confidence penalty, never a silent pass." +"DG-007","उपयोगकर्ता लॉगिन के साथ एक वेब ऐप बनाएं","coding","n/a","none","degraded","none","degradation","Hindi (Devanagari). Spec verification row 'Non-English honesty' (probe P2): v0.2 answers score 100/ready with zero findings. Expected v0.3: degradation_note non-null, heuristic_coverage not 'full', and status never silent 'ready'. Uncovered script: English-only rules are skipped explicitly with a confidence penalty, never a silent pass." +"DG-008","मेरा कोड बहुत धीमा चल रहा है, इसे कैसे सुधारें?","coding","n/a","none","degraded","none","degradation","Hindi (Devanagari). Spec verification row 'Non-English honesty' (probe P2): v0.2 answers score 100/ready with zero findings. Expected v0.3: degradation_note non-null, heuristic_coverage not 'full', and status never silent 'ready'. Uncovered script: English-only rules are skipped explicitly with a confidence penalty, never a silent pass." +"DG-009","Создай веб-приложение с авторизацией пользователей","coding","n/a","none","degraded","none","degradation","Russian (Cyrillic). Spec verification row 'Non-English honesty' (probe P2): v0.2 answers score 100/ready with zero findings. Expected v0.3: degradation_note non-null, heuristic_coverage not 'full', and status never silent 'ready'. Uncovered script: English-only rules are skipped explicitly with a confidence penalty, never a silent pass." +"DG-010","Почему этот код работает медленно и как его оптимизировать?","coding","n/a","none","degraded","none","degradation","Russian (Cyrillic). Spec verification row 'Non-English honesty' (probe P2): v0.2 answers score 100/ready with zero findings. Expected v0.3: degradation_note non-null, heuristic_coverage not 'full', and status never silent 'ready'. Uncovered script: English-only rules are skipped explicitly with a confidence penalty, never a silent pass." +"DG-011","Crée une application web avec connexion utilisateur et tableau de bord","coding","n/a","none","degraded","none","degradation","French (accented Latin). Spec verification row 'Non-English honesty' (probe P2): v0.2 answers score 100/ready with zero findings. Expected v0.3: degradation_note non-null, heuristic_coverage not 'full', and status never silent 'ready'. Accented Latin is technically the Latin script, but the heuristics are English-only: an ASCII-ratio/stop-word coverage probe must mark coverage as partial and set the note - Latin script alone must not silently pass." +"DG-012","¿Por qué mi código se ejecuta tan lento y cómo puedo optimizarlo?","coding","n/a","none","degraded","none","degradation","Spanish (accented Latin). Spec verification row 'Non-English honesty' (probe P2): v0.2 answers score 100/ready with zero findings. Expected v0.3: degradation_note non-null, heuristic_coverage not 'full', and status never silent 'ready'. Accented Latin is technically the Latin script, but the heuristics are English-only: an ASCII-ratio/stop-word coverage probe must mark coverage as partial and set the note - Latin script alone must not silently pass." +"DG-013","Fix this bug 修复这个错误 in the payment flow","coding","n/a","none","degraded","none","degradation","Mixed English + Han. Spec verification row 'Non-English honesty' (probe P2): v0.2 answers score 100/ready with zero findings. Expected v0.3: degradation_note non-null, heuristic_coverage not 'full', and status never silent 'ready'. English signal words ('fix', 'bug') coexist with Han script; the uncovered-script probe still forces the degraded path regardless of any partial English detection." +"DG-014","Preciso de um aplicativo web com login de usuário e relatórios","coding","n/a","none","degraded","none","degradation","Portuguese (accented Latin). Spec verification row 'Non-English honesty' (probe P2): v0.2 answers score 100/ready with zero findings. Expected v0.3: degradation_note non-null, heuristic_coverage not 'full', and status never silent 'ready'. Accented Latin is technically the Latin script, but the heuristics are English-only: an ASCII-ratio/stop-word coverage probe must mark coverage as partial and set the note - Latin script alone must not silently pass." +"TP-WRT-01","Write a blog post about remote work","writing","n/a","audience; purpose; structure/format; completeness","needs_clarification","high; high; medium; low","true_positive","Gap-bearing writing task (reference score 30); the listed writing-vocabulary gaps must fire and warning-mode status must match." +"TP-WRT-02","Write an email to the engineering team announcing the Q3 roadmap","writing","n/a","structure/format; completeness","usable_with_warnings","medium; low","true_positive","Gap-bearing writing task (reference score 80); the listed writing-vocabulary gaps must fire and warning-mode status must match." +"TP-WRT-03","Rewrite my resume summary","writing","n/a","audience; purpose; structure/format; source material; completeness","needs_clarification","high; high; medium; high; low","true_positive","Gap-bearing writing task (reference score 5); the listed writing-vocabulary gaps must fire and warning-mode status must match." +"TP-WRT-04","Improve the clarity of the essay I pasted above; admissions officers are the readers and it must stay under 650 words","writing","n/a","purpose","usable_with_warnings","high","true_positive","Gap-bearing writing task (reference score 75); the listed writing-vocabulary gaps must fire and warning-mode status must match." +"TP-WRT-05","Write a linkedin post about our Series A","writing","n/a","audience; purpose; structure/format; completeness","needs_clarification","high; high; medium; low","true_positive","Gap-bearing writing task (reference score 30); the listed writing-vocabulary gaps must fire and warning-mode status must match." +"TP-WRT-06","Rewrite this paragraph to be more concise","writing","n/a","audience; purpose; source material; completeness","needs_clarification","high; high; high; low","true_positive","Gap-bearing writing task (reference score 20); the listed writing-vocabulary gaps must fire and warning-mode status must match." +"BD-WRT-01","Turn my bullet outline into a full proposal for the steering committee; goal is to win headcount for Q1; include budget, risks, and milestones; the outline is at the bottom","writing","n/a","none","ready","none","boundary","Writing threshold probe: score exactly 100 (ready_at boundary with a residual medium gap)." +"TN-WRT-01","Draft a persuasive one-page memo to the executive team asking for budget for the migration project; must cover current costs, risks, and the timeline","writing","n/a","none","ready","none","true_negative","Well-specified writing task: every one of the six writing gaps is satisfied, so no finding may fire (score 100)." +"TN-WRT-02","Turn the meeting notes below into a structured recap for the client stakeholders, max 300 words, bullets with owners and deadlines","writing","n/a","none","ready","none","true_negative","Well-specified writing task: every one of the six writing gaps is satisfied, so no finding may fire (score 100)." +"TN-WRT-03","Help me draft an email to my manager requesting a two-week extension on the deadline; context: the vendor delivered the data late; keep it under 150 words and mention the new ship date","writing","n/a","none","ready","none","true_negative","Well-specified writing task: every one of the six writing gaps is satisfied, so no finding may fire (score 100)." +"TN-WRT-04","Draft a 300-word press release announcing the 2.0 launch for tech journalists; goal is maximum pickup; include the headline, a quote from the CEO, and pricing; background: we shipped the beta in May and doubled retention","writing","n/a","none","ready","none","true_negative","Well-specified writing task: every one of the six writing gaps is satisfied, so no finding may fire (score 100)." +"TN-WRT-05","Write a short email to my team explaining why the deadline moved, so that nobody reruns the old pipeline; mention the migration finished Tuesday and the new dashboard URL","writing","n/a","none","ready","none","true_negative","Well-specified writing task: every one of the six writing gaps is satisfied, so no finding may fire (score 100)." +"TN-WRT-06","Turn the survey responses I pasted below into a one-page summary for the board, highlighting the three biggest complaints; the goal is to justify the support-team expansion","writing","n/a","none","ready","none","true_negative","Well-specified writing task: every one of the six writing gaps is satisfied, so no finding may fire (score 100)." +"TN-WRT-07","Rewrite the cover letter I attached for the product-manager role at Stripe; recruiters skim in 30 seconds, so lead with outcomes; keep it to one page and include the two strongest metrics up top","writing","n/a","none","ready","none","true_negative","Well-specified writing task: every one of the six writing gaps is satisfied, so no finding may fire (score 100)." +"TP-DAT-01","Analyze my sales data","data-analysis","n/a","dataset/source; question/goal; output format; tooling; volume; reproducibility","needs_clarification","high; high; medium; medium; low; low","true_positive","Gap-bearing data-analysis task (reference score 10); the listed data-analysis gaps must fire and warning-mode status must match." +"TP-DAT-02","Explore the support ticket data and find why CSAT dropped","data-analysis","n/a","dataset/source; output format; tooling; volume; reproducibility","needs_clarification","high; medium; medium; low; low","true_positive","Gap-bearing data-analysis task (reference score 35); the listed data-analysis gaps must fire and warning-mode status must match." +"TP-DAT-03","What does this dataset say about churn?","data-analysis","n/a","dataset/source; output format; tooling; volume; reproducibility","needs_clarification","high; medium; medium; low; low","true_positive","Gap-bearing data-analysis task (reference score 35); the listed data-analysis gaps must fire and warning-mode status must match." +"TP-DAT-04","Compare the NPS scores between the two regions using the CSV files I emailed yesterday; deliver a two-slide summary for the ops review; roughly 5,000 responses total","data-analysis","n/a","tooling; reproducibility","usable_with_warnings","medium; low","true_positive","Gap-bearing data-analysis task (reference score 80); the listed data-analysis gaps must fire and warning-mode status must match." +"TP-DAT-05","Help me understand my website traffic","data-analysis","n/a","dataset/source; question/goal; output format; tooling; volume; reproducibility","needs_clarification","high; high; medium; medium; low; low","true_positive","Gap-bearing data-analysis task (reference score 10); the listed data-analysis gaps must fire and warning-mode status must match." +"TP-DAT-06","Pull insights from the NPS verbatims for the leadership offsite; group them into themes; there are about 2,300 comments in the sheet the HR team shared (nps_verbatims_2026.xlsx); a themes table plus pull-quotes works","data-analysis","n/a","tooling; reproducibility","usable_with_warnings","medium; low","true_positive","Gap-bearing data-analysis task (reference score 80); the listed data-analysis gaps must fire and warning-mode status must match." +"TP-DAT-07","Crunch the numbers in the spreadsheet (about 12,000 rows) and find insights for the board meeting","data-analysis","n/a","dataset/source; output format; tooling; reproducibility","needs_clarification","high; medium; medium; low","true_positive","Gap-bearing data-analysis task (reference score 40); the listed data-analysis gaps must fire and warning-mode status must match." +"BD-DAT-01","Analyze the churn in the postgres subscriptions table; I want to know which plan tier cancels most; a bar chart is fine; we run this monthly so make the SQL reusable","data-analysis","n/a","volume","ready","low","boundary","Data-analysis threshold probe: score 95 - ready band edge with a single residual gap; the satisfied conditions exercise one gap axis at a time." +"BD-DAT-02","Visualize the user growth trend from the GA4 export; the CFO wants a slide-ready chart; about 18 months of daily data; I'll rerun this every quarter so document the query steps","data-analysis","n/a","tooling","ready","medium","boundary","Data-analysis threshold probe: score 85 - ready band edge with a single residual gap; the satisfied conditions exercise one gap axis at a time." +"BD-DAT-03","Look at the attached Q3 transactions export (about 80,000 rows) and tell me whether refunds spiked after the pricing change; using pandas only; deliver a short summary with the weekly trend table","data-analysis","n/a","none","ready","none","boundary","Data-analysis threshold probe: score 100 - ready band edge with a single residual gap; the satisfied conditions exercise one gap axis at a time." +"TN-DAT-01","Run a cohort retention analysis on the event stream data in BigQuery (flattened_events table), about 400M rows, and produce the week-1/4/12 retention curves; no external libraries beyond dbt; the growth team will rerun it for every launch","data-analysis","n/a","none","ready","none","true_negative","Well-specified data-analysis task: all six data-analysis gaps satisfied (score 100)." +"TN-DAT-02","Analyze the A/B test results in the attached experiment export; determine whether variant B beats control on 7-day retention; deliver a one-page summary with the confidence interval; use statsmodels only; 120,000 users per arm; the data team will rerun this for every experiment","data-analysis","n/a","none","ready","none","true_negative","Well-specified data-analysis task: all six data-analysis gaps satisfied (score 100)." +"TN-DAT-03","Build a correlation matrix from the feature usage table in the warehouse (analytics.feature_events, about 90 days of data) and highlight pairs above 0.7; deliverable: a heatmap and the SQL to reproduce it; dbt model so analytics can rerun weekly","data-analysis","n/a","none","ready","none","true_negative","Well-specified data-analysis task: all six data-analysis gaps satisfied (score 100)." +"TN-DAT-04","Look at the page-load times in the RUM export I pasted below; I want to know whether the new CDN helped p95; a simple before/after chart is enough; run it in python; about two weeks of beacons; the query must be rerunnable against next month's export","data-analysis","n/a","none","ready","none","true_negative","Well-specified data-analysis task: all six data-analysis gaps satisfied (score 100)." +"TN-DAT-05","Profile the query performance of the reporting views (about 60M rows in the sales schema); goal is to cut the Monday refresh below 20 minutes; deliver the slowest three queries with rewritten SQL; BigQuery only; document each rewrite so it can be reviewed","data-analysis","n/a","none","ready","none","true_negative","Well-specified data-analysis task: all six data-analysis gaps satisfied (score 100)." +"PF-001","The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (0) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (1) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (2) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (3) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (4) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (5) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (6) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (7) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (8) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (9) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (10) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (11) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (12) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (13) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (14) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (15) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (16) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (17) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (18) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (19) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (20) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (21) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (22) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (23) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (24) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (25) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (26) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (27) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (28) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (29) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (30) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (31) Now the actual request: build a python CLI that deduplicates a CSV by the email column and writes a manifest of removed rows.The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (0) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (1) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (2) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (3) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (4) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (5) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (6) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (7) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (8) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (9) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (10) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (11) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (12) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (13) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (14) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (15) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (16) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (17) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (18) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (19) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (20) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (21) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (22) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (23) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (24) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (25) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (26) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (27) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (28) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (29) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (30) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (31) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (32) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (33) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (34) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (35) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (36) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (37) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (38) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (39) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (40) ","coding","build","none","ready","none","performance","9741 chars (< the 10,000-char cap): analyzed in full with no truncation. Filler is engine-verified signal-free; the embedded build request satisfies language and output format, so the expected result is ready/100. Assert p50/p99 within the latency budget (spec: p99 <= 25 ms at the cap)." +"PF-002","The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (0) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (1) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (2) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (3) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (4) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (5) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (6) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (7) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (8) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (9) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (10) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (11) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (12) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (13) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (14) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (15) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (16) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (17) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (18) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (19) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (20) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (21) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (22) Now the actual request: build a python CLI that deduplicates a CSV by the email column and writes a manifest of removed rows.The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (0) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (1) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (2) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (3) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (4) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (5) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (6) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (7) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (8) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (9) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (10) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (11) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (12) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (13) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (14) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (15) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (16) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (17) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (18) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (19) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (20) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (21) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (22) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (23) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (24) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (25) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (26) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (27) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (28) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (29) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (30) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (31) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (32) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (33) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (34) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (35) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (36) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (37) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (38) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (39) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (40) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (41) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (42) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (43) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (44) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (45) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (46) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (47) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (48) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (49) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (50) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (51) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (52) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (53) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (54) The platform team reviewed the roadmap for the next quarter and confirmed the milestones with the stakeholders at the offsite. (55) ","coding","build","none","ready","none","performance","10533 chars (> the 10,000-char cap): the request sits inside the first 10,000 chars, so the analyzed prefix yields the same ready/100 result as PF-001, and the truncation must be VISIBLE in the result (spec: truncation is never silent). Assert the truncation indicator, no error, and the latency budget holds at the cap." diff --git a/eval/measure_fp.py b/eval/measure_fp.py new file mode 100644 index 0000000..d9d435d --- /dev/null +++ b/eval/measure_fp.py @@ -0,0 +1,324 @@ +#!/usr/bin/env python3 +"""Measure InputGuard against the labeled clarity-evaluation set. + +Loads ``eval/cases.csv``, runs every row through ``inputguard``'s analyzer, +and prints per-case_type and overall false-positive / false-negative rates. + +This is a data-measurement tool, not a CI assertion: labels may legitimately +disagree with current analyzer behavior — that disagreement IS the +measurement. The script exits 0 after printing its summary regardless of the +rates. Only a broken run (missing file, malformed CSV, missing columns) +exits 2. + +Stdlib-only, Python 3.9+. Run from anywhere: + + python3 eval/measure_fp.py + python3 eval/measure_fp.py --limit 10 # quick smoke run + python3 eval/measure_fp.py --json # machine-readable summary +""" + +from __future__ import annotations + +import argparse +import csv +import json +import sys +import time +from pathlib import Path +from typing import Any, Dict, List, NoReturn, Optional, Sequence, Tuple + +# Make the repo checkout importable (and preferred over any site-packages +# install) so the tool measures the code in this tree without pip install. +REPO_ROOT = Path(__file__).resolve().parent.parent +if str(REPO_ROOT) not in sys.path: + sys.path.insert(0, str(REPO_ROOT)) + +REQUIRED_COLUMNS: Tuple[str, ...] = ( + "id", + "text", + "domain", + "expected_intent", + "expected_gaps", + "expected_status", + "expected_severities", + "case_type", + "labeler_note", +) + +CASE_TYPE_ORDER: Tuple[str, ...] = ( + "true_positive", + "true_negative", + "boundary", + "degradation", + "performance", +) + +SEVERITY_RANK: Dict[str, int] = {"low": 0, "medium": 1, "high": 2} +RANK_TO_SEVERITY: Dict[int, str] = {rank: name for name, rank in SEVERITY_RANK.items()} + +NONE_SENTINELS = ("", "none", "n/a") + + +def fail(message: str) -> NoReturn: + print(f"measure_fp: {message}", file=sys.stderr) + raise SystemExit(2) + + +def parse_multi_value(raw: str) -> List[str]: + """Split a ``;``-separated field; ``none``/``n/a``/empty mean no entries.""" + value = (raw or "").strip() + if value.lower() in NONE_SENTINELS: + return [] + return [part.strip() for part in value.split(";") if part.strip()] + + +def load_cases(path: Path) -> List[Dict[str, str]]: + if not path.is_file(): + fail(f"cases file not found: {path}") + with path.open(newline="", encoding="utf-8") as handle: + reader = csv.DictReader(handle) + fieldnames = reader.fieldnames or [] + missing = [column for column in REQUIRED_COLUMNS if column not in fieldnames] + if missing: + fail(f"{path} is missing required columns: {', '.join(missing)}") + rows = [dict(row) for row in reader] + if not rows: + fail(f"{path} contains no data rows") + return rows + + +def actual_gap_severities(findings: Sequence[Any]) -> Dict[str, int]: + """Map gap key -> max severity rank across that gap's findings.""" + best: Dict[str, int] = {} + for finding in findings: + key = finding.gap if finding.gap is not None else finding.code + rank = SEVERITY_RANK.get(finding.severity, -1) + if rank > best.get(key, -1): + best[key] = rank + return best + + +def classify_case(row: Dict[str, str], guard: Any) -> Dict[str, Any]: + """Run one labeled row through the analyzer and classify it. + + A row matches when: status equals expected, intent equals expected (when + the label asserts one), the actual gap set equals the expected set + (order-insensitive), and each shared gap's max severity equals the + expected severity. Any deviation is recorded as the row's reasons. + """ + expected_gaps = parse_multi_value(row["expected_gaps"]) + expected_severities = parse_multi_value(row["expected_severities"]) + expected_gap_set = set(expected_gaps) + intent_asserted = row["expected_intent"].strip().lower() not in NONE_SENTINELS + + record: Dict[str, Any] = { + "id": row["id"], + "case_type": row["case_type"], + "evaluated": True, + "match": False, + "status_match": False, + "intent_match": False, + "spurious_gaps": [], # analyzer flagged, label did not list -> false positive + "missed_gaps": [], # label listed, analyzer did not flag -> false negative + "severity_mismatches": [], + "actual_status": None, + "actual_intent": None, + "actual_gaps": [], + "wall_ms": None, + "notes": [], + "reasons": [], + } + + start = time.perf_counter() + try: + result = guard.analyze(row["text"], domain=row["domain"]) + except Exception as exc: # unevaluable row (e.g. domain not yet implemented) + record["evaluated"] = False + record["wall_ms"] = (time.perf_counter() - start) * 1000.0 + record["missed_gaps"] = list(expected_gaps) + message = f"analyzer raised {type(exc).__name__}: {exc}" + record["notes"].append(message) + record["reasons"].append(f"{message} — row counted as unmatched") + return record + + record["wall_ms"] = (time.perf_counter() - start) * 1000.0 + record["actual_status"] = result.status + record["actual_intent"] = result.detected_intent + record["actual_gaps"] = list(result.gaps) + + actual_gap_set = set(result.gaps) + record["spurious_gaps"] = sorted(actual_gap_set - expected_gap_set) + record["missed_gaps"] = [gap for gap in expected_gaps if gap not in actual_gap_set] + + actual_severities = actual_gap_severities(result.findings) + for gap, expected_severity in zip(expected_gaps, expected_severities): + if gap not in actual_gap_set: + continue + actual_rank = actual_severities.get(gap, -1) + expected_rank = SEVERITY_RANK.get(expected_severity.strip(), -1) + if actual_rank != expected_rank: + record["severity_mismatches"].append( + f"{gap}: expected {expected_severity}, got " + f"{RANK_TO_SEVERITY.get(actual_rank, 'unknown')}" + ) + + record["status_match"] = result.status == row["expected_status"] + record["intent_match"] = (not intent_asserted) or ( + result.detected_intent == row["expected_intent"] + ) + + if row["case_type"] == "degradation": + # Degradation rubric (labeling guide): the analyzer must report an + # explicit degradation note and never a silent `ready`. The note is + # additive on any underlying status; until that path ships the + # attribute is simply absent and the row stays unmatched. + note = getattr(result, "degradation_note", None) + record["notes"].append( + "degradation_note present" if note else "degradation_note absent" + ) + + reasons: List[str] = [] + if not record["status_match"]: + reasons.append(f"status {row['expected_status']} -> {result.status}") + if intent_asserted and not record["intent_match"]: + reasons.append(f"intent {row['expected_intent']} -> {result.detected_intent}") + if record["spurious_gaps"]: + reasons.append("spurious gaps: " + ", ".join(record["spurious_gaps"])) + if record["missed_gaps"]: + reasons.append("missed gaps: " + ", ".join(record["missed_gaps"])) + reasons.extend(f"severity {item}" for item in record["severity_mismatches"]) + + record["match"] = not reasons + record["reasons"] = reasons + return record + + +def summarize(results: Sequence[Dict[str, Any]]) -> Dict[str, Any]: + """Aggregate per-case results into per-case_type and overall rates.""" + groups: Dict[str, List[Dict[str, Any]]] = {name: [] for name in CASE_TYPE_ORDER} + for record in results: + groups.setdefault(record["case_type"], []).append(record) + + def group_summary(records: Sequence[Dict[str, Any]]) -> Dict[str, Any]: + total = len(records) + matches = sum(1 for r in records if r["match"]) + status_matches = sum(1 for r in records if r["status_match"]) + fp_cases = sum(1 for r in records if r["spurious_gaps"]) + fn_cases = sum(1 for r in records if r["missed_gaps"]) + unevaluated = sum(1 for r in records if not r["evaluated"]) + return { + "total": total, + "matches": matches, + "status_matches": status_matches, + "false_positive_cases": fp_cases, + "false_negative_cases": fn_cases, + "unevaluated": unevaluated, + "false_positive_rate": fp_cases / total if total else 0.0, + "false_negative_rate": fn_cases / total if total else 0.0, + "match_rate": matches / total if total else 0.0, + } + + summary = {name: group_summary(records) for name, records in groups.items() if records} + summary["overall"] = group_summary(list(results)) + return summary + + +def print_table(summary: Dict[str, Any], results: Sequence[Dict[str, Any]]) -> None: + header = ( + f"{'case_type':<15} {'total':>5} {'match':>6} {'status_ok':>9} " + f"{'fp_cases':>8} {'fn_cases':>8} {'fp_rate':>8} {'fn_rate':>8}" + ) + print("=== Summary (case-level rates) ===") + print(header) + print("-" * len(header)) + for name, stats in summary.items(): + print( + f"{name:<15} {stats['total']:>5} {stats['matches']:>6} " + f"{stats['status_matches']:>9} {stats['false_positive_cases']:>8} " + f"{stats['false_negative_cases']:>8} " + f"{stats['false_positive_rate']:>7.1%} {stats['false_negative_rate']:>7.1%}" + ) + print( + "\nfp: rows where the analyzer flagged a gap the label does not list " + "(false positives)\n" + "fn: rows where a labeled gap was not flagged (false negatives)" + ) + + degradation = summary.get("degradation") + if degradation: + noted = sum( + 1 + for r in results + if r["case_type"] == "degradation" + and any("degradation_note present" in n for n in r["notes"]) + ) + print( + f"\ndegradation honesty: degradation_note present on {noted}/" + f"{degradation['total']} degradation rows" + ) + + perf = [r for r in results if r["case_type"] == "performance"] + if perf: + timings = ", ".join(f"{r['id']} {r['wall_ms']:.0f} ms" for r in perf) + print(f"performance wall time (single pass): {timings}") + + mismatched = [r for r in results if not r["match"]] + if mismatched: + print(f"\n=== Mismatched rows ({len(mismatched)}) ===") + for record in mismatched: + detail_source = record.get("reasons") or record.get("notes") or [] + detail = "; ".join(detail_source) + print(f"{record['id']} [{record['case_type']}] {detail}") + + +def main(argv: Optional[Sequence[str]] = None) -> int: + parser = argparse.ArgumentParser( + description="Measure inputguard against the labeled clarity-evaluation set.", + ) + parser.add_argument( + "--cases", + type=Path, + default=Path(__file__).resolve().parent / "cases.csv", + help="path to the labeled cases CSV (default: eval/cases.csv)", + ) + parser.add_argument( + "--limit", + type=int, + default=None, + help="evaluate only the first N rows (smoke runs)", + ) + parser.add_argument( + "--json", + action="store_true", + help="print the summary and per-case results as JSON", + ) + args = parser.parse_args(argv) + + import inputguard # noqa: E402 (after sys.path setup above) + + rows = load_cases(args.cases) + if args.limit is not None: + if args.limit < 1: + fail("--limit must be >= 1") + rows = rows[: args.limit] + + guard = inputguard.InputGuard() + results = [classify_case(row, guard) for row in rows] + summary = summarize(results) + + if args.json: + print(json.dumps({"summary": summary, "results": results}, indent=2)) + else: + print( + f"InputGuard clarity evaluation — {len(rows)} rows from " + f"{args.cases} · analyzer {inputguard.__version__} · default policy" + ) + print_table(summary, results) + + # A measurement tool reports, it does not gate: labels disagreeing with + # current behavior is the measurement, so always exit 0 on a clean run. + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/tests/test_eval_loader.py b/tests/test_eval_loader.py new file mode 100644 index 0000000..30af459 --- /dev/null +++ b/tests/test_eval_loader.py @@ -0,0 +1,72 @@ +"""Smoke tests for the versioned clarity-evaluation set (``eval/``). + +These tests verify the dataset's mechanics only: the CSV parses, the +required columns exist, row ids are unique, and the measurement script runs +end to end on a small sample. They deliberately do NOT assert that labels +agree with analyzer output — label disagreement is the measurement +(``eval/measure_fp.py``), and freezing behavior against labels here would +defeat its purpose. +""" + +from __future__ import annotations + +import csv +import subprocess +import sys +from pathlib import Path +from typing import List + +REPO_ROOT = Path(__file__).resolve().parent.parent +CASES_CSV = REPO_ROOT / "eval" / "cases.csv" +MEASURE_FP = REPO_ROOT / "eval" / "measure_fp.py" + +REQUIRED_COLUMNS = ( + "id", + "text", + "domain", + "expected_intent", + "expected_gaps", + "expected_status", + "expected_severities", + "case_type", + "labeler_note", +) + + +def _load_rows() -> List[dict]: + with CASES_CSV.open(newline="", encoding="utf-8") as handle: + return list(csv.DictReader(handle)) + + +def test_cases_csv_parses_with_required_columns() -> None: + with CASES_CSV.open(newline="", encoding="utf-8") as handle: + reader = csv.DictReader(handle) + fieldnames = reader.fieldnames or [] + missing = [column for column in REQUIRED_COLUMNS if column not in fieldnames] + assert not missing, f"cases.csv missing columns: {missing}" + rows = _load_rows() + assert len(rows) > 0, "cases.csv contains no data rows" + assert all(row["text"].strip() for row in rows), "every case must carry prompt text" + + +def test_case_ids_are_unique() -> None: + rows = _load_rows() + ids = [row["id"] for row in rows] + duplicates = {row_id for row_id in ids if ids.count(row_id) > 1} + assert not duplicates, f"duplicate case ids: {sorted(duplicates)}" + + +def test_measure_fp_runs_end_to_end_on_sample() -> None: + # Smoke run on a 10-case sample: the tool must complete, exit 0, and + # print its summary table. Exit code and rates are not asserted against + # the labels — that is the measurement, not a gate. + result = subprocess.run( + [sys.executable, str(MEASURE_FP), "--limit", "10"], + cwd=str(REPO_ROOT), + capture_output=True, + text=True, + timeout=120, + ) + assert result.returncode == 0, f"measure_fp.py failed:\n{result.stderr}" + assert "=== Summary (case-level rates) ===" in result.stdout + assert "overall" in result.stdout From c2ccfd4a41645ac142193fe8e28ef93dbd47750a Mon Sep 17 00:00:00 2001 From: "obvious-autobuild[bot]" <262744130+obvious-autobuild[bot]@users.noreply.github.com> Date: Thu, 17 Sep 2026 20:01:17 +0000 Subject: [PATCH 06/18] =?UTF-8?q?fix:=20enforce=20registry=20contract=20?= =?UTF-8?q?=E2=80=94=20check(text)=20signature,=20unique=20intents,=20regi?= =?UTF-8?q?stration=20guards=20(#8)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * fix(registry-contract-spec-b1): revert Rule.check to the spec's check(self, text) signature The merged foundation (PR #3) dispatches rule.check(text, intent), but the pinned spec (art_bTvdPdJS section 1) defines check(self, text). The intent parameter is redundant by construction: analyze() dispatches REGISTRY.rules_for_intent(detected_intent), which filters on rule.domain == intent, so every rule invoked already knows the intent as its own domain member — and all 19 built-in adapters ignore it. The mismatch is not cosmetic: registration only checks callable(rule.check), so a rule written per the spec registers cleanly and then crashes analyze() mid-run with "TypeError: check() takes 2 positional arguments but 3 were given" (review probe P1). Spec and code must agree before wave-2/3 rule authors copy either version. - inputguard/registry.py: Rule protocol back to check(self, text); module and member docstrings updated to match. - inputguard/analyzer.py: dispatch rule.check(normalized). - inputguard/rules/: all 19 built-in adapters drop the unused intent parameter; InsufficientContextRule docstring no longer cites (text, intent). - tests/test_registry.py: helper and decorator-form test rules updated. Part of release PR #2. Co-authored-by: Kalisetti Nihanth Naidu * fix(registry-intent-scoping-b2): globally-unique intents, registration guards, loud rule attribution Closes the cross-domain leak and the registration holes the adversarial review proved against the merged foundation (art_hC18m78C). All guards fire at registration time, before any mutation. - B2: register_domain raises ValueError when an intent name is already declared by another registered domain, naming both domains. Rules dispatch on intent name alone (rules_for_intent filters on rule.domain), so a shared intent name ran one domain's rules inside the other's analysis (review probe P6b: a legal rule fired inside a coding analysis). Duplicate intent names within one signals argument raise too — the ownership map must be unambiguous. - C4: the self-contradictory error ("declares domain 'coding', which no registered domain declares as an intent. Registered domains: 'coding'") now states the actual constraint and enumerates the valid intent names; the Rule protocol documents that domain holds the intent name the rule is registered under, never a domain name. - C3: duplicate rule ids within one register_domain call raise — the old code silently dropped the second rule (review probe P2). - N3: _validate checks the check() signature via inspect.signature binding; a rule whose check cannot accept the single positional text argument is a registration error, not a mid-analyze TypeError (review probe P1's crash shape). - C2-lite: analyze() wraps rule dispatch and re-raises RuntimeError with the rule id and its registration origin ("at :", captured at _add time and exposed via RuleRegistry.rule_origin), chaining the original exception. Blast-radius contract documented on the Rule protocol and RuleRegistry: a rule exception aborts analyze() by design, registration is permanent, REGISTRY is process-global. - C6: the extension API (register_rule, register_domain, REGISTRY, Rule) is exported at package top level, additive on the untouched v0.2 four-export surface. Part of release PR #2. Co-authored-by: Kalisetti Nihanth Naidu * test(registry-contract-guards): pin every remediated registry behavior in the suite Permanent tests for the review's findings so the gates enforce the chosen contract (the review found none of these covered): - spec-compliant check(text) rule registers and fires end to end (P1) - wrong-arity and zero-argument check rejected at registration, nothing enters the registry (N3) - cross-domain intent collision raises naming both domains (B2/P6b), leakage probe cannot recur - duplicate intent within one signals argument raises (N1 registry half) - same-call duplicate rule ids raise and leave no partial registration (C3/P2) - rule exception aborts analyze() with rule id + registration origin and the original cause chained (C2-lite/P4) - extension API importable at top level; v0.2 four-export surface intact (C6) Suite: 123 passed (114 baseline + 9 new). Part of release PR #2. Co-authored-by: Kalisetti Nihanth Naidu * test: adapt PR #4 follow-ups test rule to the spec check(text) signature Rebase adaptation: PR #4's unknown-gap test registered a rule with the retired check(text, intent) signature; the N3 registration guard now correctly rejects it. The test's subject (unknown-gap fallback follow-up) is orthogonal to the signature — the rule moves to the spec's check(self, text). Suite: 202 passed. Part of release PR #2. Co-authored-by: Kalisetti Nihanth Naidu --------- Co-authored-by: Obvious --- inputguard/__init__.py | 15 ++- inputguard/analyzer.py | 15 ++- inputguard/registry.py | 180 ++++++++++++++++++++++++++++--- inputguard/rules/coding.py | 26 ++--- inputguard/rules/debug.py | 6 +- inputguard/rules/explanation.py | 4 +- inputguard/rules/feature.py | 6 +- inputguard/rules/optimization.py | 6 +- tests/test_followups.py | 2 +- tests/test_registry.py | 156 ++++++++++++++++++++++++++- 10 files changed, 371 insertions(+), 45 deletions(-) diff --git a/inputguard/__init__.py b/inputguard/__init__.py index f90a4ef..f62271f 100644 --- a/inputguard/__init__.py +++ b/inputguard/__init__.py @@ -1,6 +1,19 @@ from inputguard.analyzer import InputGuard +from inputguard.registry import REGISTRY, Rule, register_domain, register_rule from inputguard.types import AnalysisResult, RuleFinding __version__ = "0.2.0" -__all__ = ["InputGuard", "AnalysisResult", "RuleFinding", "__version__"] +__all__ = [ + # v0.2 public API — unchanged compat contract. + "InputGuard", + "AnalysisResult", + "RuleFinding", + "__version__", + # v0.3 extension API (additive): register custom rules and domains + # through the same path the built-ins take. + "REGISTRY", + "Rule", + "register_rule", + "register_domain", +] diff --git a/inputguard/analyzer.py b/inputguard/analyzer.py index 94074ca..d518443 100644 --- a/inputguard/analyzer.py +++ b/inputguard/analyzer.py @@ -75,7 +75,20 @@ def analyze(self, user_input: str, domain: str = "coding") -> AnalysisResult: findings: List[RuleFinding] = [] seen_codes: Set[str] = set() for rule in REGISTRY.rules_for_intent(detected_intent): - finding = rule.check(normalized, detected_intent) + try: + finding = rule.check(normalized) + except Exception as exc: + # Loud failure with attribution (review C2-lite): a rule + # exception is never swallowed or silently degraded around — + # it aborts analyze(), naming the rule and where it was + # registered, with the original traceback chained. + raise RuntimeError( + f"inputguard rule {rule.id!r} " + f"(registered {REGISTRY.rule_origin(rule.id)}) raised " + f"{type(exc).__name__}: {exc}. A rule exception aborts " + f"analyze() by design — rules are the author's " + f"responsibility after registration; fix or remove the rule." + ) from exc if finding is None or finding.code in seen_codes: continue seen_codes.add(finding.code) diff --git a/inputguard/registry.py b/inputguard/registry.py index 3762bcc..6261bf2 100644 --- a/inputguard/registry.py +++ b/inputguard/registry.py @@ -4,11 +4,13 @@ engine: rules and domains are registered data, not code paths. - A rule is any object with the four members (``id``, ``domain``, - ``severity``, ``gap``) and a ``check(text, intent)`` method — the + ``severity``, ``gap``) and a ``check(text)`` method — the :class:`Rule` protocol. - A domain is a named analysis scope (``"coding"`` ships built in) that declares its intent signals in priority order via - :meth:`RuleRegistry.register_domain`. + :meth:`RuleRegistry.register_domain`. Intent names are globally unique + across domains (enforced at registration) because rules dispatch on + intent name alone. Registration happens at import/startup time; ``analyze()`` only reads the registry, preserving the thread-safe, dependency-free pipeline v0.2 @@ -17,6 +19,7 @@ from __future__ import annotations +import inspect from typing import ( Dict, Iterable, @@ -49,6 +52,29 @@ _REQUIRED_MEMBERS = ("id", "domain", "severity", "gap", "check") +# Placeholder passed to check's signature at registration to prove the +# analyzer's one-positional-argument dispatch can bind. +_PROBE_TEXT = "probe text" + + +def _registration_origin() -> str: + """Call site of the registration, for loud rule attribution. + + The first stack frame outside this module — where ``register_rule`` / + ``register_domain`` was actually invoked. Registration is startup-time, + so the walk costs nothing during ``analyze()``. + """ + frame = inspect.currentframe() + try: + caller = frame + while caller is not None and caller.f_code.co_filename == __file__: + caller = caller.f_back + if caller is None: + return "unknown origin" + return f"at {caller.f_code.co_filename}:{caller.f_lineno}" + finally: + del frame # break reference cycles through frame objects + @runtime_checkable class Rule(Protocol): @@ -57,10 +83,31 @@ class Rule(Protocol): ``id`` must be unique across the registry, ``severity`` must be one of ``'low' | 'medium' | 'high'`` (validated at registration), and ``gap`` groups the rule's findings for scoring dedup (``None`` dedupes by code - instead). ``check`` receives lowercased, whitespace-collapsed text plus - the intent detected for this ``analyze()`` call, and returns at most one + instead). + + ``domain`` holds the **intent name** the rule is registered under — + never a domain name. ``analyze()`` dispatches + ``rules_for_intent(detected_intent)``, so a rule fires only when its + intent is the one detected; for the built-in coding domain the valid + values are ``'build'``, ``'debug'``, ``'optimization'``, + ``'explanation'``, and ``'feature'``. Intent names are globally unique + across domains (``register_domain`` raises on a collision), so an + intent name unambiguously identifies the rules that run for it. + + ``check`` receives lowercased, whitespace-collapsed text, and + returns at most one :class:`~inputguard.types.RuleFinding` — ``None`` when the rule does not fire. + + Contracts a rule author accepts at registration: + + - A ``check`` exception is not contained: it aborts the ``analyze()`` + call in flight, re-raised with the rule id and its registration + origin. Shape is validated at registration; behaviour is not — a + rule is the author's responsibility after registration. + - Registration is permanent for the process lifetime (there is no + unregister), and the module-level ``REGISTRY`` is a process-global + singleton shared by everything that imports inputguard. """ id: str @@ -68,7 +115,7 @@ class Rule(Protocol): severity: str gap: Optional[str] - def check(self, text: str, intent: str) -> Optional[RuleFinding]: + def check(self, text: str) -> Optional[RuleFinding]: ... # pragma: no cover — protocol body @@ -92,12 +139,19 @@ def _normalize_signals( raise ValueError(f"Domain {name!r}: signals must declare at least one intent.") normalized: List[Tuple[str, Tuple[str, ...]]] = [] + seen_intents: Set[str] = set() fallbacks = 0 for intent, terms in items: if not isinstance(intent, str) or not intent: raise ValueError( f"Domain {name!r}: intent names must be non-empty strings, got {intent!r}." ) + if intent in seen_intents: + raise ValueError( + f"Domain {name!r}: duplicate intent name {intent!r} in signals — " + f"declare each intent exactly once." + ) + seen_intents.add(intent) if isinstance(terms, str) or not isinstance(terms, Iterable): raise TypeError( f"Domain {name!r}: terms for intent {intent!r} must be an iterable " @@ -129,11 +183,23 @@ class RuleRegistry: lookups and iteration over registered rules, so parallel ``analyze()`` calls stay consistent (the v0.2 thread-safety contract). Registering while another thread is mid-``analyze()`` is not supported. + + Blast radius: registration is permanent for the process lifetime — + there is no unregister — and the module-level ``REGISTRY`` is a + process-global singleton shared by everything that imports inputguard. + The registry validates a rule's shape at registration (members, + severity, ``check`` arity); it cannot guard a rule's behaviour. A + ``check`` that raises aborts the ``analyze()`` call in flight with + loud attribution (rule id + registration origin) — a broken rule is + its author's registration responsibility. """ def __init__(self) -> None: self._rules: Dict[str, Rule] = {} self._domains: Dict[str, SignalSpec] = {} + # rule id -> human-readable registration call site, for loud + # attribution when a rule raises inside analyze(). + self._origins: Dict[str, str] = {} # -- registration ---------------------------------------------------- @@ -147,8 +213,11 @@ def register_rule(self, rule: Union[type, Rule]) -> Union[type, Rule]: Raises ``ValueError`` on a duplicate rule id, an unknown severity, or — once any domain is registered — a domain no registered domain declares as an intent. Raises ``TypeError`` when the object is not - shaped like a :class:`Rule` or the class cannot be constructed with - no arguments. + shaped like a :class:`Rule`, the class cannot be constructed with + no arguments, or ``check`` cannot be called with the single + positional argument the analyzer passes (``rule.check(text)``) — + a wrong-signature rule is a registration error, never a mid-``analyze()`` + crash. """ if isinstance(rule, type): instance = _instantiate(rule) @@ -174,9 +243,15 @@ def register_domain( example via the ``@register_rule`` decorator) are reused, and the rest are registered through the same path. + Intent names are **globally unique across domains**: rules dispatch on + intent name alone, so an intent name declared by two domains would run + one domain's rules inside the other's analysis. + Raises ``ValueError`` on a duplicate domain name, malformed signals - (no fallback intent), or a rule whose domain the domain does not - declare. + (no fallback intent, duplicate intent names within ``signals``), an + intent name already declared by another registered domain, duplicate + rule ids within the ``rules`` argument, or a rule whose domain the + domain does not declare. Nothing mutates unless every check passes. """ if not isinstance(name, str) or not name: raise ValueError(f"Domain name must be a non-empty string, got {name!r}.") @@ -187,6 +262,21 @@ def register_domain( normalized_signals = _normalize_signals(name, signals) declared = {intent for intent, _ in normalized_signals} + # Globally-unique intent names: a collision would silently leak rules + # across domains in both directions (rules_for_intent filters on the + # intent name alone), so it is a registration-time error. + for intent in sorted(declared): + owner = self._intent_owner(intent) + if owner is not None: + raise ValueError( + f"Domain {name!r} cannot declare intent {intent!r}: it is " + f"already declared by domain {owner!r}. Intent names must be " + f"globally unique across domains — rules dispatch on intent " + f"name alone, so a shared intent name would run one domain's " + f"rules inside the other domain's analysis. Rename the intent " + f"in one of the two domains." + ) + coerced = [self._coerce(entry) for entry in rules] for rule in coerced: self._validate(rule) @@ -194,8 +284,17 @@ def register_domain( if rule.domain not in declared: raise ValueError( f"Rule {rule.id!r} declares domain {rule.domain!r}, which is not " - f"an intent of domain {name!r}. Declared intents: {sorted(declared)}." + f"an intent of domain {name!r}. Rule.domain must hold one of the " + f"domain's intent names — never the domain name itself. " + f"Declared intents: {sorted(declared)}." ) + coerced_ids = [rule.id for rule in coerced] + if len(set(coerced_ids)) != len(coerced_ids): + duplicated = sorted({rid for rid in coerced_ids if coerced_ids.count(rid) > 1}) + raise ValueError( + f"Duplicate rule id within one register_domain call: " + f"{', '.join(repr(d) for d in duplicated)}. Rule ids must be unique." + ) for rule in coerced: existing = self._rules.get(rule.id) if existing is not None and existing is not rule: @@ -231,7 +330,12 @@ def domain_names(self) -> Tuple[str, ...]: return tuple(self._domains) def rules_for_intent(self, intent: str) -> List[Rule]: - """Rules whose domain is this intent scope, in registration order.""" + """Rules whose domain is this intent scope, in registration order. + + Intent names are globally unique across domains (enforced at + registration), so an intent name dispatches exactly one domain's + rules — never a mix from several domains. + """ return [rule for rule in self._rules.values() if rule.domain == intent] def rule_ids(self) -> Tuple[str, ...]: @@ -246,6 +350,15 @@ def get_rule(self, rule_id: str) -> Optional[Rule]: """Return the rule registered under ``rule_id``, or ``None``.""" return self._rules.get(rule_id) + def rule_origin(self, rule_id: str) -> str: + """Human-readable registration call site for ``rule_id``. + + ``"at :"`` — the first frame outside this module when + the rule was registered — or ``"unknown origin"`` if the id is not + registered. Used for loud rule attribution in analyzer errors. + """ + return self._origins.get(rule_id, "unknown origin") + # -- internals --------------------------------------------------------- def _coerce(self, entry: Union[type, Rule]) -> Rule: @@ -266,7 +379,7 @@ def _validate(self, rule: Rule) -> None: raise TypeError( f"Rule {rule!r} is missing required member(s): " f"{', '.join(repr(m) for m in missing)}. A rule needs id, domain, " - "severity, gap, and a check(text, intent) method." + "severity, gap, and a check(text) method." ) if not isinstance(rule.id, str) or not rule.id: raise ValueError(f"Rule id must be a non-empty string, got {rule.id!r}.") @@ -286,8 +399,28 @@ def _validate(self, rule: Rule) -> None: if not callable(rule.check): raise TypeError( f"Rule {rule.id!r}: check must be callable — " - "check(text, intent) -> Optional[RuleFinding]." + "check(text) -> Optional[RuleFinding]." ) + # Arity validation (review N3): the analyzer dispatches exactly one + # positional argument, rule.check(normalized_text). A signature that + # cannot accept it must fail here, at registration — not later, + # mid-analyze, on arbitrary user input. + try: + signature = inspect.signature(rule.check) + signature.bind(_PROBE_TEXT) + except TypeError as exc: + raise TypeError( + f"Rule {rule.id!r}: check{signature} cannot be called as " + f"rule.check(text) — the analyzer passes exactly one positional " + f"argument, the normalized text. Define check(self, text) per the " + f"Rule protocol. ({exc})" + ) from exc + except ValueError: + # Signature introspection unavailable (e.g. some C callables): + # skip the arity check rather than reject an inspectable-in-practice + # rule. A genuinely wrong signature still fails loudly at + # analyze() with rule attribution. + pass def _add(self, rule: Rule) -> None: if rule.id in self._rules: @@ -296,19 +429,34 @@ def _add(self, rule: Rule) -> None: # rules register during package import, before any domain exists), then # enforced for everything registered afterwards. if self._domains and rule.domain not in self._declared_intents(): + valid_intents = ", ".join(repr(i) for i in sorted(self._declared_intents())) raise ValueError( - f"Rule {rule.id!r} declares domain {rule.domain!r}, which no " - f"registered domain declares as an intent. Registered domains: " - f"{', '.join(repr(d) for d in self._domains)}." + f"Rule {rule.id!r} declares domain {rule.domain!r}, but Rule.domain " + f"must hold the intent name the rule is registered under — never a " + f"domain name. Valid intents: {valid_intents or 'none'}. " + f"Registered domains: {', '.join(repr(d) for d in self._domains)}." ) self._rules[rule.id] = rule + self._origins[rule.id] = _registration_origin() def _declared_intents(self) -> Set[str]: declared: Set[str] = set() for signal_spec in self._domains.values(): declared.update(intent for intent, _ in signal_spec) + return declared + def _intent_owner(self, intent: str) -> Optional[str]: + """The domain that declared ``intent``, or ``None``. + + Globally-unique intent names (enforced at registration) make this + unambiguous: an intent belongs to exactly one domain. + """ + for domain, signal_spec in self._domains.items(): + if any(existing == intent for existing, _ in signal_spec): + return domain + return None + REGISTRY = RuleRegistry() diff --git a/inputguard/rules/coding.py b/inputguard/rules/coding.py index 60ca560..85f8fc1 100644 --- a/inputguard/rules/coding.py +++ b/inputguard/rules/coding.py @@ -316,7 +316,7 @@ class MissingLanguageRule: severity = "high" gap = "programming language" - def check(self, text: str, intent: str) -> Optional[RuleFinding]: + def check(self, text: str) -> Optional[RuleFinding]: return check_missing_language(text) @@ -329,7 +329,7 @@ class MissingApiStructureRule: severity = "high" gap = "api structure" - def check(self, text: str, intent: str) -> Optional[RuleFinding]: + def check(self, text: str) -> Optional[RuleFinding]: return check_missing_api_structure(text) @@ -342,7 +342,7 @@ class MissingDataModelRule: severity = "high" gap = "data model" - def check(self, text: str, intent: str) -> Optional[RuleFinding]: + def check(self, text: str) -> Optional[RuleFinding]: return check_missing_data_model(text) @@ -355,7 +355,7 @@ class MissingIntegrationSpecificsRule: severity = "medium" gap = "integration specifics" - def check(self, text: str, intent: str) -> Optional[RuleFinding]: + def check(self, text: str) -> Optional[RuleFinding]: return check_missing_integration_specifics(text) @@ -368,7 +368,7 @@ class MissingAuthTypeRule: severity = "high" gap = "authentication type" - def check(self, text: str, intent: str) -> Optional[RuleFinding]: + def check(self, text: str) -> Optional[RuleFinding]: return check_missing_auth_type(text) @@ -381,7 +381,7 @@ class MissingOutputFormatRule: severity = "medium" gap = "output format" - def check(self, text: str, intent: str) -> Optional[RuleFinding]: + def check(self, text: str) -> Optional[RuleFinding]: return check_missing_output_format(text) @@ -394,7 +394,7 @@ class IntentWithoutDetailRule: severity = "high" gap = "programming language" - def check(self, text: str, intent: str) -> Optional[RuleFinding]: + def check(self, text: str) -> Optional[RuleFinding]: return _check_intent_without_detail(text) @@ -403,11 +403,11 @@ class InsufficientContextRule: """Catch-all safety net for build-intent input (registry adapter). v0.2's run_coding_rules passed this rule the findings collected so far and - it fired only when that list was empty. A Rule sees only (text, intent), - so the adapter re-runs the other built-in build rules — they are pure - functions, so the verdict is identical. Findings from user-registered - rules are not visible here: the catch-all suppresses on the built-in - build rules only. + it fired only when that list was empty. A Rule sees only the normalized + text, so the adapter re-runs the other built-in build rules — they are + pure functions, so the verdict is identical. Findings from + user-registered rules are not visible here: the catch-all suppresses on + the built-in build rules only. """ id = "insufficient_context" @@ -415,7 +415,7 @@ class InsufficientContextRule: severity = "high" gap = "task context" - def check(self, text: str, intent: str) -> Optional[RuleFinding]: + def check(self, text: str) -> Optional[RuleFinding]: normalized = _normalize(text) seen_codes: Set[str] = set() prior: List[RuleFinding] = [] diff --git a/inputguard/rules/debug.py b/inputguard/rules/debug.py index 3067ae7..f474dbc 100644 --- a/inputguard/rules/debug.py +++ b/inputguard/rules/debug.py @@ -120,7 +120,7 @@ class MissingErrorMessageRule: severity = "high" gap = "error description" - def check(self, text: str, intent: str) -> Optional[RuleFinding]: + def check(self, text: str) -> Optional[RuleFinding]: return check_missing_error_message(text) @@ -133,7 +133,7 @@ class MissingExpectedVsActualRule: severity = "high" gap = "expected vs actual behavior" - def check(self, text: str, intent: str) -> Optional[RuleFinding]: + def check(self, text: str) -> Optional[RuleFinding]: return check_missing_expected_vs_actual(text) @@ -146,7 +146,7 @@ class MissingDebugCodeContextRule: severity = "medium" gap = "code context" - def check(self, text: str, intent: str) -> Optional[RuleFinding]: + def check(self, text: str) -> Optional[RuleFinding]: return check_missing_debug_code_context(text) diff --git a/inputguard/rules/explanation.py b/inputguard/rules/explanation.py index 8a5a5cf..bc933e6 100644 --- a/inputguard/rules/explanation.py +++ b/inputguard/rules/explanation.py @@ -101,7 +101,7 @@ class MissingCodeReferenceRule: severity = "high" gap = "code reference" - def check(self, text: str, intent: str) -> Optional[RuleFinding]: + def check(self, text: str) -> Optional[RuleFinding]: return check_missing_code_reference(text) @@ -114,7 +114,7 @@ class MissingExplanationDepthRule: severity = "low" gap = "explanation depth" - def check(self, text: str, intent: str) -> Optional[RuleFinding]: + def check(self, text: str) -> Optional[RuleFinding]: return check_missing_explanation_depth(text) diff --git a/inputguard/rules/feature.py b/inputguard/rules/feature.py index 7afa551..642ab18 100644 --- a/inputguard/rules/feature.py +++ b/inputguard/rules/feature.py @@ -121,7 +121,7 @@ class MissingExistingStackRule: severity = "high" gap = "existing stack" - def check(self, text: str, intent: str) -> Optional[RuleFinding]: + def check(self, text: str) -> Optional[RuleFinding]: return check_missing_existing_stack(text) @@ -134,7 +134,7 @@ class MissingFeatureScopeRule: severity = "high" gap = "feature scope" - def check(self, text: str, intent: str) -> Optional[RuleFinding]: + def check(self, text: str) -> Optional[RuleFinding]: return check_missing_feature_scope(text) @@ -147,7 +147,7 @@ class MissingCompletionCriteriaRule: severity = "low" gap = "completion criteria" - def check(self, text: str, intent: str) -> Optional[RuleFinding]: + def check(self, text: str) -> Optional[RuleFinding]: return check_missing_completion_criteria(text) diff --git a/inputguard/rules/optimization.py b/inputguard/rules/optimization.py index 031346c..b1d82ed 100644 --- a/inputguard/rules/optimization.py +++ b/inputguard/rules/optimization.py @@ -120,7 +120,7 @@ class MissingOptimizationTargetRule: severity = "high" gap = "optimization target" - def check(self, text: str, intent: str) -> Optional[RuleFinding]: + def check(self, text: str) -> Optional[RuleFinding]: return check_missing_optimization_target(text) @@ -133,7 +133,7 @@ class MissingPerformanceBaselineRule: severity = "medium" gap = "performance baseline" - def check(self, text: str, intent: str) -> Optional[RuleFinding]: + def check(self, text: str) -> Optional[RuleFinding]: return check_missing_performance_baseline(text) @@ -146,7 +146,7 @@ class MissingOptimizationConstraintRule: severity = "low" gap = "optimization constraint" - def check(self, text: str, intent: str) -> Optional[RuleFinding]: + def check(self, text: str) -> Optional[RuleFinding]: return check_missing_optimization_constraint(text) diff --git a/tests/test_followups.py b/tests/test_followups.py index 6e21a5d..2f93711 100644 --- a/tests/test_followups.py +++ b/tests/test_followups.py @@ -179,7 +179,7 @@ class DeadlineRule: severity = "medium" gap = "deadline" - def check(self, text: str, intent: str): + def check(self, text: str): return finding rules_before = {rule.id: rule for rule in REGISTRY.rules()} diff --git a/tests/test_registry.py b/tests/test_registry.py index 7e0da55..e60b294 100644 --- a/tests/test_registry.py +++ b/tests/test_registry.py @@ -66,7 +66,7 @@ class TestRule: TestRule.domain = domain TestRule.severity = severity TestRule.gap = gap - TestRule.check = lambda self, text, intent: finding + TestRule.check = lambda self, text: finding return TestRule() @@ -75,11 +75,14 @@ def registry_isolation(): """Snapshot the registry around tests that mutate it.""" rules_before = dict(REGISTRY._rules) domains_before = dict(REGISTRY._domains) + origins_before = dict(REGISTRY._origins) yield REGISTRY REGISTRY._rules.clear() REGISTRY._rules.update(rules_before) REGISTRY._domains.clear() REGISTRY._domains.update(domains_before) + REGISTRY._origins.clear() + REGISTRY._origins.update(origins_before) # --- the 19 built-ins dogfood the registry path --------------------------- @@ -147,7 +150,7 @@ class AlwaysAskForDeadlineRule: severity = "medium" gap = "deadline" - def check(self, text, intent): + def check(self, text): return RuleFinding( code=self.id, message="Test rule fired.", @@ -304,3 +307,152 @@ def test_parallel_analyze_stays_consistent(): assert len(parallel) == 65 for offset in range(0, len(parallel), len(texts)): assert parallel[offset : offset + len(texts)] == sequential + + +# --- registry contract remediations (adversarial review art_hC18m78C) ------ + + +def test_spec_compliant_check_text_rule_registers_and_fires(registry_isolation): + """Review P1: a rule written per the pinned spec used to crash analyze().""" + + class SpecRule: + id = "test_spec_signature_rule" + domain = "debug" + severity = "medium" + gap = "deadline" + + def check(self, text): + if "fix the bug" in text: + return RuleFinding( + code=self.id, + message="Spec-signature rule fired.", + severity=self.severity, + gap=self.gap, + ) + return None + + register_rule(SpecRule()) + result = InputGuard().analyze("fix the bug in my app") + assert "test_spec_signature_rule" in [f.code for f in result.findings] + + +def test_wrong_arity_check_rejected_at_registration(registry_isolation): + """Review N3: a wrong-signature rule is a registration error, not a + mid-analyze TypeError. The retired foundation signature is the case.""" + + class OldSignatureRule: + id = "test_old_signature" + domain = "debug" + severity = "low" + gap = None + + def check(self, text, intent): + return None + + with pytest.raises(TypeError, match="test_old_signature"): + register_rule(OldSignatureRule()) + assert "test_old_signature" not in REGISTRY.rule_ids() + + +def test_zero_argument_check_rejected_at_registration(registry_isolation): + + class NoArgRule: + id = "test_noarg_check" + domain = "debug" + severity = "low" + gap = None + + def check(self): + return None + + with pytest.raises(TypeError, match="test_noarg_check"): + register_rule(NoArgRule()) + assert "test_noarg_check" not in REGISTRY.rule_ids() + + +def test_intent_collision_across_domains_raises(registry_isolation): + """Review B2/P6b: 'legal' reusing coding's 'debug' intent used to register + silently and leak rules across domains. The error names both domains.""" + legal_signals = {"debug": ("lawsuit", "court filing"), "plead": ()} + with pytest.raises(ValueError, match="globally unique") as exc_info: + register_domain("legal", legal_signals) + assert "legal" in str(exc_info.value) + assert "coding" in str(exc_info.value) + + +def test_cross_domain_rule_leakage_is_now_impossible(registry_isolation): + """Review P6b: a legal rule firing inside a coding analysis must not recur.""" + legal_signals = {"debug": ("lawsuit", "court filing"), "plead": ()} + with pytest.raises(ValueError, match="globally unique"): + register_domain("legal", legal_signals) + + result = InputGuard().analyze("fix the bug in my app") + assert "legal_fires" not in [f.code for f in result.findings] + assert "legal" not in REGISTRY.domain_names() + + +def test_duplicate_intent_within_one_domain_raises(registry_isolation): + """Review N1 (registry half): one call cannot declare an intent twice.""" + duplicated = [("build", ("make",)), ("build", ()), ("review", ("check",))] + with pytest.raises(ValueError, match="duplicate intent"): + register_domain("testdupint", duplicated) + + +def test_same_call_duplicate_rule_ids_raise(registry_isolation): + """Review C3/P2: duplicate ids in one register_domain call used to be + silently dropped (the high-severity second rule vanished).""" + + class HighSeverity: + severity = "high" + gap = "probe gap" + domain = "probe thing" + + def __init__(self, rule_id): + self.id = rule_id + + def check(self, text): + return None + + with pytest.raises(ValueError, match="register_domain call.*'dup_x'"): + register_domain( + "testdupdom", + {"probe thing": ("widget",), "probe review": ()}, + rules=[HighSeverity("dup_x"), HighSeverity("dup_x")], + ) + assert "testdupdom" not in REGISTRY.domain_names() + + +def test_rule_exception_gets_loud_attribution(registry_isolation): + """Review C2-lite/P4: a rule exception aborts analyze() naming the rule + and its registration origin, with the original exception chained.""" + register_domain("testboom", {"boom thing": ("widget",), "boom review": ()}) + + class ExplodingRule: + id = "test_exploding_rule" + domain = "boom thing" + severity = "low" + gap = None + + def check(self, text): + raise RuntimeError("boom from user rule") + + register_rule(ExplodingRule()) + with pytest.raises(RuntimeError, match="test_exploding_rule.*registered at") as exc_info: + InputGuard().analyze("widget please", domain="testboom") + assert "boom from user rule" in str(exc_info.value.__cause__) + + +def test_extension_api_exported_at_top_level(): + """Review C6: the extension API is importable from the package root.""" + import inputguard + + for name in ("register_rule", "register_domain", "REGISTRY", "Rule"): + assert hasattr(inputguard, name) + assert name in inputguard.__all__ + # the v0.2 compat surface is untouched and leads __all__ + assert inputguard.__all__[:4] == [ + "InputGuard", + "AnalysisResult", + "RuleFinding", + "__version__", + ] From 68f2c30c173c7961dd0392245aa582d9cbe7115d Mon Sep 17 00:00:00 2001 From: "obvious-autobuild[bot]" <262744130+obvious-autobuild[bot]@users.noreply.github.com> Date: Thu, 17 Sep 2026 20:07:20 +0000 Subject: [PATCH 07/18] fix: match terms at word boundaries across detector and all rule modules (#6) * fix(word-boundary-generalization): shared word-boundary matcher across detector and all rule modules Raw substring containment (detector.py:57-58) read "fixture" as the debug signal "fix" (probe P1) and flagged innocent input at 35/needs_clarification. All term lookups now route through inputguard/matching.py, which matches terms as standalone words using the boundary class the coding matcher already used, extended with inflectional endings so suffix-shaped hits (bugs, errors, debugged, refactoring, profiling, stack traces) keep firing. Coding rules keep their v0.2 multiword token fallback via token_fallback=True; other callers get strict adjacent-phrase matching. No-alnum terms (/, =>, ->) keep substring semantics. Finders compile once per term-set (lru_cache). Co-authored-by: Kalisetti Nihanth Naidu * test(fp-regressions): pin word-boundary matcher contract and probe-P1 fix Regression coverage for the shared matcher: the P1 "fixture" input no longer reads as debug and is not hard-flagged; embedded words (fixture, refix, prefix, praised) stop matching; inflected forms (bugs, errors, debugged, refactoring, profiling, stack traces) keep firing so the boundary tightening adds no false negatives; phrases match adjacently with final-word inflections; the coding token fallback stays opt-in; symbol terms keep substring semantics. Co-authored-by: Kalisetti Nihanth Naidu --------- Co-authored-by: Obvious Co-authored-by: Kalisetti Nihanth Naidu --- inputguard/detector.py | 6 +- inputguard/matching.py | 197 +++++++++++++++++++++++++++++++ inputguard/rules/coding.py | 23 +--- inputguard/rules/debug.py | 4 +- inputguard/rules/explanation.py | 4 +- inputguard/rules/feature.py | 4 +- inputguard/rules/optimization.py | 4 +- tests/test_word_boundary.py | 176 +++++++++++++++++++++++++++ 8 files changed, 396 insertions(+), 22 deletions(-) create mode 100644 inputguard/matching.py create mode 100644 tests/test_word_boundary.py diff --git a/inputguard/detector.py b/inputguard/detector.py index 31db70e..a86fe96 100644 --- a/inputguard/detector.py +++ b/inputguard/detector.py @@ -4,6 +4,8 @@ from typing import Iterable, Mapping, Optional, Tuple +from inputguard.matching import contains_any + DEBUG_SIGNALS = { "error", "exception", "traceback", "not working", "isn't working", "doesn't work", "won't work", "broken", "failing", "fails", "failed", @@ -69,7 +71,9 @@ def normalize(text: str) -> str: def _contains_any(text: str, terms) -> bool: - return any(term in text for term in terms) + # v0.3: word-boundary matching shared with the rule modules — "fixture" + # is no longer read as the debug signal "fix" (probe P1). + return contains_any(text, terms) def detect_intent( diff --git a/inputguard/matching.py b/inputguard/matching.py new file mode 100644 index 0000000..36d3ef8 --- /dev/null +++ b/inputguard/matching.py @@ -0,0 +1,197 @@ +"""Shared word-boundary matching for every signal and rule wordlist. + +Before v0.3 the detector and four rule modules matched terms with raw +substring containment (``term in text``), so the debug signal ``"fix"`` read +``"fixture"`` as a debug request and innocent inputs were flagged (probe P1: +``"I love this fixture in the test suite..."`` scored 35 as debug). All term +lookups now go through this module, which matches terms as standalone words +under three guarantees: + +- **Boundaries:** a term matches only at a word boundary on each side, so + ``"fix"`` matches ``"fix this"`` but not ``"fixture"``, ``"refix"``, or + ``"praised"`` (from ``"raised"``). The boundary class is the one the v0.2 + coding matcher already used (``(?``, ``->``, ``::``) have +no word to bound and keep plain substring semantics. + +Text is expected pipeline-normalized (lowercased, whitespace-collapsed); +case is folded defensively so ad-hoc callers get identical verdicts. + +Matching is deterministic and thread-safe: finders are compiled once per +term-set through ``functools.lru_cache`` (compiled finders are immutable and +the cached values are idempotent, so concurrent first calls are safe). +""" + +from __future__ import annotations + +import re +from functools import lru_cache +from typing import Callable, FrozenSet, Iterable, List, Tuple + +__all__ = ["contains_any", "contains_term"] + +# The v0.2 coding matcher's boundary class: a neighboring identifier +# character ("my_error") still counts as a mention; a longer word does not. +_LEFT_BOUND = r"(? "debugged"). +_DOUBLABLE = frozenset("bgmnpt") + + +def _fold(text: str) -> str: + # Internal callers pass normalized text; fold defensively for ad-hoc calls. + return text if text.islower() else text.lower() + + +def _stem_variants(term: str) -> Tuple[str, ...]: + """Stems whose suffixed forms belong to ``term`` ("debug" -> "debugged").""" + variants = {term} + if term.endswith("e"): + variants.add(term[:-1]) # "store" -> "storing" + if term[-1] in _DOUBLABLE: + variants.add(term + term[-1]) # "debug" -> "debugged" + return tuple(sorted(variants)) + + +def _single_word_body(term: str) -> str: + return "|".join(re.escape(stem) + _ENDINGS for stem in _stem_variants(term)) + + +def _single_word_pattern(term: str) -> str: + return _LEFT_BOUND + "(?:" + _single_word_body(term) + ")" + _RIGHT_BOUND + + +def _phrase_body(tokens: Tuple[str, ...]) -> str: + # Adjacent words (v0.2 substring parity); the final word carries the + # inflections so plural phrases ("stack traces") still match the term. + head = " ".join(re.escape(token) for token in tokens[:-1]) + tail = re.escape(tokens[-1]) + _ENDINGS + return head + " " + tail if head else tail + + +def _phrase_pattern(tokens: Tuple[str, ...]) -> str: + return _LEFT_BOUND + "(?:" + _phrase_body(tokens) + ")" + _RIGHT_BOUND + + +def _exact_word_pattern(term: str) -> str: + return _LEFT_BOUND + re.escape(term) + _RIGHT_BOUND + + +def contains_term(text: str, term: str) -> bool: + """Whether ``term`` occurs in ``text`` as a standalone word or phrase.""" + return _term_in(_fold(text), term.strip()) + + +def _term_in(text: str, term: str) -> bool: + if not term: + return False + if not any(ch.isalnum() for ch in term): + # "/", "=>", "->", "::" — no word to bound; v0.2 substring semantics. + return term in text + if " " in term: + return re.search(_phrase_pattern(tuple(term.split())), text) is not None + return re.search(_single_word_pattern(term), text) is not None + + +@lru_cache(maxsize=None) +def _build_finder( + terms: FrozenSet[str], token_fallback: bool +) -> Callable[[str], bool]: + singles: List[str] = [] + phrases: List[Tuple[str, ...]] = [] + plain: List[str] = [] + for term in terms: + stripped = term.strip() + if not stripped: + continue + if not any(ch.isalnum() for ch in stripped): + plain.append(stripped) + elif " " in stripped: + phrases.append(tuple(stripped.split())) + else: + singles.append(stripped) + + # One alternation scan covers every single-word term in the set. + single_re = ( + re.compile( + _LEFT_BOUND + + "(?:" + + "|".join(sorted(_single_word_body(s) for s in singles)) + + ")" + + _RIGHT_BOUND + ) + if singles + else None + ) + phrase_re = ( + re.compile( + _LEFT_BOUND + + "(?:" + + "|".join(sorted(_phrase_body(tokens) for tokens in phrases)) + + ")" + + _RIGHT_BOUND + ) + if phrases + else None + ) + # v0.2 coding parity: fallback tokens match exactly, without inflections. + fallback_res = ( + [ + tuple(re.compile(_exact_word_pattern(token)) for token in tokens) + for tokens in phrases + ] + if token_fallback + else [] + ) + plain_tuple = tuple(plain) + + def find(text: str) -> bool: + if single_re is not None and single_re.search(text): + return True + if any(term in text for term in plain_tuple): + return True + if phrase_re is not None and phrase_re.search(text): + return True + if token_fallback: + for tokens_re in fallback_res: + if all(token_re.search(text) for token_re in tokens_re): + return True + return False + + return find + + +def contains_any( + text: str, terms: Iterable[str], token_fallback: bool = False +) -> bool: + """Whether any of ``terms`` occurs in ``text`` as a standalone word or phrase. + + ``token_fallback=True`` restores the v0.2 coding-rule behavior of matching + a multiword term whose words are all present but not adjacent. + """ + if isinstance(terms, str): + raise TypeError( + "contains_any expects an iterable of terms, not a single string — " + f"got {terms!r}." + ) + return _build_finder(frozenset(terms), token_fallback)(_fold(text)) diff --git a/inputguard/rules/coding.py b/inputguard/rules/coding.py index 85f8fc1..2b3b9ff 100644 --- a/inputguard/rules/coding.py +++ b/inputguard/rules/coding.py @@ -3,6 +3,7 @@ import re from typing import List, Optional, Set +from inputguard.matching import contains_any from inputguard.registry import register_rule from inputguard.types import RuleFinding @@ -110,23 +111,11 @@ def _normalize(text: str) -> str: return re.sub(r"\s+", " ", text.lower()).strip() -def _contains_term(text: str, term: str) -> bool: - if term == "/": - return "/" in text - pattern = r"(? bool: - return any(_contains_term(text, t) for t in terms) +def _contains_any(text: str, terms) -> bool: + # v0.3: word-boundary matching via the shared matcher. token_fallback + # keeps the v0.2 coding-rule behavior for multiword terms ("def ", + # "sign in with"): their words may appear non-adjacent. + return contains_any(text, terms, token_fallback=True) def check_missing_language(text: str) -> Optional[RuleFinding]: diff --git a/inputguard/rules/debug.py b/inputguard/rules/debug.py index f474dbc..7431f15 100644 --- a/inputguard/rules/debug.py +++ b/inputguard/rules/debug.py @@ -4,6 +4,7 @@ from typing import List, Optional from inputguard.detector import DEBUG_SIGNALS +from inputguard.matching import contains_any from inputguard.registry import register_rule from inputguard.types import RuleFinding @@ -43,7 +44,8 @@ def _normalize(text: str) -> str: def _contains_any(text: str, terms) -> bool: - return any(term in text for term in terms) + # v0.3: word-boundary matching via the shared matcher (probe P1 fix). + return contains_any(text, terms) def check_missing_error_message(text: str) -> Optional[RuleFinding]: diff --git a/inputguard/rules/explanation.py b/inputguard/rules/explanation.py index bc933e6..69498de 100644 --- a/inputguard/rules/explanation.py +++ b/inputguard/rules/explanation.py @@ -4,6 +4,7 @@ from typing import List, Optional from inputguard.detector import EXPLANATION_SIGNALS +from inputguard.matching import contains_any from inputguard.registry import register_rule from inputguard.types import RuleFinding @@ -37,7 +38,8 @@ def _normalize(text: str) -> str: def _contains_any(text: str, terms) -> bool: - return any(term in text for term in terms) + # v0.3: word-boundary matching via the shared matcher (probe P1 fix). + return contains_any(text, terms) def check_missing_code_reference(text: str) -> Optional[RuleFinding]: diff --git a/inputguard/rules/feature.py b/inputguard/rules/feature.py index 642ab18..51b6948 100644 --- a/inputguard/rules/feature.py +++ b/inputguard/rules/feature.py @@ -4,6 +4,7 @@ from typing import List, Optional from inputguard.detector import FEATURE_SIGNALS +from inputguard.matching import contains_any from inputguard.registry import register_rule from inputguard.types import RuleFinding @@ -44,7 +45,8 @@ def _normalize(text: str) -> str: def _contains_any(text: str, terms) -> bool: - return any(term in text for term in terms) + # v0.3: word-boundary matching via the shared matcher (probe P1 fix). + return contains_any(text, terms) def check_missing_existing_stack(text: str) -> Optional[RuleFinding]: diff --git a/inputguard/rules/optimization.py b/inputguard/rules/optimization.py index b1d82ed..fdfc587 100644 --- a/inputguard/rules/optimization.py +++ b/inputguard/rules/optimization.py @@ -4,6 +4,7 @@ from typing import List, Optional from inputguard.detector import OPTIMIZATION_SIGNALS +from inputguard.matching import contains_any from inputguard.registry import register_rule from inputguard.types import RuleFinding @@ -43,7 +44,8 @@ def _normalize(text: str) -> str: def _contains_any(text: str, terms) -> bool: - return any(term in text for term in terms) + # v0.3: word-boundary matching via the shared matcher (probe P1 fix). + return contains_any(text, terms) def check_missing_optimization_target(text: str) -> Optional[RuleFinding]: diff --git a/tests/test_word_boundary.py b/tests/test_word_boundary.py new file mode 100644 index 0000000..8130c0a --- /dev/null +++ b/tests/test_word_boundary.py @@ -0,0 +1,176 @@ +"""Word-boundary matching regressions (the probe P1 false-positive class). + +The v0.2 matcher read ``"fixture"`` as the debug signal ``"fix"`` (probe P1: +``"I love this fixture in the test suite, what does it do"`` scored 35 as a +debug request). These tests pin the shared matcher's contract +(``inputguard.matching``): terms match as standalone words, inflected and +suffixed forms still match (no new false negatives at term boundaries), and +unrelated embeddings (``"fixture"``, ``"refix"``, ``"prefix"``) do not. +""" + +from __future__ import annotations + +import pytest + +from inputguard import InputGuard +from inputguard.detector import DEBUG_SIGNALS +from inputguard.matching import contains_any, contains_term + + +P1_INPUT = "I love this fixture in the test suite, what does it do" +DEBUG_CODES = {"missing_error_message", "missing_expected_vs_actual", "missing_debug_code_context"} + + +class TestProbeP1Regression: + """The false positive that motivated the matcher: fixture reads as fix.""" + + def test_p1_input_no_longer_reads_as_debug(self): + result = InputGuard().analyze(P1_INPUT) + assert result.detected_intent != "debug" + assert not any(f.code in DEBUG_CODES for f in result.findings) + + def test_p1_input_not_hard_flagged(self): + # The input is honestly an (underspecified) explanation request, so + # advisory findings are correct — but it must not land in a hard + # flag state in the default warning mode (v0.2: score 35, + # needs_clarification). + result = InputGuard().analyze(P1_INPUT) + assert result.clarity_score >= 60 + assert result.status in ("ready", "usable_with_warnings") + + def test_p1_input_does_not_match_any_debug_signal(self): + assert contains_any(P1_INPUT.lower(), DEBUG_SIGNALS) is False + + +class TestBoundaryPolarity: + """Embeddings that must stop matching (the killed false-positive class).""" + + @pytest.mark.parametrize( + "text", + [ + "the fixture loads slowly", + "my fixtures are missing", + "refix the layout", + "prefix the variable", + ], + ) + def test_embedded_words_do_not_match(self, text): + assert contains_term(text, "fix") is False + + def test_left_embedding_does_not_match(self): + # "raised" embedded in "praised" (the survey's second substring FP). + assert contains_term("praised the release", "raised") is False + + def test_fixture_does_not_match_fix(self): + assert contains_term("the fixture loads slowly", "fix") is False + + def test_suffix_embeddings_are_required(self): + # "fixture" is not a suffix of "fix" — only inflectional endings count. + assert contains_term("fixture", "fix") is False + assert contains_term("fixtures", "fix") is False + + +class TestGenuineHitsStillMatch: + """No new false negatives: real term occurrences keep firing.""" + + @pytest.mark.parametrize( + ("text", "term"), + [ + ("fix the login bug", "fix"), + ("fix, this breaks", "fix"), + ("(fix) applied", "fix"), + ("FIX THIS NOW", "fix"), + ("my_error is misleading", "error"), + ("route /users", "/"), + ("x=>y mapping", "=>"), + ("a->b pointer", "->"), + ], + ) + def test_genuine_hits(self, text, term): + assert contains_term(text, term) is True + + +class TestInflectionsStillMatch: + """Inflectional forms the substring matcher caught must keep matching.""" + + @pytest.mark.parametrize( + ("text", "term"), + [ + ("my code has bugs", "bug"), + ("it raises errors", "error"), + ("i debugged the loop", "debug"), + ("still debugging it", "debug"), + ("the bugfix was debugged twice", "debug"), + ("it crashed twice", "crash"), + ("refactoring the hot loop", "refactor"), + ("we refactored it", "refactor"), + ("profile this code", "profil"), + ("profiling shows the bottleneck", "profil"), + ("the profiler output", "profil"), + ("stack traces attached", "stack trace"), + ("error messages below", "error message"), + ], + ) + def test_inflected_forms(self, text, term): + assert contains_term(text, term) is True + + +class TestMultiwordPhrases: + """Phrases match adjacently (with final-word inflections); the coding + rules' v0.2 token fallback stays opt-in.""" + + def test_adjacent_phrase_matches(self): + assert contains_any("it is not working today", {"not working"}) is True + + def test_non_adjacent_words_do_not_match_by_default(self): + # v0.2 substring never matched these either — adjacency is required + # unless the caller opts into the coding token fallback. + assert contains_any("add authentication to the app", {"add to the"}) is False + + def test_token_fallback_is_opt_in(self): + terms = {"error message"} + text = "the error and the message" + assert contains_any(text, terms) is False + assert contains_any(text, terms, token_fallback=True) is True + + def test_token_fallback_matches_v02_coding_rules(self): + # The v0.2 coding matcher's fallback: multiword term words may appear + # non-adjacent (e.g. "def " token prefixes). + assert contains_any("in def process_order we loop", {"def "}) is True + + +class TestAnalyzeSpotChecks: + """End-to-end: intents and findings survive the matcher change.""" + + @pytest.mark.parametrize( + ("text", "expected_intent"), + [ + ("fix the login bug, it crashes with an error", "debug"), + ("my code has bugs that raise errors", "debug"), + ("the parser failed after I debugged it, stack traces show typeerror", "debug"), + ("we are refactoring the hot loop", "optimization"), + ("profiling shows the bottleneck", "optimization"), + ("Build a REST API using FastAPI. Store users in PostgreSQL with JWT auth and Stripe integration.", "build"), + ], + ) + def test_intent_preserved(self, text, expected_intent): + assert InputGuard().analyze(text).detected_intent == expected_intent + + def test_feature_phrase_does_not_shadow_build(self): + # "add to the" must not fire on non-adjacent words and steal the + # intent from build (regression: coding auth rule stopped running). + result = InputGuard().analyze("add authentication to the app") + assert result.detected_intent == "build" + assert "missing_auth_type" in [f.code for f in result.findings] + + +class TestMatcherContract: + def test_string_terms_rejected(self): + with pytest.raises(TypeError): + contains_any("text", "fix") + + def test_empty_term_never_matches(self): + assert contains_term("anything", "") is False + + def test_case_folded_for_ad_hoc_callers(self): + assert contains_any("FIX the bug", {"fix"}) is True From bee08f02d5aec3f4d424aee9a8209c55806d467e Mon Sep 17 00:00:00 2001 From: "obvious-autobuild[bot]" <262744130+obvious-autobuild[bot]@users.noreply.github.com> Date: Thu, 17 Sep 2026 20:18:29 +0000 Subject: [PATCH 08/18] feat(policy): frozen validated Policy with pinned v0.2 defaults, banding, input cap, and score breakdown (#9) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * feat(policy-object): frozen validated Policy with rule filters and word floor Every scoring tunable moves into a frozen, validated Policy dataclass whose defaults are the v0.2 constants: ready_at 85, usable_at 60, strict_clarify_at 65, penalties 5/15/25, min_words 3, max_chars 10000, borderline_at 74. Validation rejects mis-ordered bands, out-of-range values, uncompilable allow_patterns, and disabled rule ids that are not registered. InputGuard and analyze() take an optional policy (a per-call policy overrides the guard's). disabled_rules skips registered rules by id; allow_patterns regex matches skip flagging entirely; min_words gates the insufficient_context vague rule below its floor. Penalties flow from Policy through calculate_score(findings, policy=None). Closes review concern C5 (art_hC18m78C): one severity vocabulary (policy.SEVERITIES) shared by registration and a new single choke point, scorer.require_known_severity, which the analyzer applies to every emitted finding; registry.KNOWN_SEVERITIES is cross-pinned to it by test. Zero new dependencies; stdlib only. Co-authored-by: Kalisetti Nihanth Naidu * feat(two-layer-banding): policy-driven status bands and a borderline near-miss signal get_status(score, mode, policy=None) now reads its bands from Policy (ready_at 85, usable_at 60, strict_clarify_at 65 — the v0.2 constants, pinned byte-exactly in tests/test_banding.py), replacing the four hard-coded scorer thresholds. Severity still decides what fires; bands decide what happens. AnalysisResult gains an additive borderline: bool field (default False, in to_dict) set when the score lands in [policy.borderline_at, ready_at) — the distinct near-miss "worth one more pass" signal from the spec's result states. interpretation_note is untouched, so existing serialized results are byte-identical under default policy. Co-authored-by: Kalisetti Nihanth Naidu * feat(input-cap): enforce Policy.max_chars with visible truncation analyze() now slices input to policy.max_chars before intent detection and rule execution — analysis cost is bounded regardless of input size (v0.2 took ~3.8 s on a 1.6 MB input; capped analysis of the same input runs in milliseconds). Truncation is never silent: AnalysisResult gains an additive truncated: bool field (default False, in to_dict), and the allowlist is evaluated against the same capped input so no scan escapes the budget. Default cap 10_000 matches the Policy default, pinned by test. Co-authored-by: Kalisetti Nihanth Naidu * feat(score-breakdown): additive score_breakdown audit field on the result calculate_score_with_breakdown(findings, policy=None) scores findings and returns the audit trail — {"base": 100, "penalties": [{"code", "severity", "points"(negative)}...], "final"} — one entry per distinct gap (or code when gap is None) in first-occurrence order, matching the result's gaps order. calculate_score keeps its v0.2 signature and delegates to it, so both stay in lockstep by construction. AnalysisResult gains an additive score_breakdown field (default None, deep- copied in to_dict so a serialized snapshot can never be rewritten by later mutation). With default policy the penalties mirror the v0.2 deductions exactly, so serialized defaults differ only by the new keys. Co-authored-by: Kalisetti Nihanth Naidu * test(pin-v02-defaults): pin every Policy default byte-exactly to v0.2 Dedicated pinning tests guard the calibration-drift risk the spec flags: each Policy default (ready_at 85, usable_at 60, strict_clarify_at 65, penalties 5/15/25, min_words 3, max_chars 10000, borderline_at 74) is asserted three ways — dataclass field defaults, constructed instance values, and the literal source line (underscore separators normalized) — plus a no-unpinned-fields guard so future fields cannot ship uncalibrated. Behavioral pins reproduce the observed v0.2 reference outputs under default policy (75/usable_with_warnings, 35/needs_clarification, 50, strict blocked/needs_clarification) and the v0.2 result and recommendation key shapes. Co-authored-by: Kalisetti Nihanth Naidu * test: integrate policy keys into sibling additive-contract tests The followups parity test and the language ordered-key-list test pin the exhaustive to_dict key set as of their own merge; the policy-calibration fields (borderline, truncated, score_breakdown) are additive per the compat contract, so the expected sets extend rather than the fields disappearing. Co-authored-by: Kalisetti Nihanth Naidu --------- Co-authored-by: Obvious Co-authored-by: Kalisetti Nihanth Naidu --- inputguard/__init__.py | 4 + inputguard/analyzer.py | 120 ++++++++++++---- inputguard/policy.py | 180 ++++++++++++++++++++++++ inputguard/scorer.py | 116 +++++++++++---- inputguard/types.py | 20 ++- tests/test_banding.py | 139 ++++++++++++++++++ tests/test_followups.py | 4 + tests/test_input_cap.py | 90 ++++++++++++ tests/test_language.py | 6 +- tests/test_policy.py | 257 ++++++++++++++++++++++++++++++++++ tests/test_policy_defaults.py | 120 ++++++++++++++++ tests/test_score_breakdown.py | 147 +++++++++++++++++++ 12 files changed, 1142 insertions(+), 61 deletions(-) create mode 100644 inputguard/policy.py create mode 100644 tests/test_banding.py create mode 100644 tests/test_input_cap.py create mode 100644 tests/test_policy.py create mode 100644 tests/test_policy_defaults.py create mode 100644 tests/test_score_breakdown.py diff --git a/inputguard/__init__.py b/inputguard/__init__.py index f62271f..0ad497f 100644 --- a/inputguard/__init__.py +++ b/inputguard/__init__.py @@ -1,4 +1,5 @@ from inputguard.analyzer import InputGuard +from inputguard.policy import Policy from inputguard.registry import REGISTRY, Rule, register_domain, register_rule from inputguard.types import AnalysisResult, RuleFinding @@ -16,4 +17,7 @@ "Rule", "register_rule", "register_domain", + # v0.3 policy calibration (additive): the frozen tuning data for scoring, + # status bands, rule filters, and the input cap. + "Policy", ] diff --git a/inputguard/analyzer.py b/inputguard/analyzer.py index d518443..5b68c44 100644 --- a/inputguard/analyzer.py +++ b/inputguard/analyzer.py @@ -1,6 +1,7 @@ from __future__ import annotations -from typing import List, Set +import re +from typing import List, Optional, Set import inputguard.rules # noqa: F401 — importing registers the coding domain and the 19 built-in rules from inputguard.detector import detect_intent, normalize @@ -14,13 +15,21 @@ partial_coverage_note, probe_script, ) +from inputguard.policy import Policy from inputguard.recommender import get_recommendations from inputguard.registry import REGISTRY -from inputguard.scorer import calculate_score, get_status +from inputguard.scorer import ( + calculate_score_with_breakdown, + get_status, + require_known_severity, +) from inputguard.types import AnalysisResult, RuleFinding _VALID_MODES = {"warning", "strict"} +# The one built-in rule that flags input as vague. Policy.min_words gates it: +# below the floor, input is valid-but-short — never "vague" (spec art_bTvdPdJS §2). +_VAGUE_RULE_ID = "insufficient_context" _INTERPRETATION_NOTE = ( "This input is ambiguous in multiple ways. Addressing each gap below " "before sending will prevent the AI from making assumptions that lead " @@ -29,20 +38,52 @@ class InputGuard: - def __init__(self, mode: str = "warning") -> None: + """The clarity engine: ``analyze(user_input, domain=...) -> AnalysisResult``. + + ``mode`` selects the status banding (warning never blocks, strict does). + ``policy`` is optional :class:`~inputguard.Policy` calibration — status + bands, severity penalties, rule filters, and the input cap. It defaults to + the v0.2 constants; every guard and every ``analyze()`` call may carry its + own (a per-call ``policy`` overrides the guard's). + """ + + def __init__(self, mode: str = "warning", policy: Optional[Policy] = None) -> None: if mode not in _VALID_MODES: raise ValueError( f"Invalid mode: {mode!r}. Expected one of: 'warning', 'strict'." ) + if policy is not None and not isinstance(policy, Policy): + raise TypeError( + f"policy must be a Policy instance, got {type(policy).__name__}." + ) self.mode = mode - - def analyze(self, user_input: str, domain: str = "coding") -> AnalysisResult: + self.policy: Policy = Policy() if policy is None else policy + + def analyze( + self, + user_input: str, + domain: str = "coding", + policy: Optional[Policy] = None, + ) -> AnalysisResult: if not isinstance(user_input, str): raise TypeError( f"user_input must be a string, got {type(user_input).__name__}." ) if not user_input.strip(): raise ValueError("user_input must be a non-empty, non-whitespace string.") + if policy is not None and not isinstance(policy, Policy): + raise TypeError( + f"policy must be a Policy instance, got {type(policy).__name__}." + ) + + effective_policy = self.policy if policy is None else policy + + # Enforce the input cap: analysis is bounded, always. Truncation is + # visible on the result (AnalysisResult.truncated), never silent. + max_chars = effective_policy.max_chars + truncated = len(user_input) > max_chars + if truncated: + user_input = user_input[:max_chars] domain_signals = REGISTRY.get_domain_signals(domain) normalized = normalize(user_input) @@ -58,7 +99,7 @@ def analyze(self, user_input: str, domain: str = "coding") -> AnalysisResult: # modes; the note says honestly what the tool does not know. score = max(0, 100 - DEGRADATION_PENALTY) return AnalysisResult( - status=get_status(score, self.mode), + status=get_status(score, self.mode, effective_policy), clarity_score=score, detected_intent=DEGRADED_INTENT, gaps=[], @@ -74,28 +115,43 @@ def analyze(self, user_input: str, domain: str = "coding") -> AnalysisResult: findings: List[RuleFinding] = [] seen_codes: Set[str] = set() - for rule in REGISTRY.rules_for_intent(detected_intent): - try: - finding = rule.check(normalized) - except Exception as exc: - # Loud failure with attribution (review C2-lite): a rule - # exception is never swallowed or silently degraded around — - # it aborts analyze(), naming the rule and where it was - # registered, with the original traceback chained. - raise RuntimeError( - f"inputguard rule {rule.id!r} " - f"(registered {REGISTRY.rule_origin(rule.id)}) raised " - f"{type(exc).__name__}: {exc}. A rule exception aborts " - f"analyze() by design — rules are the author's " - f"responsibility after registration; fix or remove the rule." - ) from exc - if finding is None or finding.code in seen_codes: - continue - seen_codes.add(finding.code) - findings.append(finding) - - score = calculate_score(findings) - status = get_status(score, self.mode) + if not _matches_allow_pattern(user_input, effective_policy): + word_count = len(normalized.split()) + for rule in REGISTRY.rules_for_intent(detected_intent): + if rule.id in effective_policy.disabled_rules: + continue + if rule.id == _VAGUE_RULE_ID and word_count < effective_policy.min_words: + continue + try: + finding = rule.check(normalized) + except Exception as exc: + # Loud failure with attribution (review C2-lite): a rule + # exception is never swallowed or silently degraded around — + # it aborts analyze(), naming the rule and where it was + # registered, with the original traceback chained. + raise RuntimeError( + f"inputguard rule {rule.id!r} " + f"(registered {REGISTRY.rule_origin(rule.id)}) raised " + f"{type(exc).__name__}: {exc}. A rule exception aborts " + f"analyze() by design — rules are the author's " + f"responsibility after registration; fix or remove the rule." + ) from exc + if finding is None: + continue + # C5 choke point: every emitted finding's severity must be in + # the same Policy-owned vocabulary registration validated. + require_known_severity(finding.severity) + if finding.code in seen_codes: + continue + seen_codes.add(finding.code) + findings.append(finding) + + score, breakdown = calculate_score_with_breakdown(findings, effective_policy) + status = get_status(score, self.mode, effective_policy) + # Two layers: severity decided what fired; the policy's bands decide + # what happens. The borderline band is the near-miss signal just below + # ready — distinct "worth one more pass" messaging (spec art_bTvdPdJS). + borderline = effective_policy.borderline_at <= score < effective_policy.ready_at gaps: List[str] = [] seen: Set[str] = set() @@ -130,4 +186,12 @@ def analyze(self, user_input: str, domain: str = "coding") -> AnalysisResult: if probe.heuristic_coverage == COVERAGE_PARTIAL else None ), + borderline=borderline, + truncated=truncated, + score_breakdown=breakdown, ) + + +def _matches_allow_pattern(text: str, policy: Policy) -> bool: + """True when the analyzed input matches any allowlist pattern — such input is never flagged.""" + return any(re.search(pattern, text) for pattern in policy.allow_patterns) diff --git a/inputguard/policy.py b/inputguard/policy.py new file mode 100644 index 0000000..ea302de --- /dev/null +++ b/inputguard/policy.py @@ -0,0 +1,180 @@ +"""Policy: InputGuard calibration as data, not constants. + +A :class:`Policy` bundles every tunable the ``analyze()`` pipeline consumes — +status bands, severity penalties, the input cap, rule filters, and the +borderline near-miss band — in one frozen, validated object. Every default is +the v0.2 constant, pinned byte-exactly by ``tests/test_policy_defaults.py`` so +existing behavior cannot drift. + +Two layers stay separate: per-rule severity decides *what fires*; the policy's +bands decide *what happens* to the score. +""" + +from __future__ import annotations + +import re +from dataclasses import dataclass +from typing import FrozenSet, Optional, Tuple + +__all__ = ["Policy"] + + +SEVERITIES: Tuple[str, ...] = ("low", "medium", "high") +"""The one severity vocabulary, owned here and shared by both validation sites +(review concern C5, art_hC18m78C): registration validates a rule's *declared* +severity against it (``registry.KNOWN_SEVERITIES``, cross-pinned by +``tests/test_policy.py``), and scoring validates every *emitted* finding's +severity against it (``scorer.require_known_severity``) — so the two checks +cannot drift apart. Deliberately not a ``Policy`` field: a per-instance +vocabulary would be silently ignored by the penalty lookup. Extend only +additively, with a matching ``penalty_*`` field.""" + + +@dataclass(frozen=True) +class Policy: + """Calibration for one :class:`~inputguard.InputGuard`. + + Every default is the v0.2 constant (pinned by ``tests/test_policy_defaults.py``). + Instances are frozen, validated at construction, and safe to share across + threads and guards. + + Fields + ------ + ready_at: + A score at or above this is ``ready`` in both modes. + usable_at: + Warning-mode floor for ``usable_with_warnings``. + strict_clarify_at: + Strict-mode floor for ``needs_clarification``; below it, strict blocks. + penalty_low / penalty_medium / penalty_high: + Points deducted per distinct gap, by the highest-severity finding in + that gap. Must satisfy ``penalty_low <= penalty_medium <= penalty_high``. + min_words: + Inputs with fewer words are never flagged as vague by the built-in + ``insufficient_context`` rule — short input is valid-but-short, not + vague. That rule keeps a structural minimum of 3 words, so values + below 3 cannot force shorter input to be flagged. + max_chars: + Input cap. ``analyze()`` inspects at most this many characters; + truncation is reported on the result (``AnalysisResult.truncated``), + never silent. + borderline_at: + Near-miss band: a score in ``[borderline_at, ready_at)`` sets the + result's ``borderline`` signal — distinct "worth one more pass" + messaging just below the ready floor. Set ``borderline_at`` equal to + ``ready_at`` to disable the signal. + disabled_rules: + Rule ids to skip entirely during ``analyze()``. Every id must be a + registered rule id (validated at construction). + allow_patterns: + Regex patterns; analyzed input matching any of them is never flagged — + rules are skipped and the result is ``ready``. Patterns are matched + against the analyzed (length-capped) input via ``re.search``. + """ + + ready_at: int = 85 + usable_at: int = 60 + strict_clarify_at: int = 65 + penalty_low: int = 5 + penalty_medium: int = 15 + penalty_high: int = 25 + min_words: int = 3 + max_chars: int = 10_000 + borderline_at: int = 74 + disabled_rules: FrozenSet[str] = frozenset() + allow_patterns: Tuple[str, ...] = () + + def __post_init__(self) -> None: + # A bare string would silently explode into characters under + # frozenset()/tuple() — reject it before coercion. + if isinstance(self.disabled_rules, str): + raise ValueError( + "Policy.disabled_rules must be a collection of rule ids, e.g. " + f"frozenset({{'missing_language'}}), got the string {self.disabled_rules!r}." + ) + if isinstance(self.allow_patterns, str): + raise ValueError( + "Policy.allow_patterns must be a collection of regex strings, " + f"e.g. ('^re:',), got the string {self.allow_patterns!r}." + ) + # Coerce to immutable containers so a later mutation of the caller's + # set or list cannot change a "frozen" policy mid-flight. + object.__setattr__(self, "disabled_rules", frozenset(self.disabled_rules)) + object.__setattr__(self, "allow_patterns", tuple(self.allow_patterns)) + _validate(self) + + +def _check_int(policy: Policy, name: str, *, minimum: int, maximum: Optional[int]) -> None: + value = getattr(policy, name) + if isinstance(value, bool) or not isinstance(value, int): + raise ValueError(f"Policy.{name} must be an int, got {value!r}.") + if maximum is None: + if value < minimum: + raise ValueError(f"Policy.{name} must be >= {minimum}, got {value}.") + elif not minimum <= value <= maximum: + raise ValueError(f"Policy.{name} must be in [{minimum}, {maximum}], got {value}.") + + +def _validate(policy: Policy) -> None: + for name in ("ready_at", "usable_at", "strict_clarify_at", "borderline_at"): + _check_int(policy, name, minimum=0, maximum=100) + if not policy.usable_at <= policy.strict_clarify_at < policy.ready_at: + raise ValueError( + "Policy bands are mis-ordered: they must satisfy " + "usable_at <= strict_clarify_at < ready_at, got " + f"usable_at={policy.usable_at}, strict_clarify_at={policy.strict_clarify_at}, " + f"ready_at={policy.ready_at}." + ) + if not policy.usable_at <= policy.borderline_at <= policy.ready_at: + raise ValueError( + "Policy.borderline_at must satisfy usable_at <= borderline_at <= ready_at, got " + f"usable_at={policy.usable_at}, borderline_at={policy.borderline_at}, " + f"ready_at={policy.ready_at}. Set borderline_at equal to ready_at to " + "disable the borderline signal." + ) + + for name in ("penalty_low", "penalty_medium", "penalty_high"): + _check_int(policy, name, minimum=0, maximum=None) + if not policy.penalty_low <= policy.penalty_medium <= policy.penalty_high: + raise ValueError( + "Policy penalties are mis-ordered: they must satisfy " + "penalty_low <= penalty_medium <= penalty_high, got " + f"penalty_low={policy.penalty_low}, penalty_medium={policy.penalty_medium}, " + f"penalty_high={policy.penalty_high}." + ) + + _check_int(policy, "min_words", minimum=0, maximum=None) + _check_int(policy, "max_chars", minimum=1, maximum=None) + + for pattern in policy.allow_patterns: + if not isinstance(pattern, str): + raise ValueError( + f"Policy.allow_patterns entries must be regex strings, got {pattern!r}." + ) + try: + re.compile(pattern) + except re.error as exc: + raise ValueError( + f"Policy.allow_patterns entry {pattern!r} is not a valid regex: {exc}" + ) from exc + + for rule_id in policy.disabled_rules: + if not isinstance(rule_id, str): + raise ValueError( + f"Policy.disabled_rules entries must be rule id strings, got {rule_id!r}." + ) + + if policy.disabled_rules: + # Local import: the registry is populated as a side effect of importing + # inputguard.rules, so the known-id check must run at construction time, + # not module-import time (and registry imports nothing from this module). + from inputguard.registry import REGISTRY + + registered = REGISTRY.rule_ids() + unknown = sorted(set(policy.disabled_rules) - set(registered)) + if unknown: + raise ValueError( + "Policy.disabled_rules references unknown rule ids: " + f"{', '.join(repr(r) for r in unknown)}. Registered rule ids: " + f"{', '.join(repr(r) for r in registered)}." + ) diff --git a/inputguard/scorer.py b/inputguard/scorer.py index d6894ad..19e7226 100644 --- a/inputguard/scorer.py +++ b/inputguard/scorer.py @@ -1,70 +1,124 @@ from __future__ import annotations -from typing import Dict, List +from typing import Any, Dict, List, Optional, Tuple +from inputguard.policy import Policy, SEVERITIES from inputguard.types import RuleFinding SEVERITY_PENALTIES = {"low": 5, "medium": 15, "high": 25} +"""The v0.2 default penalty table — now mirrored by :class:`~inputguard.Policy`'s +penalty defaults (cross-pinned by ``tests/test_policy_defaults.py``). Kept for +backward compatibility of the module surface.""" -_SEVERITY_RANK = {"low": 0, "medium": 1, "high": 2} +_SEVERITY_RANK = {name: rank for rank, name in enumerate(SEVERITIES)} -# Warning mode thresholds -_WARN_READY = 85 -_WARN_USABLE = 60 -# Strict mode thresholds -_STRICT_READY = 85 -_STRICT_NEEDS = 65 +def require_known_severity(severity: str) -> None: + """The single choke point validating a severity against ``Policy.SEVERITIES``. + + Two validations share this one vocabulary (review concern C5, + art_hC18m78C): registration validates a rule's *declared* severity + (``registry.KNOWN_SEVERITIES``, cross-pinned by ``tests/test_policy.py``), + and the analyzer/scorer validate every *emitted* finding's severity here — + so the two checks cannot drift apart. + """ + if severity not in _SEVERITY_RANK: + raise ValueError( + f"Unknown severity: {severity!r}. Expected one of: " + f"{', '.join(repr(s) for s in SEVERITIES)}." + ) def _severity_rank(severity: str) -> int: """Rank lookup that fails loudly: an unknown severity must never silently score zero.""" - try: - return _SEVERITY_RANK[severity] - except KeyError: - raise ValueError( - f"Unknown severity: {severity!r}. Expected one of: 'low', 'medium', 'high'." - ) from None + require_known_severity(severity) + return _SEVERITY_RANK[severity] -def calculate_score(findings: List[RuleFinding]) -> int: +def calculate_score(findings: List[RuleFinding], policy: Optional[Policy] = None) -> int: """100 minus one penalty per distinct gap (or code, when gap is None), clamped to [0, 100]. - Unknown severities raise ``ValueError`` — a typo'd severity would - otherwise distort every score silently (the v0.2 ``.get(severity, 0)`` + Penalties come from ``policy`` (defaults reproduce the v0.2 constants + byte-exactly). Unknown severities raise ``ValueError`` — a typo'd severity + would otherwise distort every score silently (the v0.2 ``.get(severity, 0)`` behavior). """ + score, _ = calculate_score_with_breakdown(findings, policy) + return score + + +def calculate_score_with_breakdown( + findings: List[RuleFinding], policy: Optional[Policy] = None +) -> Tuple[int, Dict[str, Any]]: + """Score the findings and also return the additive audit breakdown. + + The breakdown is ``{"base": 100, "penalties": [...], "final": score}`` + where each penalty records ``{"code", "severity", "points"}`` (negative) — + one entry per distinct gap (or code, when gap is None), in first-occurrence + order. Ops teams use it to audit exactly which finding cost which points. + Unknown severities raise ``ValueError`` before any scoring happens. + """ + p = Policy() if policy is None else policy # Validate every severity before scoring so a bad finding fails loudly # even when a later finding would otherwise mask it in the dedup loop. for f in findings: _severity_rank(f.severity) - highest_by_gap: Dict[str, str] = {} + highest_by_gap: Dict[str, RuleFinding] = {} for f in findings: key = f.gap if f.gap is not None else f.code current = highest_by_gap.get(key) - if current is None or _severity_rank(f.severity) > _severity_rank(current): - highest_by_gap[key] = f.severity + if current is None or _severity_rank(f.severity) > _severity_rank( + current.severity + ): + highest_by_gap[key] = f score = 100 - for severity in highest_by_gap.values(): - score -= SEVERITY_PENALTIES[severity] - - return max(0, min(100, score)) - - -def get_status(score: int, mode: str) -> str: + penalties: List[Dict[str, Any]] = [] + for f in highest_by_gap.values(): + points = _penalty_for(f.severity, p) + penalties.append({"code": f.code, "severity": f.severity, "points": -points}) + score -= points + + final = max(0, min(100, score)) + return final, {"base": 100, "penalties": penalties, "final": final} + + +def _penalty_for(severity: str, policy: Policy) -> int: + """Per-distinct-gap penalty for a severity already validated by ``_severity_rank``.""" + if severity == "low": + return policy.penalty_low + if severity == "medium": + return policy.penalty_medium + if severity == "high": + return policy.penalty_high + raise ValueError( + f"Unknown severity: {severity!r}. Expected one of: 'low', 'medium', 'high'." + ) + + +def get_status(score: int, mode: str, policy: Optional[Policy] = None) -> str: + """Map a score to a status using the mode's policy bands. + + Bands come from ``policy`` (defaults reproduce the v0.2 constants + byte-exactly): warning mode — ``ready`` at ``ready_at`` (85), then + ``usable_with_warnings`` at ``usable_at`` (60); strict mode — ``ready`` + at ``ready_at`` (85), then ``needs_clarification`` at ``strict_clarify_at`` + (65), else ``blocked``. Two layers stay separate: severity decided what + fired; these bands decide what happens. + """ + p = Policy() if policy is None else policy if mode == "warning": - if score >= _WARN_READY: + if score >= p.ready_at: return "ready" - if score >= _WARN_USABLE: + if score >= p.usable_at: return "usable_with_warnings" return "needs_clarification" if mode == "strict": - if score >= _STRICT_READY: + if score >= p.ready_at: return "ready" - if score >= _STRICT_NEEDS: + if score >= p.strict_clarify_at: return "needs_clarification" return "blocked" raise ValueError(f"Unknown mode: {mode!r}. Expected 'warning' or 'strict'.") diff --git a/inputguard/types.py b/inputguard/types.py index 573046b..73bde72 100644 --- a/inputguard/types.py +++ b/inputguard/types.py @@ -1,7 +1,7 @@ from __future__ import annotations from dataclasses import dataclass, field, asdict -from typing import List, Optional +from typing import Any, Dict, List, Optional @dataclass(frozen=True) @@ -12,6 +12,15 @@ class RuleFinding: gap: Optional[str] = None +def _copy_breakdown(breakdown: Optional[Dict[str, Any]]) -> Optional[Dict[str, Any]]: + """Copy the breakdown so a result and its to_dict never share mutable state.""" + if breakdown is None: + return None + copied = dict(breakdown) + copied["penalties"] = [dict(p) for p in breakdown["penalties"]] + return copied + + @dataclass(frozen=True) class AnalysisResult: status: str @@ -32,6 +41,12 @@ class AnalysisResult: detected_language: str = "en" heuristic_coverage: str = "full" degradation_note: Optional[str] = None + # v0.3 policy calibration (additive): near-miss signal — True when the + # score lands in [policy.borderline_at, policy.ready_at), the "worth one + # more pass" band just below ready. + borderline: bool = False + truncated: bool = False + score_breakdown: Optional[Dict[str, Any]] = None def to_dict(self) -> dict: return { @@ -46,6 +61,9 @@ def to_dict(self) -> dict: "detected_language": self.detected_language, "heuristic_coverage": self.heuristic_coverage, "degradation_note": self.degradation_note, + "borderline": self.borderline, + "truncated": self.truncated, + "score_breakdown": _copy_breakdown(self.score_breakdown), } def is_clear(self) -> bool: diff --git a/tests/test_banding.py b/tests/test_banding.py new file mode 100644 index 0000000..f183cda --- /dev/null +++ b/tests/test_banding.py @@ -0,0 +1,139 @@ +"""Two-layer banding tests: policy-driven status bands and the borderline signal. + +Covers the banding slice of the v0.3 spec (art_bTvdPdJS §2 and Result states): +severity decides what fires, Policy bands decide what happens. Byte-exact +default pinning lives in tests/test_policy_defaults.py. +""" + +from __future__ import annotations + +import dataclasses + +import pytest + +from inputguard import InputGuard, Policy +from inputguard.scorer import get_status + +# Observed via runtime probe: "do something now" scores 75 (one high finding). +SINGLE_HIGH_INPUT = "do something now" + + +# --- default bands reproduce v0.2 warning mode ------------------------------- + + +def test_default_warning_bands_byte_exact(): + assert get_status(100, "warning") == "ready" + assert get_status(85, "warning") == "ready" + assert get_status(84, "warning") == "usable_with_warnings" + assert get_status(60, "warning") == "usable_with_warnings" + assert get_status(59, "warning") == "needs_clarification" + assert get_status(0, "warning") == "needs_clarification" + + +def test_default_strict_bands_byte_exact(): + assert get_status(100, "strict") == "ready" + assert get_status(85, "strict") == "ready" + assert get_status(84, "strict") == "needs_clarification" + assert get_status(65, "strict") == "needs_clarification" + assert get_status(64, "strict") == "blocked" + assert get_status(0, "strict") == "blocked" + + +def test_default_policy_get_status_matches_no_policy(): + for mode in ("warning", "strict"): + for score in (0, 59, 60, 64, 65, 74, 84, 85, 100): + assert get_status(score, mode) == get_status(score, mode, Policy()) + + +# --- bands are tunable data --------------------------------------------------- + + +def test_custom_bands_change_warning_mode(): + policy = Policy(ready_at=90, usable_at=40) + assert get_status(85, "warning", policy) == "usable_with_warnings" + assert get_status(90, "warning", policy) == "ready" + assert get_status(39, "warning", policy) == "needs_clarification" + + +def test_custom_bands_change_strict_mode(): + policy = Policy(ready_at=95, strict_clarify_at=75) + assert get_status(80, "strict", policy) == "needs_clarification" + assert get_status(95, "strict", policy) == "ready" + assert get_status(74, "strict", policy) == "blocked" + + +def test_relaxed_bands_never_block_in_warning_mode(): + # Warning mode never returns blocked for any score, whatever the bands. + policy = Policy(usable_at=0, strict_clarify_at=0) + for score in range(0, 101): + assert get_status(score, "warning", policy) != "blocked" + + +# --- the borderline near-miss signal ------------------------------------------ + + +def test_default_borderline_band_is_74_to_84(): + guard = InputGuard() + result = guard.analyze(SINGLE_HIGH_INPUT) + assert result.clarity_score == 75 + assert result.borderline is True + assert result.status == "usable_with_warnings" + + +def test_scores_outside_the_band_are_not_borderline(): + guard = InputGuard() + below = guard.analyze("fix the bug in my app") # 35 + assert below.clarity_score == 35 + assert below.borderline is False + clean = guard.analyze( + "Build a web app using React and FastAPI. " + "It needs user authentication with JWT. " + "Store tasks in a PostgreSQL database with title, description, and due date fields. " + "Expose REST API endpoints: GET /tasks, POST /tasks, DELETE /tasks/{id}." + ) # 100 + assert clean.clarity_score == 100 + assert clean.borderline is False + + +def test_borderline_at_bounds(): + # Band floor inclusive: borderline_at == score -> borderline. + assert ( + InputGuard(policy=Policy(borderline_at=75)) + .analyze(SINGLE_HIGH_INPUT) + .borderline + is True + ) + # Disabled: borderline_at == ready_at means no score can be borderline. + disabled = InputGuard(policy=Policy(borderline_at=85)) + result = disabled.analyze(SINGLE_HIGH_INPUT) + assert result.borderline is False + + +def test_borderline_works_in_strict_mode_too(): + # Score 75 in strict mode is needs_clarification (65-84) and inside [74, 85). + result = InputGuard(mode="strict").analyze(SINGLE_HIGH_INPUT) + assert result.status == "needs_clarification" + assert result.borderline is True + + +def test_borderline_field_is_additive_on_to_dict(): + result = InputGuard().analyze(SINGLE_HIGH_INPUT).to_dict() + assert result["borderline"] is True + # Every v0.2 key is still present. + assert set(result).issuperset( + { + "status", + "clarity_score", + "detected_intent", + "gaps", + "recommendations", + "findings", + "interpretation_note", + } + ) + + +def test_frozen_result_rejects_mutation(): + result = InputGuard().analyze(SINGLE_HIGH_INPUT) + with pytest.raises(dataclasses.FrozenInstanceError): + result.borderline = False # type: ignore[misc] diff --git a/tests/test_followups.py b/tests/test_followups.py index 2f93711..f468deb 100644 --- a/tests/test_followups.py +++ b/tests/test_followups.py @@ -298,6 +298,10 @@ def test_english_score_parity_with_v02(text): "detected_language", "heuristic_coverage", "degradation_note", + # policy-calibration PR: borderline band, input cap, score breakdown + "borderline", + "truncated", + "score_breakdown", } v02_shaped = AnalysisResult( status=result.status, diff --git a/tests/test_input_cap.py b/tests/test_input_cap.py new file mode 100644 index 0000000..12b1800 --- /dev/null +++ b/tests/test_input_cap.py @@ -0,0 +1,90 @@ +"""Input-cap tests: enforced max_chars truncation, visible in the result. + +Covers the input-length slice of the v0.3 spec (art_bTvdPdJS §2): analyzed +input is bounded, always, and truncation is reported on the result — never +silent. The v0.2 code had no cap at all (a 1.6 MB input took ~3.8 s). +""" + +from __future__ import annotations + +import time + +from inputguard import InputGuard, Policy + +CLEAN_INPUT = ( + "Build a web app using React and FastAPI. " + "It needs user authentication with JWT. " + "Store tasks in a PostgreSQL database with title, description, and due date fields. " + "Expose REST API endpoints: GET /tasks, POST /tasks, DELETE /tasks/{id}." +) + + +def test_input_at_cap_is_not_truncated(): + guard = InputGuard() + assert guard.analyze("a" * 10_000).truncated is False + + +def test_input_one_char_over_cap_is_truncated(): + guard = InputGuard() + assert guard.analyze("a" * 10_001).truncated is True + + +def test_truncation_visible_in_to_dict(): + assert InputGuard().analyze("x" * 10_001).to_dict()["truncated"] is True + + +def test_default_cap_is_the_v02_compatible_10k(): + assert Policy().max_chars == 10_000 + + +def test_custom_cap_truncates_at_its_boundary(): + guard = InputGuard(policy=Policy(max_chars=50)) + assert guard.analyze("x" * 51).truncated is True + assert guard.analyze("x" * 50).truncated is False + + +def test_truncated_analysis_equals_analyzing_the_kept_prefix(): + # Only the truncation flag may differ — the cap is a pure prefix slice. + prefix = "fix the bug in my app" + huge = prefix + " and some filler that would change nothing " * 400 + guard = InputGuard() + full = guard.analyze(huge).to_dict() + head = guard.analyze(huge[:10_000]).to_dict() + assert full["truncated"] is True + assert head["truncated"] is False + full.pop("truncated") + head.pop("truncated") + assert full == head + + +def test_findings_come_from_the_kept_prefix(): + # Flag-worthy content inside the cap still fires after truncation. + guard = InputGuard(policy=Policy(max_chars=100)) + result = guard.analyze("fix the bug in my app" + " z" * 200) + assert result.truncated is True + assert result.findings[0].code == "missing_error_message" + + +def test_allowlist_evaluated_against_the_capped_input(): + # The allow pattern must match within the analyzed (capped) region. + guard = InputGuard(policy=Policy(max_chars=10, allow_patterns=("^fix",))) + result = guard.analyze("fix the bug in my app" + "z" * 500) + assert result.truncated is True + assert result.findings == [] + assert result.status == "ready" + + +def test_clean_short_input_reports_no_truncation(): + result = InputGuard().analyze(CLEAN_INPUT) + assert result.truncated is False + assert result.status == "ready" + + +def test_huge_input_analysis_is_bounded_by_the_cap(): + # v0.2 took ~3.8 s on a 1.6 MB input; the cap must bound the scan. + huge = ("make my database query faster " * 55_000)[:1_600_000] + start = time.perf_counter() + result = InputGuard().analyze(huge) + elapsed = time.perf_counter() - start + assert result.truncated is True + assert elapsed < 1.0 diff --git a/tests/test_language.py b/tests/test_language.py index d235d0f..997a1ab 100644 --- a/tests/test_language.py +++ b/tests/test_language.py @@ -321,7 +321,8 @@ def test_to_dict_includes_additive_fields(): assert isinstance(d["degradation_note"], str) # Merged additive contract: v0.2 keys keep their names and relative # order, follow_ups slots in after recommendations (questions engine), - # and the three probe keys trail. + # the three probe keys trail, and the policy-calibration keys append + # after those (borderline band, input cap, score breakdown). assert list(d) == [ "status", "clarity_score", @@ -334,6 +335,9 @@ def test_to_dict_includes_additive_fields(): "detected_language", "heuristic_coverage", "degradation_note", + "borderline", + "truncated", + "score_breakdown", ] json.dumps(d, ensure_ascii=False) # JSON-serializable as before diff --git a/tests/test_policy.py b/tests/test_policy.py new file mode 100644 index 0000000..0fc1622 --- /dev/null +++ b/tests/test_policy.py @@ -0,0 +1,257 @@ +"""Policy object tests: validation, filtering, severity vocabulary, and plumbing. + +Covers the policy-object slice of the v0.3 spec (art_bTvdPdJS §2) plus the C5 +closure from the adversarial API review (art_hC18m78C): one Policy-owned +severity vocabulary shared by registration and scoring. Byte-exact default +pinning lives in tests/test_policy_defaults.py. +""" + +from __future__ import annotations + +import dataclasses +from concurrent.futures import ThreadPoolExecutor + +import pytest + +from inputguard import InputGuard, Policy, RuleFinding +from inputguard.policy import SEVERITIES +from inputguard.registry import KNOWN_SEVERITIES, register_rule +from inputguard.scorer import require_known_severity +from test_registry import _test_rule, registry_isolation # noqa: F401 — pytest fixture + +# Observed v0.2 / release-branch behavior (runtime probes), reused as fixtures. +VAGUE_INPUT = "do something now" # build intent, one high finding: insufficient_context +SINGLE_HIGH_INPUT = "build a REST API using Python" # build intent, one high finding + + +# --- construction, coercion, and validation -------------------------------- + + +def test_policy_is_frozen(): + policy = Policy() + with pytest.raises(dataclasses.FrozenInstanceError): + policy.ready_at = 90 # type: ignore[misc] + + +def test_disabled_rules_coerced_to_frozenset_and_isolated(): + source = {"missing_language"} + policy = Policy(disabled_rules=source) # type: ignore[arg-type] + assert isinstance(policy.disabled_rules, frozenset) + source.add("missing_api_structure") # mutating the source must not leak in + assert policy.disabled_rules == frozenset({"missing_language"}) + + +def test_allow_patterns_coerced_to_tuple_and_isolated(): + source = ["^re:"] + policy = Policy(allow_patterns=source) # type: ignore[arg-type] + assert isinstance(policy.allow_patterns, tuple) + source.append("x") + assert policy.allow_patterns == ("^re:",) + + +def test_string_disabled_rules_rejected(): + # A bare string would silently become a set of characters. + with pytest.raises(ValueError, match="collection of rule ids"): + Policy(disabled_rules="missing_language") # type: ignore[arg-type] + + +def test_string_allow_patterns_rejected(): + with pytest.raises(ValueError, match="collection of regex"): + Policy(allow_patterns="^re:") # type: ignore[arg-type] + + +@pytest.mark.parametrize( + "kwargs, match", + [ + ({"usable_at": 70, "strict_clarify_at": 65}, "mis-ordered"), + ({"strict_clarify_at": 90}, "mis-ordered"), + ({"usable_at": 90}, "mis-ordered"), + ({"ready_at": 101}, "ready_at"), + ({"usable_at": -1}, "usable_at"), + ({"borderline_at": 90, "ready_at": 85}, "borderline_at"), + ({"borderline_at": 50}, "borderline_at"), + ({"penalty_high": -1}, "penalty_high"), + ({"penalty_low": 30}, "penalties are mis-ordered"), + ({"penalty_medium": 30, "penalty_high": 25}, "penalties are mis-ordered"), + ({"min_words": -1}, "min_words"), + ({"max_chars": 0}, "max_chars"), + ({"ready_at": "85"}, "must be an int"), + ({"disabled_rules": {"no_such_rule_id"}}, "unknown rule ids"), + ({"allow_patterns": ("(",)}, "not a valid regex"), + ({"allow_patterns": (5,)}, "must be regex strings"), + ({"disabled_rules": {5}}, "must be rule id strings"), + ], +) +def test_invalid_policy_rejected_with_value_error(kwargs, match): + with pytest.raises(ValueError, match=match): + Policy(**kwargs) + + +def test_band_edges_are_valid(): + policy = Policy(usable_at=0, strict_clarify_at=0, ready_at=100, borderline_at=0) + assert policy.ready_at == 100 + + +def test_borderline_disabled_by_setting_equal_to_ready_at(): + policy = Policy(borderline_at=85) + assert policy.borderline_at == policy.ready_at + + +def test_known_rule_ids_accepted_in_disabled_rules(): + policy = Policy( + disabled_rules=frozenset({"missing_error_message", "missing_language"}) + ) + assert policy.disabled_rules == frozenset( + {"missing_error_message", "missing_language"} + ) + + +# --- C5: one severity vocabulary, shared by registration and scoring -------- + + +def test_registry_severity_vocabulary_cross_pins_to_policy(): + # The registration check (registry.KNOWN_SEVERITIES) and the scoring choke + # point (policy.SEVERITIES) must never drift apart. + assert KNOWN_SEVERITIES == SEVERITIES + assert SEVERITIES == ("low", "medium", "high") + + +def test_require_known_severity_rejects_unknown(): + with pytest.raises(ValueError, match="Unknown severity: 'typo'"): + require_known_severity("typo") + + +def test_rule_emitting_unknown_severity_fails_loudly_at_analyze(registry_isolation): + # Declared severity is valid; the EMITTED finding's severity is not. The + # analyzer choke point must reject it — the tie registration could not give us. + bad = _test_rule( + rule_id="bad_severity_rule", + domain="debug", + severity="high", + gap="error description", + finding=RuleFinding(code="bad_finding", message="m", severity="critical"), + ) + register_rule(bad) + with pytest.raises(ValueError, match="Unknown severity: 'critical'"): + InputGuard().analyze("fix the bug in my app") + + +# --- disabled_rules filtering ------------------------------------------------ + + +def test_disabled_rule_is_skipped_during_analyze(): + guard = InputGuard( + policy=Policy(disabled_rules=frozenset({"missing_api_structure"})) + ) + result = guard.analyze(SINGLE_HIGH_INPUT) + assert result.findings == [] # the only finding this input produces + assert result.clarity_score == 100 + assert result.status == "ready" + + +def test_disabled_rules_default_is_no_filtering(): + result = InputGuard().analyze(SINGLE_HIGH_INPUT) + assert [f.code for f in result.findings] == ["missing_api_structure"] + + +def test_disabled_rules_apply_per_call_too(): + guard = InputGuard() + result = guard.analyze( + SINGLE_HIGH_INPUT, + policy=Policy(disabled_rules=frozenset({"missing_api_structure"})), + ) + assert result.findings == [] + + +# --- allow_patterns filtering ------------------------------------------------- + + +def test_allow_pattern_skips_flagging(): + guard = InputGuard(policy=Policy(allow_patterns=(r"\bfixture\b",))) + result = guard.analyze("I love this fixture in the test suite, what does it do") + assert result.status == "ready" + assert result.clarity_score == 100 + assert result.findings == [] + assert result.gaps == [] + + +def test_allow_pattern_no_match_leaves_analysis_unchanged(): + matched = InputGuard(policy=Policy(allow_patterns=("fix",))).analyze( + "fix the bug in my app" + ) + unmatched = InputGuard(policy=Policy(allow_patterns=("deploy",))).analyze( + "fix the bug in my app" + ) + assert matched.status == "ready" + assert matched.findings == [] + assert unmatched.findings # still analyzed + assert unmatched.status != "ready" + + +# --- min_words: short input is valid-but-short, never "vague" ----------------- + + +def test_min_words_below_floor_never_vague(): + result = InputGuard().analyze(VAGUE_INPUT) + assert [f.code for f in result.findings] == ["insufficient_context"] + assert result.clarity_score == 75 + + raised = InputGuard(policy=Policy(min_words=10)).analyze(VAGUE_INPUT) + assert raised.findings == [] + assert raised.clarity_score == 100 + assert raised.status == "ready" + + +def test_min_words_default_matches_explicit_default_policy(): + default_guard = InputGuard() + pinned = InputGuard(policy=Policy()) + assert default_guard.analyze(VAGUE_INPUT).to_dict() == pinned.analyze( + VAGUE_INPUT + ).to_dict() + + +def test_min_words_above_floor_still_flags_long_vague_input(): + result = InputGuard(policy=Policy(min_words=4)).analyze("do something now please") + assert [f.code for f in result.findings] == ["insufficient_context"] + + +# --- plumbing: guard attribute, per-call override, type errors ---------------- + + +def test_guard_exposes_its_policy(): + policy = Policy(penalty_high=30) + assert InputGuard(policy=policy).policy is policy + assert InputGuard().policy == Policy() + + +def test_per_call_policy_overrides_guard_policy(): + guard = InputGuard(policy=Policy(min_words=10)) + assert guard.analyze(VAGUE_INPUT).findings == [] # guard policy applies + allowed = guard.analyze(VAGUE_INPUT, policy=Policy()) # per-call override + assert [f.code for f in allowed.findings] == ["insufficient_context"] + + +def test_constructor_rejects_non_policy(): + with pytest.raises(TypeError, match="Policy instance"): + InputGuard(policy=Policy) # class, not instance + with pytest.raises(TypeError, match="Policy instance"): + InputGuard(policy={"ready_at": 85}) # type: ignore[dict-item] + + +def test_analyze_rejects_non_policy(): + guard = InputGuard() + with pytest.raises(TypeError, match="Policy instance"): + guard.analyze(SINGLE_HIGH_INPUT, policy=85) # type: ignore[arg-type] + + +# --- thread safety with a shared custom policy -------------------------------- + + +def test_shared_policy_parallel_analyze_consistent(): + policy = Policy(disabled_rules=frozenset({"missing_debug_code_context"})) + guard = InputGuard(policy=policy) + text = "fix the bug in my app" + expected = guard.analyze(text).to_dict() + with ThreadPoolExecutor(max_workers=8) as pool: + results = list(pool.map(lambda _: guard.analyze(text).to_dict(), range(64))) + assert results == [expected] * 64 diff --git a/tests/test_policy_defaults.py b/tests/test_policy_defaults.py new file mode 100644 index 0000000..8aca8c7 --- /dev/null +++ b/tests/test_policy_defaults.py @@ -0,0 +1,120 @@ +"""Byte-exact pinning of every Policy default against the v0.2 constants. + +The v0.2 constants are load-bearing for existing users (spec art_bTvdPdJS, +"Calibration drift" risk): these tests fail if any default drifts. Sources: +SEVERITY_PENALTIES and the four threshold constants in the v0.2 scorer, and +the three-word structural floor of the insufficient_context rule. +""" + +from __future__ import annotations + +import dataclasses +import inspect +import re + +import pytest + +from inputguard import InputGuard, Policy + +V02_THRESHOLD_CONSTANTS = { + "ready_at": 85, # warning and strict ready floor + "usable_at": 60, # warning-mode band floor + "strict_clarify_at": 65, # strict-mode needs_clarification floor + "penalty_low": 5, + "penalty_medium": 15, + "penalty_high": 25, + "min_words": 3, # insufficient_context structural floor + "max_chars": 10_000, # v0.3 default cap (v0.2 had none) + "borderline_at": 74, # near-miss band just below ready +} + + +def test_every_policy_default_is_pinned(): + fields = {f.name: f.default for f in dataclasses.fields(Policy)} + for name, value in V02_THRESHOLD_CONSTANTS.items(): + assert fields[name] == value, f"Policy.{name} drifted from v0.2 ({value})" + + +def test_no_unpinned_numeric_fields(): + # Every int field must appear in the pin table so future fields cannot + # ship uncalibrated. + int_fields = { + f.name + for f in dataclasses.fields(Policy) + if f.type in ("int",) or getattr(f.type, "__name__", "") == "int" + } + assert int_fields == set(V02_THRESHOLD_CONSTANTS) + + +def test_dataclass_defaults_are_the_source_of_truth(): + # Constructing with zero arguments must equal the pinned table — no + # hidden overrides inside __post_init__. + policy = Policy() + for name, value in V02_THRESHOLD_CONSTANTS.items(): + assert getattr(policy, name) == value + + +def test_defaults_are_frozen(): + policy = Policy() + with pytest.raises(dataclasses.FrozenInstanceError): + policy.ready_at = 90 # type: ignore[misc] + + +def test_source_defaults_match_the_pinned_table(): + # Guard against the dataclass defaults themselves drifting (the pinning + # tests above read instantiated values; this reads the source). Underscore + # digit separators are normalized before comparison. + source = inspect.getsource(Policy) + for name, value in V02_THRESHOLD_CONSTANTS.items(): + match = re.search(rf"^\s*{name}: int = ([0-9_]+),?\s*$", source, re.MULTILINE) + assert match is not None, f"source default for {name} missing or reformatted" + assert int(match.group(1).replace("_", "")) == value + + +# --- behavioral pins: v0.2 reference outputs under default policy -------------- + + +def _result(text, **kwargs): + return InputGuard(**kwargs).analyze(text).to_dict() + + +def test_reference_outputs_unchanged_warning_mode(): + # Scores observed via runtime probes on the release branch. + assert _result("do something now")["clarity_score"] == 75 + assert _result("do something now")["status"] == "usable_with_warnings" + assert _result("fix the bug in my app")["clarity_score"] == 35 + assert _result("fix the bug in my app")["status"] == "needs_clarification" + assert _result("build a REST API")["clarity_score"] == 50 + + +def test_reference_outputs_unchanged_strict_mode(): + # 50 < 65 strict floor -> blocked; 75 in [65, 85) -> needs_clarification. + assert _result("build a REST API", mode="strict")["status"] == "blocked" + assert _result("do something now", mode="strict")["status"] == "needs_clarification" + + +def test_default_result_keys_are_a_superset_of_v02(): + result = _result("build a REST API") + assert set(result).issuperset( + { + "status", + "clarity_score", + "detected_intent", + "gaps", + "recommendations", + "findings", + "interpretation_note", + } + ) + + +def test_v02_recommendation_keys_unchanged(): + recs = _result("build a REST API")["recommendations"] + assert recs, "expected at least one recommendation" + for rec in recs: + assert set(rec) == { + "gap", + "what_is_missing", + "what_to_provide", + "why_it_matters", + } diff --git a/tests/test_score_breakdown.py b/tests/test_score_breakdown.py new file mode 100644 index 0000000..89096ff --- /dev/null +++ b/tests/test_score_breakdown.py @@ -0,0 +1,147 @@ +"""Score-breakdown tests: the additive audit trail on AnalysisResult. + +Covers the breakdown slice of the v0.3 spec (art_bTvdPdJS §2/§5): base, +per-distinct-gap penalties (negative points), final — owned by the +policy-calibration PR. All expected values below were observed via runtime +probes against the release branch. +""" + +from __future__ import annotations + +import json + +import pytest + +from inputguard import InputGuard, Policy, RuleFinding +from inputguard.scorer import calculate_score, calculate_score_with_breakdown + + +def test_breakdown_shape_end_to_end(): + result = InputGuard().analyze("fix the bug in my app") + assert result.clarity_score == 35 + assert result.score_breakdown == { + "base": 100, + "penalties": [ + {"code": "missing_error_message", "severity": "high", "points": -25}, + { + "code": "missing_expected_vs_actual", + "severity": "high", + "points": -25, + }, + { + "code": "missing_debug_code_context", + "severity": "medium", + "points": -15, + }, + ], + "final": 35, + } + + +def test_breakdown_final_matches_clarity_score(): + for text in ("build a REST API", "do something now", "fix the bug in my app"): + result = InputGuard().analyze(text) + assert result.score_breakdown is not None + assert result.score_breakdown["final"] == result.clarity_score + + +def test_breakdown_dedupes_shared_gap(): + # missing_language and intent_without_language share gap + # "programming language" — one penalty entry for that gap, first wins. + result = InputGuard().analyze("I need to build an app") + assert result.clarity_score == 60 + assert result.score_breakdown == { + "base": 100, + "penalties": [ + {"code": "missing_language", "severity": "high", "points": -25}, + {"code": "missing_output_format", "severity": "medium", "points": -15}, + ], + "final": 60, + } + + +def test_breakdown_penalty_order_matches_findings_order(): + result = InputGuard().analyze("fix the bug in my app") + assert result.score_breakdown is not None + # Every finding here has a distinct gap, so penalty order == findings order. + assert [p["code"] for p in result.score_breakdown["penalties"]] == [ + f.code for f in result.findings + ] + + +def test_breakdown_penalties_follow_policy_penalties(): + guard = InputGuard(policy=Policy(penalty_low=1, penalty_medium=3, penalty_high=10)) + result = guard.analyze("build a REST API") # two high findings + assert result.score_breakdown == { + "base": 100, + "penalties": [ + {"code": "missing_language", "severity": "high", "points": -10}, + {"code": "missing_api_structure", "severity": "high", "points": -10}, + ], + "final": 80, + } + assert result.clarity_score == 80 + + +def test_breakdown_clean_input_has_no_penalties(): + result = InputGuard().analyze( + "Build a web app using React and FastAPI. " + "It needs user authentication with JWT. " + "Store tasks in a PostgreSQL database with title, description, and due date fields. " + "Expose REST API endpoints: GET /tasks, POST /tasks, DELETE /tasks/{id}." + ) + assert result.score_breakdown == {"base": 100, "penalties": [], "final": 100} + + +def test_breakdown_present_in_to_dict_and_json_serializable(): + d = InputGuard().analyze("build a REST API").to_dict() + assert d["score_breakdown"]["final"] == 50 + assert json.dumps(d) # must not raise + + +def test_to_dict_breakdown_is_a_copy(): + # The dict to_dict() returns must be detached from the result's own + # breakdown object: mutating the result afterwards cannot rewrite history. + result = InputGuard().analyze("build a REST API") + snapshot = result.to_dict() + assert result.score_breakdown is not None + result.score_breakdown["penalties"][0]["points"] = 0 # type: ignore[index] + assert snapshot["score_breakdown"]["penalties"][0]["points"] == -25 + assert result.to_dict()["score_breakdown"]["penalties"][0]["points"] == 0 + + +def test_breakdown_clamps_at_zero(): + findings = [ + RuleFinding(code=f"c{i}", message="m", severity="high") for i in range(5) + ] + score, breakdown = calculate_score_with_breakdown(findings) + assert score == 0 + assert breakdown["final"] == 0 + assert len(breakdown["penalties"]) == 5 + assert sum(p["points"] for p in breakdown["penalties"]) == -125 + + +def test_breakdown_gap_none_dedupes_by_code(): + findings = [ + RuleFinding(code="a", message="m", severity="low"), + RuleFinding(code="a", message="m", severity="high"), + ] + score, breakdown = calculate_score_with_breakdown(findings) + assert score == 75 + assert breakdown["penalties"] == [ + {"code": "a", "severity": "high", "points": -25} + ] + + +def test_calculate_score_keeps_v02_signature_and_matches_breakdown(): + findings = [RuleFinding(code="a", message="m", severity="medium", gap="g")] + assert calculate_score(findings) == 85 + score, breakdown = calculate_score_with_breakdown(findings) + assert score == 85 + assert breakdown["final"] == 85 + + +def test_breakdown_unknown_severity_raises_before_scoring(): + findings = [RuleFinding(code="a", message="m", severity="typo")] + with pytest.raises(ValueError, match="Unknown severity: 'typo'"): + calculate_score_with_breakdown(findings) From 1d5c8c90664ba0a2b8987c69bfe29eaa51280c8f Mon Sep 17 00:00:00 2001 From: "obvious-autobuild[bot]" <262744130+obvious-autobuild[bot]@users.noreply.github.com> Date: Thu, 17 Sep 2026 16:40:44 -0400 Subject: [PATCH 09/18] =?UTF-8?q?feat:=20first-party=20writing=20domain=20?= =?UTF-8?q?=E2=80=94=20six=20rules,=20compose=20intent,=20complete=20advic?= =?UTF-8?q?e=20(#10)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * feat(writing-domain): first-party writing domain — six rules, compose intent, complete advice Register the writing domain through the same registry path user domains take: one globally unique fallback intent ("compose") and six rules covering the eval vocabulary's six gaps — audience, purpose, structure/format, source material, context, completeness — at the vocabulary-pinned severities (high/high/medium/high/medium/low). All term matching goes through the shared word-boundary matcher (inputguard.matching, PR #6). Rules follow the trigger-and-satisfy shape; the source-material rule fires only when existing text is referenced but not provided. Recommendations and follow-up entries ship for all six gaps in the same change, keeping the registry completeness invariant whole. Co-authored-by: Kalisetti Nihanth Naidu * test(writing-domain): per-rule, completeness, integration, no-leakage, and eval-label coverage Five layers: per-rule fire/silent behavior; the writing-domain completeness invariant (every writing gap has a four-key recommendation and one or two follow-up questions); analyze() integration end to end (underspecified -> needs_clarification at score 30, fully specified -> ready, partially specified -> usable_with_warnings); no-leakage in both directions at registry and analyze level; and the 14 labeled English writing rows from eval/cases.csv pinned as a calibration net. The exact-set registry expectations extend from 19 coding rules to 25 rules across two domains — the guards are strengthened, not weakened. Co-authored-by: Kalisetti Nihanth Naidu --------- Co-authored-by: Obvious Co-authored-by: Kalisetti Nihanth Naidu --- inputguard/followups.py | 21 ++ inputguard/recommender.py | 36 +++ inputguard/rules/__init__.py | 22 +- inputguard/rules/writing.py | 431 +++++++++++++++++++++++++++++++++++ tests/test_registry.py | 17 +- tests/test_writing_domain.py | 303 ++++++++++++++++++++++++ 6 files changed, 822 insertions(+), 8 deletions(-) create mode 100644 inputguard/rules/writing.py create mode 100644 tests/test_writing_domain.py diff --git a/inputguard/followups.py b/inputguard/followups.py index ff8e590..a1fba74 100644 --- a/inputguard/followups.py +++ b/inputguard/followups.py @@ -94,6 +94,27 @@ "completion criteria": ( "What does 'done' look like — what should you be able to do when the feature works?", ), + "audience": ( + "Who is going to read this — a manager, your team, clients, or the public?", + "How familiar will readers be with the topic?", + ), + "purpose": ( + "What should this piece accomplish — inform, persuade, announce, or request something?", + ), + "structure/format": ( + "How long should it be, and how should it be organized — bullets, sections, or flowing prose?", + ), + "source material": ( + "Where is the text to work from — can you paste it or attach the file?", + "Should I preserve its structure, or can I reorganize it freely?", + ), + "context": ( + "What is the topic or situation this piece should cover?", + ), + "completeness": ( + "Is there anything it must include — specific points, names, or numbers?", + "Any hard limits, like a word count or a required section?", + ), } # Slot names a template may reference. A template naming anything else is a diff --git a/inputguard/recommender.py b/inputguard/recommender.py index 56c42bf..d17dc50 100644 --- a/inputguard/recommender.py +++ b/inputguard/recommender.py @@ -112,6 +112,42 @@ "what_to_provide": "Describe what done looks like. For example: 'the feature is complete when a user can type in the search box and see matching results appear within 200ms' or 'done means the user receives an email notification within 30 seconds of placing an order'. This prevents the AI from stopping too early or going too far.", "why_it_matters": "Without a clear definition of done, the AI may ship a half-finished feature or over-engineer something well past what you needed.", }, + "audience": { + "gap": "audience", + "what_is_missing": "You haven't said who will read this.", + "what_to_provide": "Name the reader and what they already know. For example: 'for the executive team, keep it high-level', 'for beginners who have never used the tool', 'for the client stakeholders, no jargon'. Even one phrase like 'for my manager' sharpens the tone and level of detail.", + "why_it_matters": "The same topic reads completely differently for a CEO, a new hire, or a customer. Without a named reader, the AI aims the piece at nobody in particular.", + }, + "purpose": { + "gap": "purpose", + "what_is_missing": "You haven't said what this piece should accomplish.", + "what_to_provide": "State the goal in one phrase: 'to persuade the steering committee to fund Q1 headcount', 'to announce the launch', 'to explain why the deadline moved'. A goal like 'convince', 'inform', or 'ask for' is enough to aim the writing.", + "why_it_matters": "Informing and persuading lead to different structures, evidence, and tone. Without a goal, the AI produces a generic piece that does neither well.", + }, + "structure/format": { + "gap": "structure/format", + "what_is_missing": "You haven't said how long the piece should be or how it should be organized.", + "what_to_provide": "Give a length and a shape. For example: 'under 300 words', 'one page', 'bullets with a short intro', 'three sections: situation, options, recommendation'. Any constraint — even just 'keep it short' — works.", + "why_it_matters": "Without length or organization guidance, the AI picks its own shape. You will often get a bloated draft you then have to cut down yourself.", + }, + "source material": { + "gap": "source material", + "what_is_missing": "You asked to work on existing text but didn't provide it.", + "what_to_provide": "Paste the text you want reworked — the draft, notes, or paragraph — directly into the message, or point to where it lives ('the outline is at the bottom'). Include any version details that matter.", + "why_it_matters": "The AI cannot see text that isn't in the message. Without it, you get a generic rewrite of an imaginary document instead of an improvement to yours.", + }, + "context": { + "gap": "context", + "what_is_missing": "You haven't said what the piece is about or what situation it responds to.", + "what_to_provide": "Name the subject and the situation. For example: 'about remote work for our company blog', 'regarding the Q3 roadmap', 'the email should tell the team the migration finished'. One sentence of background is enough.", + "why_it_matters": "Without a topic or situation, there is nothing to write about. The AI either invents one or asks you everything you could have said up front.", + }, + "completeness": { + "gap": "completeness", + "what_is_missing": "You haven't listed anything the piece must include.", + "what_to_provide": "Name the must-haves: 'include the headline, a quote from the CEO, and pricing', 'must cover current costs, risks, and the timeline', 'mention the new ship date'. Also note any hard limits, like a word count.", + "why_it_matters": "Must-have details left out of the request get left out of the draft. Naming them up front saves a second pass to work them in.", + }, } diff --git a/inputguard/rules/__init__.py b/inputguard/rules/__init__.py index fc32218..203c47f 100644 --- a/inputguard/rules/__init__.py +++ b/inputguard/rules/__init__.py @@ -1,9 +1,11 @@ -"""Built-in rule modules and the registry wiring for the coding domain. +"""Built-in rule modules and the registry wiring for the coding and +writing domains. -Importing this package registers the coding domain — its intent signals and -all 19 built-in rules — through the exact same registry path a user rule -takes. The v0.2 ``run_*_rules`` functions stay exported for backward -compatibility, but the analyzer dispatches through the registry now. +Importing this package registers both first-party domains — their intent +signals and all built-in rules — through the exact same registry path a +user rule takes. The v0.2 ``run_*_rules`` functions stay exported for +backward compatibility, but the analyzer dispatches through the registry +now. """ from inputguard.detector import ( @@ -22,6 +24,7 @@ OPTIMIZATION_RULES, run_optimization_rules, ) +from inputguard.rules.writing import WRITING_RULES, WRITING_SIGNALS __all__ = [ "run_coding_rules", @@ -45,3 +48,12 @@ *FEATURE_RULES, ), ) + +# The writing domain: a single fallback intent ("compose" — globally unique; +# no coding intent name is reused, so no cross-domain rule leakage) and the +# six first-party writing rules, registered through the same path. +REGISTRY.register_domain( + "writing", + WRITING_SIGNALS, + rules=WRITING_RULES, +) diff --git a/inputguard/rules/writing.py b/inputguard/rules/writing.py new file mode 100644 index 0000000..a0b8007 --- /dev/null +++ b/inputguard/rules/writing.py @@ -0,0 +1,431 @@ +"""First-party writing domain rules. + +The writing domain covers essays, emails, documents, and social posts. +Its gap vocabulary (pinned by the eval set's ``gap_vocabulary`` sheet) is +six gaps, one rule each: + +============ ========== ========================== +Gap Severity Rule +============ ========== ========================== +audience high missing_audience +purpose high missing_purpose +structure/ medium missing_structure_format +format +source high missing_source_material +material +context medium missing_writing_context +completeness low missing_completeness +============ ========== ========================== + +Rules follow the established trigger-and-satisfy shape: each fires only on +writing tasks and only when the gap's satisfier is absent from the prompt. + +Design note — one intent, not many. The six gaps apply to any writing task +(drafting, rewriting, feedback), so the domain declares a single fallback +intent (``compose``) instead of fragmenting into sub-intents whose inputs +would silently skip the rules. The compose-vs-revise discrimination the +coding domains put in intent signals lives here inside the rules that need +it (``missing_source_material`` fires only when the task references +existing material). + +All term matching goes through the shared word-boundary matcher +(``inputguard.matching``): ``"write"`` matches ``"writing"`` but not +``"written in Go"``, and ``"board"`` matches ``"the board"`` but not +``"keyboard"``. Phrase terms match adjacent words; the non-adjacent token +fallback is deliberately left off — these are new rules, so there is no +v0.2 verdict to preserve. +""" + +from __future__ import annotations + +import re +from typing import List, Optional, Tuple + +from inputguard.matching import contains_any +from inputguard.registry import register_rule +from inputguard.types import RuleFinding + + +# A writing task is present when the prompt names a writing artifact or a +# writing action. Every rule gates on this before looking at its own gap — +# it keeps coding prompts analyzed under domain="writing" (a user error) +# from producing misleading findings. +WRITING_TASK_TERMS = { + "write", "draft", "compose", + "blog", "post", "email", "memo", "letter", "essay", "resume", + "proposal", "paragraph", "article", "recap", "summary", "summaries", + "announcement", "press release", "cover letter", "outline", + "linkedin", "newsletter", "bio", "story", "tweet", "thread", + "note", "notes", "report", "document", "review", + "rewrite", "revise", "reword", "rephrase", "proofread", "polish", + "condense", "shorten", "tighten", "edit", +} + +# The named readership that satisfies the audience gap. Reader nouns are +# matched anywhere in the text ("admissions officers are the readers" has +# no "for" prefix, and it must satisfy). +AUDIENCE_TERMS = { + "audience", "readers", "reader", "readership", + "beginners", "beginner", "newcomers", + "manager", "managers", "my manager", "your manager", "my boss", + "supervisor", "supervisors", + "executives", "executive team", "leadership", "leadership team", + "board", "board members", "steering committee", "committee", + "stakeholders", "stakeholder", + "clients", "client", "customers", "customer", + "team", "my team", "engineering team", "engineers", "engineer", + "developers", "developer", + "journalists", "journalist", "press", "media", + "recruiters", "recruiter", "hiring manager", "hiring managers", + "admissions officers", "admissions", "officers", + "students", "student", "teachers", "teacher", "professors", + "professor", "instructors", "instructor", + "investors", "investor", "shareholders", "shareholder", + "subscribers", "subscriber", "followers", "follower", + "colleagues", "colleague", "coworkers", "teammates", "peers", + "donors", "donor", "members", "community", + "applicants", "applicant", "candidates", "candidate", + "reviewers", "reviewer", "editor", "editors", + "users", "user base", "general public", +} + +# A stated goal that satisfies the purpose gap ("announcing the Q3 +# roadmap", "asking for budget", "so that nobody reruns the old pipeline"). +PURPOSE_TERMS = { + "to convince", "convince", "convincing", "persuade", "persuading", + "persuasive", + "announce", "announcing", "announced", "announcement", + "request", "requesting", "requested", + "ask for", "asking for", "asks for", + "goal is", "goal of", "aim is", "aimed at", "purpose is", + "objective is", + "so that", "in order to", + "recap", + "lead with", "leads with", + "explain", "explaining", "explains", + "apologize", "apologizing", "apology", + "thank", "thanking", "thanks", + "celebrate", "celebrating", "celebration", + "invite", "inviting", "invitation", + "pitch", "pitching", "propose", "proposing", + "inform", "informing", "update", "updating", + "introduce", "introducing", + "complain", "complaining", "complaint", + "seeking", "seek", "justify", "justifying", + "follow up", "follow-up", "following up", + "respond", "responding", "reply", + "remind", "reminding", "reminder", + "promote", "promoting", "marketing", +} + +# Length or organization guidance that satisfies the structure/format gap. +# Bare artifact types ("blog post") do not satisfy — the vocabulary is +# explicit about that — so "page" and "word" only match in count forms. +STRUCTURE_TERMS = { + "bullet", "bullets", "bullet point", "bullet points", + "sections", "section", "headings", "heading", "subheadings", + "subheading", "numbered", "numbered list", + "concise", "short", "shorter", "brief", "condensed", + "length", "word count", "word limit", "page limit", + "table format", "as a table", "in a table", + "paragraph", "paragraphs", +} +# Count forms: "300 words", "one page", "two pages", "under 150 words", +# "3 bullets", "2 sections" — including hyphenated modifiers ("300-word"). +_STRUCTURE_COUNT_PATTERNS = ( + r"\b\d+\s*[-\s]?\s*words?\b", + r"\b(?:one|two|three|four|five|six|ten|single|half|\d+)\s*[-\s]?\s*pages?\b", + r"\b\d+\s*[-\s]?\s*(?:paragraphs?|bullets?|sections?|slides?)\b", +) + +# The task references existing material it wants worked on. Fires the +# source-material rule; the gap is then satisfied only if the material is +# actually provided. +MATERIAL_REFERENCE_TERMS = { + "rewrite", "revise", "reword", "rephrase", "proofread", "polish", + "improve", "condense", "shorten", "tighten", "trim", "expand", + "fix up", "clean up", "restructure", "reorganize", + "summarize", "summarise", "edit", "edited", "editing", "edits", + "turn the", "turn my", "turn this", + "my resume", "my essay", "my draft", "my outline", "my paragraph", + "my summary", "my cover letter", "my letter", "my email", "my post", + "my blog", "my notes", "my bio", "my article", "my proposal", + "my story", "my memo", "my document", "my report", + "this paragraph", "this essay", "this draft", "this outline", + "this email", "this document", "this summary", "this post", + "this blog", "this letter", "this resume", "this memo", "this note", + "this report", "this article", + "the essay", "the draft", "the outline", "the paragraph", + "the notes", "the meeting notes", "the resume", "the cover letter", + "the summary", "the document", "the manuscript", + "the survey responses", "the bullet outline", +} +# The material is actually provided: pasted into the prompt, attached, or +# pointed at concretely. +MATERIAL_PROVIDED_TERMS = { + "pasted", "attached", "below", "above", "at the bottom", "at the end", + "copied", "the following", "as follows", "included", "quoted", + "in the attachment", "in the document", "in the file", "the transcript", +} + +# A named subject or situation that satisfies the context gap: explicit +# context markers, topic introducers, the artifact itself, or a concrete +# situation noun ("the migration", "the Q3 roadmap", "the vendor"). +CONTEXT_TERMS = { + "context", "background", "regarding", "about", "on the topic of", + "subject:", "topic:", "re:", + "blog", "post", "email", "memo", "letter", "essay", "resume", + "proposal", "paragraph", "article", "recap", "summary", + "announcement", "press release", "cover letter", "outline", + "linkedin", "newsletter", "bio", "story", "tweet", + "note", "notes", "report", "document", "draft", "review", + "migration", "launch", "release", "deadline", "roadmap", "project", + "initiative", "campaign", "rollout", "transition", "incident", + "outage", "meeting", "event", "acquisition", "merger", "reorg", + "vendor", "offsite", "layoff", "hiring", "product", "issue", + "delay", "expansion", "funding", "series a", "restructure", +} + +# Enumerated required content or constraints that satisfy the +# completeness gap ("must cover current costs", "mention the new ship +# date", "include the headline", "max 300 words"). Bare "cover" is +# deliberately absent — "cover letter" must not satisfy. +COMPLETENESS_TERMS = { + "include", "including", "included", "includes", + "must cover", "should cover", "needs to cover", "covering", + "mention", "mentions", "mentioning", "mentioned", + "highlight", "highlighting", "highlights", + "emphasize", "emphasizing", "emphasise", "emphasising", + "such as", "specifically", "namely", "in particular", + "at least", "at minimum", + "make sure", "be sure to", "don't forget", "do not forget", + "required", "require", "requires", "requirements", + "must have", "must-have", "needs to have", "needs to include", + "max", "maximum", "at most", "no more than", "stay under", + "within", "limit", "up to", +} + + +def _normalize(text: str) -> str: + """Defensive re-normalization for direct callers of the check functions.""" + return re.sub(r"\s+", " ", text.strip().lower()) + + +def check_missing_audience(text: str) -> Optional[RuleFinding]: + text = _normalize(text) + if not contains_any(text, WRITING_TASK_TERMS): + return None + if contains_any(text, AUDIENCE_TERMS): + return None + return RuleFinding( + code="missing_audience", + message="Writing task detected but no audience or reader is specified.", + severity="high", + gap="audience", + ) + + +def check_missing_purpose(text: str) -> Optional[RuleFinding]: + text = _normalize(text) + if not contains_any(text, WRITING_TASK_TERMS): + return None + if contains_any(text, PURPOSE_TERMS): + return None + return RuleFinding( + code="missing_purpose", + message="Writing task detected but the goal — what the piece should accomplish — is not stated.", + severity="high", + gap="purpose", + ) + + +def _has_structure_signal(text: str) -> bool: + if contains_any(text, STRUCTURE_TERMS): + return True + return any(re.search(pattern, text) for pattern in _STRUCTURE_COUNT_PATTERNS) + + +def check_missing_structure_format(text: str) -> Optional[RuleFinding]: + text = _normalize(text) + if not contains_any(text, WRITING_TASK_TERMS): + return None + if _has_structure_signal(text): + return None + return RuleFinding( + code="missing_structure_format", + message="Writing task detected but no length or organization guidance is provided.", + severity="medium", + gap="structure/format", + ) + + +def check_missing_source_material(text: str) -> Optional[RuleFinding]: + text = _normalize(text) + if not contains_any(text, WRITING_TASK_TERMS): + return None + if not contains_any(text, MATERIAL_REFERENCE_TERMS): + return None + if contains_any(text, MATERIAL_PROVIDED_TERMS): + return None + return RuleFinding( + code="missing_source_material", + message="Existing material is referenced but not provided — paste or attach the text to work from.", + severity="high", + gap="source material", + ) + + +def check_missing_writing_context(text: str) -> Optional[RuleFinding]: + text = _normalize(text) + if not contains_any(text, WRITING_TASK_TERMS): + return None + if contains_any(text, CONTEXT_TERMS): + return None + return RuleFinding( + code="missing_writing_context", + message="Writing task detected but the subject or situation is not named.", + severity="medium", + gap="context", + ) + + +def check_missing_completeness(text: str) -> Optional[RuleFinding]: + text = _normalize(text) + if not contains_any(text, WRITING_TASK_TERMS): + return None + if contains_any(text, COMPLETENESS_TERMS): + return None + return RuleFinding( + code="missing_completeness", + message="Writing task detected but no required content or constraints are specified.", + severity="low", + gap="completeness", + ) + + +# Registry adapters: the check functions above stay the single home of the +# rule logic; these classes expose it through the v0.3 Rule protocol and +# register it through the same path a user rule takes. + + +@register_rule +class MissingAudienceRule: + """Registry adapter for check_missing_audience.""" + + id = "missing_audience" + domain = "compose" + severity = "high" + gap = "audience" + + def check(self, text: str) -> Optional[RuleFinding]: + return check_missing_audience(text) + + +@register_rule +class MissingPurposeRule: + """Registry adapter for check_missing_purpose.""" + + id = "missing_purpose" + domain = "compose" + severity = "high" + gap = "purpose" + + def check(self, text: str) -> Optional[RuleFinding]: + return check_missing_purpose(text) + + +@register_rule +class MissingStructureFormatRule: + """Registry adapter for check_missing_structure_format.""" + + id = "missing_structure_format" + domain = "compose" + severity = "medium" + gap = "structure/format" + + def check(self, text: str) -> Optional[RuleFinding]: + return check_missing_structure_format(text) + + +@register_rule +class MissingSourceMaterialRule: + """Registry adapter for check_missing_source_material.""" + + id = "missing_source_material" + domain = "compose" + severity = "high" + gap = "source material" + + def check(self, text: str) -> Optional[RuleFinding]: + return check_missing_source_material(text) + + +@register_rule +class MissingWritingContextRule: + """Registry adapter for check_missing_writing_context.""" + + id = "missing_writing_context" + domain = "compose" + severity = "medium" + gap = "context" + + def check(self, text: str) -> Optional[RuleFinding]: + return check_missing_writing_context(text) + + +@register_rule +class MissingCompletenessRule: + """Registry adapter for check_missing_completeness.""" + + id = "missing_completeness" + domain = "compose" + severity = "low" + gap = "completeness" + + def check(self, text: str) -> Optional[RuleFinding]: + return check_missing_completeness(text) + + +WRITING_RULES: Tuple[type, ...] = ( + MissingAudienceRule, + MissingPurposeRule, + MissingStructureFormatRule, + MissingSourceMaterialRule, + MissingWritingContextRule, + MissingCompletenessRule, +) + +# The single fallback intent: every writing-domain input is a composition +# task. Exactly one empty-terms intent is what the registry requires. +WRITING_SIGNALS: Tuple[Tuple[str, tuple], ...] = ( + ("compose", ()), +) + +# Findings this domain can emit, for the completeness invariant. +WRITING_GAPS: Tuple[str, ...] = ( + "audience", + "purpose", + "structure/format", + "source material", + "context", + "completeness", +) + + +def run_writing_rules(text: str) -> List[RuleFinding]: + """Run every writing rule directly (mirror of the coding run_* helpers).""" + findings: List[RuleFinding] = [] + seen = set() + for check in ( + check_missing_audience, + check_missing_purpose, + check_missing_structure_format, + check_missing_source_material, + check_missing_writing_context, + check_missing_completeness, + ): + result = check(text) + if result and result.code not in seen: + findings.append(result) + seen.add(result.code) + return findings diff --git a/tests/test_registry.py b/tests/test_registry.py index e60b294..4b74cad 100644 --- a/tests/test_registry.py +++ b/tests/test_registry.py @@ -45,6 +45,13 @@ "missing_existing_stack", "missing_feature_scope", "missing_completion_criteria", + # First-party writing domain (6 rules). + "missing_audience", + "missing_purpose", + "missing_structure_format", + "missing_source_material", + "missing_writing_context", + "missing_completeness", } _V02_INTENT_RUNNERS = { @@ -85,11 +92,13 @@ def registry_isolation(): REGISTRY._origins.update(origins_before) -# --- the 19 built-ins dogfood the registry path --------------------------- +# --- the 25 built-ins dogfood the registry path --------------------------- def test_all_builtin_rules_registered_through_registry(): - assert len(REGISTRY.rule_ids()) == 19 + # 19 coding + 6 writing built-ins. The exact-set guard is the point: + # a rule that stops registering (or an unregistered stray) fails here. + assert len(REGISTRY.rule_ids()) == 25 assert set(REGISTRY.rule_ids()) == EXPECTED_BUILTIN_IDS for rule in REGISTRY.rules(): assert isinstance(rule, Rule) @@ -97,7 +106,9 @@ def test_all_builtin_rules_registered_through_registry(): def test_coding_domain_registered_with_priority_signals(): - assert REGISTRY.domain_names() == ("coding",) + # Both first-party domains register at import: coding first, writing + # second. + assert REGISTRY.domain_names() == ("coding", "writing") chain = REGISTRY.get_domain_signals("coding") assert [intent for intent, _ in chain] == [ "debug", diff --git a/tests/test_writing_domain.py b/tests/test_writing_domain.py new file mode 100644 index 0000000..ca4e0c9 --- /dev/null +++ b/tests/test_writing_domain.py @@ -0,0 +1,303 @@ +"""First-party writing domain tests. + +Five layers, matching the acceptance criteria: + +- per-rule behavior: each of the six writing rules fires on a gap-bearing + writing prompt and stays silent when that gap is satisfied (or when the + prompt is not a writing task at all); +- the writing-domain completeness invariant: every registered writing gap + has a four-key recommendation entry and one or two curated follow-up + questions, and the intent the rules bind to ("compose") is globally + unique — no coding intent reuses it; +- analyze() integration end to end: an underspecified writing prompt comes + back needs_clarification with the right gaps, severities, score, and + follow-ups; a fully specified one comes back ready; +- no leakage in either direction, at the registry level and through the + public analyze() path: writing codes never appear in coding analyses and + coding codes never appear in writing ones; +- the labeled English writing rows in eval/cases.csv match their labels — + the domain's calibration net. (The FP/FN measurement harness stays the + source of truth for rates; this pins the rows the gap_vocabulary sheet + defines.) + +All existing tests are untouched — this file only adds. +""" + +from __future__ import annotations + +import csv +from pathlib import Path + +import pytest + +from inputguard import InputGuard +from inputguard.followups import _FOLLOW_UP_QUESTIONS +from inputguard.recommender import _RECOMMENDATIONS +from inputguard.registry import REGISTRY +from inputguard.rules.writing import ( + WRITING_GAPS, + WRITING_RULES, + check_missing_audience, + check_missing_completeness, + check_missing_purpose, + check_missing_source_material, + check_missing_structure_format, + check_missing_writing_context, + run_writing_rules, +) + +EVAL_CASES = Path(__file__).resolve().parent.parent / "eval" / "cases.csv" + + +# --- per-rule behavior: fires on the gap, silent when satisfied -------------- + + +@pytest.mark.parametrize( + "check, fires_on, silent_on", + [ + ( + check_missing_audience, + "Write a blog post about remote work", + "Write a blog post about remote work for the engineering team", + ), + ( + check_missing_purpose, + "Write a blog post about remote work", + "Write a blog post announcing our new remote-work policy", + ), + ( + check_missing_structure_format, + "Write a blog post about remote work", + "Write a short blog post about remote work", + ), + # Fires only when existing material is referenced but not provided; + # a rewrite-with-text (gap satisfied) and a bare drafting prompt + # (no trigger) are both silent. + ( + check_missing_source_material, + "Rewrite my resume summary", + "Rewrite the resume summary I pasted below", + ), + ( + check_missing_writing_context, + "Write something short for tomorrow", + "Write a blog post about remote work", + ), + ( + check_missing_completeness, + "Write a blog post about remote work", + "Write a blog post about remote work and include the headline stat", + ), + ], +) +def test_rule_fires_on_gap_and_stays_silent_when_satisfied(check, fires_on, silent_on): + finding = check(fires_on) + assert finding is not None, f"{check.__name__} must fire on {fires_on!r}" + + assert check(silent_on) is None, f"{check.__name__} must stay silent on {silent_on!r}" + + +def test_not_a_writing_task_produces_no_findings(): + # Coding prompts analyzed in the writing domain (a user error) must not + # produce misleading writing findings — the writing-task gate blocks + # every rule. + for text in ( + "fix the login crash", + "my endpoint returns 500 instead of 200", + "explain async", + ): + assert run_writing_rules(text) == [], f"expected no findings for {text!r}" + + +def test_every_writing_rule_binds_to_the_compose_intent_with_vocab_severity(): + # The gap_vocabulary sheet pins the severities; the intent must be + # "compose" for every rule (the single fallback intent). + expected = { + "missing_audience": "high", + "missing_purpose": "high", + "missing_structure_format": "medium", + "missing_source_material": "high", + "missing_writing_context": "medium", + "missing_completeness": "low", + } + for rule in WRITING_RULES: + assert rule.domain == "compose" + assert rule.severity == expected[rule.id] + assert rule.gap is not None + + +# --- the writing-domain completeness invariant ------------------------------ + + +def _writing_registry_rules(): + writing_ids = {rule.id for rule in WRITING_RULES} + return [r for r in REGISTRY.rules() if r.id in writing_ids] + + +def test_every_registered_writing_rule_has_a_recommendation_and_follow_ups(): + # Mirrors the global invariant in test_followups.py, scoped to the + # writing domain: a writing rule can ship only with complete advice. + for rule in _writing_registry_rules(): + assert rule.gap in _RECOMMENDATIONS, f"gap {rule.gap!r} has no recommendation entry" + assert rule.gap in _FOLLOW_UP_QUESTIONS, f"gap {rule.gap!r} has no follow-up question" + assert 1 <= len(_FOLLOW_UP_QUESTIONS[rule.gap]) <= 2 + + +def test_writing_gap_set_matches_the_domain_vocabulary(): + # The six gaps the eval set's gap_vocabulary sheet defines for writing + # are exactly what the rules declare — no more, no fewer. + declared = {rule.gap for rule in _writing_registry_rules()} + assert declared == set(WRITING_GAPS) + + +def test_compose_intent_is_globally_unique(): + # PR #8 contract: intent names are globally unique across domains. No + # coding intent may collide with the writing domain's single intent. + for intent in ("build", "debug", "optimization", "explanation", "feature"): + ids = [r.id for r in REGISTRY.rules_for_intent(intent)] + assert not writing_ids_in(ids), f"writing rules leaked into intent {intent!r}" + + +def writing_ids_in(ids): + writing_ids = {rule.id for rule in WRITING_RULES} + return [i for i in ids if i in writing_ids] + + +def test_rules_for_compose_are_exactly_the_writing_rules(): + assert {r.id for r in REGISTRY.rules_for_intent("compose")} == { + rule.id for rule in WRITING_RULES + } + + +# --- analyze() integration, end to end --------------------------------------- + + +def test_underspecified_writing_prompt_end_to_end(): + result = InputGuard().analyze("Write a blog post about remote work", domain="writing") + + assert result.status == "needs_clarification" + assert result.detected_intent == "compose" + assert set(result.gaps) == {"audience", "purpose", "structure/format", "completeness"} + # 100 - 25 (audience, high) - 25 (purpose, high) - 15 (structure, + # medium) - 5 (completeness, low) = 30. + assert result.clarity_score == 30 + severity_by_gap = {f.gap: f.severity for f in result.findings} + assert severity_by_gap == { + "audience": "high", + "purpose": "high", + "structure/format": "medium", + "completeness": "low", + } + # Advice and follow-ups are complete: one four-key recommendation and + # at least one question per gap, in gap order. + assert [r["gap"] for r in result.recommendations] == result.gaps + for rec in result.recommendations: + assert set(rec) == {"gap", "what_is_missing", "what_to_provide", "why_it_matters"} + assert len(result.follow_ups) >= len(result.gaps) + assert all(q.endswith("?") for q in result.follow_ups) + # English input: no degradation, honest full coverage. + assert result.detected_language == "en" + assert result.heuristic_coverage == "full" + assert result.degradation_note is None + + +def test_fully_specified_writing_prompt_is_ready(): + text = ( + "Turn my bullet outline into a full proposal for the steering " + "committee; goal is to win headcount for Q1; include budget, risks, " + "and milestones; the outline is at the bottom" + ) + result = InputGuard().analyze(text, domain="writing") + + assert result.status == "ready" + assert result.clarity_score == 100 + assert result.gaps == [] + assert result.findings == [] + assert result.follow_ups == [] + + +def test_partially_specified_writing_prompt_lands_in_warning_band(): + # Audience, purpose, and context are present; length/organization and + # required content are not. 100 - 15 (structure, medium) - 5 + # (completeness, low) = 80 -> usable_with_warnings in warning mode. + result = InputGuard().analyze( + "Write an email to the engineering team announcing the Q3 roadmap", + domain="writing", + ) + + assert result.status == "usable_with_warnings" + assert result.clarity_score == 80 + assert set(result.gaps) == {"structure/format", "completeness"} + + +# --- no leakage, either direction -------------------------------------------- + + +def test_coding_analyses_never_emit_writing_codes(): + guard = InputGuard() + for text in ( + "Write a blog post about remote work", + "Rewrite my resume summary", + "Write a short email to my team explaining why the deadline moved", + ): + result = guard.analyze(text) # default domain: coding + assert writing_ids_in(f.code for f in result.findings) == [], ( + f"writing codes leaked into a coding analysis of {text!r}" + ) + + +def test_writing_analyses_never_emit_coding_codes(): + coding_codes = {rule.id for rule in REGISTRY.rules_for_intent("build",)} | { + rule.id + for intent in ("build", "debug", "optimization", "explanation", "feature") + for rule in REGISTRY.rules_for_intent(intent) + } + guard = InputGuard() + for text in ( + "fix the login crash", + "make this faster", + "add auth to my app", + "explain async", + ): + result = guard.analyze(text, domain="writing") + leaked = [f.code for f in result.findings if f.code in coding_codes] + assert leaked == [], f"coding codes leaked into a writing analysis of {text!r}" + + +# --- the labeled eval rows (calibration net) --------------------------------- + + +def _writing_eval_rows(): + with open(EVAL_CASES, newline="", encoding="utf-8") as handle: + rows = [row for row in csv.DictReader(handle) if row["domain"] == "writing"] + assert rows, "eval/cases.csv must contain writing rows" + return rows + + +def _parse_multi_value(raw: str): + """Split a ``;``-separated eval field; ``none``/``n/a`` mean no entries.""" + value = (raw or "").strip() + if value.lower() in {"none", "n/a", ""}: + return [] + return [part.strip() for part in value.split(";") if part.strip()] + + +@pytest.mark.parametrize("row", _writing_eval_rows(), ids=lambda row: row["id"]) +def test_labeled_writing_rows_match_their_labels(row): + guard = InputGuard() + result = guard.analyze(row["text"], domain="writing") + + expected_gaps = _parse_multi_value(row["expected_gaps"]) + expected_severities = _parse_multi_value(row["expected_severities"]) + + assert set(result.gaps) == set(expected_gaps), ( + f"{row['id']}: gap set mismatch on {row['text']!r}" + ) + assert result.status == row["expected_status"], ( + f"{row['id']}: status mismatch on {row['text']!r}" + ) + severity_by_gap = {f.gap: f.severity for f in result.findings} + for gap, severity in zip(expected_gaps, expected_severities): + assert severity_by_gap.get(gap) == severity, ( + f"{row['id']}: severity mismatch for gap {gap!r} on {row['text']!r}" + ) From 8cfd558cf4a12066b5757f488e02ef094694140e Mon Sep 17 00:00:00 2001 From: "obvious-autobuild[bot]" <262744130+obvious-autobuild[bot]@users.noreply.github.com> Date: Thu, 17 Sep 2026 20:55:51 +0000 Subject: [PATCH 10/18] =?UTF-8?q?feat:=20first-party=20data-analysis=20dom?= =?UTF-8?q?ain=20=E2=80=94=20six=20rules,=20analysis=20intent,=20complete?= =?UTF-8?q?=20advice=20(#11)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-authored-by: Kalisetti Nihanth Naidu --- inputguard/followups.py | 19 +- inputguard/recommender.py | 38 ++- inputguard/rules/__init__.py | 28 +- inputguard/rules/data_analysis.py | 418 +++++++++++++++++++++++++++++ tests/test_data_analysis_domain.py | 313 +++++++++++++++++++++ tests/test_registry.py | 24 +- 6 files changed, 821 insertions(+), 19 deletions(-) create mode 100644 inputguard/rules/data_analysis.py create mode 100644 tests/test_data_analysis_domain.py diff --git a/inputguard/followups.py b/inputguard/followups.py index a1fba74..1550371 100644 --- a/inputguard/followups.py +++ b/inputguard/followups.py @@ -54,7 +54,24 @@ "Which login method should it use — email and password, Google or GitHub sign-in, an API key, or a magic link?", ), "output format": ( - "What should this be when it's done — a web app, a command-line tool, a REST API, a script, or a mobile app?", + "What should this be when it's done — a web app, a command-line tool, a REST API, a chart, a table, or a written summary?", + ), + "dataset/source": ( + "Which file, table, or export holds the data — by name?", + "What's in {dataset}, and which fields matter for this question?", + ), + "question/goal": ( + "What question should the analysis answer, in one sentence?", + "What decision will this analysis inform?", + ), + "tooling": ( + "Which tools should the analysis use — Python, SQL, a specific library, or anything available?", + ), + "volume": ( + "How much data is it — row count, file size, or the date range it covers?", + ), + "reproducibility": ( + "Is this a one-off look, or will it be rerun — weekly, monthly, at every launch?", ), "task context": ( "What are you trying to build or accomplish, in a sentence or two?", diff --git a/inputguard/recommender.py b/inputguard/recommender.py index d17dc50..b36b7a4 100644 --- a/inputguard/recommender.py +++ b/inputguard/recommender.py @@ -36,9 +36,41 @@ }, "output format": { "gap": "output format", - "what_is_missing": "You didn't say what kind of thing you're building or how it will be used.", - "what_to_provide": "Add something like: 'as a web app I can open in a browser', 'as a command-line tool I run in my terminal', 'as a REST API', 'as a Python script', or 'as a mobile app'. Pick whichever matches how you plan to use it.", - "why_it_matters": "A web app, a script, and an API that do the same job look completely different in code. Without this, the AI picks one and you might get the wrong one entirely.", + # Shared gap string, two senses (coding + data analysis — eval-pinned + # vocabulary): the advice names the deliverable in both worlds. + "what_is_missing": "You didn't say what kind of result you want — what should exist when this is done, and how will it be used?", + "what_to_provide": "Add something like: 'as a web app I can open in a browser', 'as a command-line tool', 'as a REST API' — or, for an analysis: 'a bar chart', 'a one-page summary', 'a trends table', or 'a dashboard'. Pick whichever matches how you plan to use it.", + "why_it_matters": "A web app, a script, an API, and a summary chart that do the same job look completely different in code. Without this, the AI picks one and you might get the wrong one entirely.", + }, + "dataset/source": { + "gap": "dataset/source", + "what_is_missing": "You haven't named the data to analyze — no file, table, export, or database is referenced.", + "what_to_provide": "Point to the data by name: 'sales_2026.csv', 'the attached export', 'the subscriptions table in postgres', or 'the GA4 export'. 'My data' or 'this dataset' isn't enough for it to know where to look.", + "why_it_matters": "Without a named source, the AI has to invent one or ask you anyway. Naming the exact file or table gets you analysis of YOUR data on the first try instead of a template with placeholder numbers.", + }, + "question/goal": { + "gap": "question/goal", + "what_is_missing": "You haven't said what question the analysis should answer or what decision it should inform.", + "what_to_provide": "State the question or goal explicitly: 'did refunds spike after the pricing change?', 'which plan tier cancels most?', 'I want a cohort retention view'. One sentence is enough.", + "why_it_matters": "Without a question, you get a generic summary of everything. With one, the analysis is focused, faster, and actually answers what you needed to know.", + }, + "tooling": { + "gap": "tooling", + "what_is_missing": "You haven't said which tool or language the analysis should use.", + "what_to_provide": "Name the tooling: 'in Python with pandas', 'SQL only', 'in BigQuery', 'statsmodels', or 'no external libraries'. If it doesn't matter, say 'any tool is fine'.", + "why_it_matters": "Different tools mean different workflows. The AI may produce a Python script when your team runs dbt, or SQL your warehouse can't run — naming the tool keeps the output usable as-is.", + }, + "volume": { + "gap": "volume", + "what_is_missing": "You haven't said how much data is involved.", + "what_to_provide": "State the scale: 'about 80,000 rows', 'two years of daily data', '5,000 survey responses', or 'a 400M-row event table'. Row counts, file sizes, or date ranges all work.", + "why_it_matters": "Scale changes the approach. A quick pandas script dies at 400M rows; a full warehouse job is overkill for 500. Saying the size gets you a method that actually runs.", + }, + "reproducibility": { + "gap": "reproducibility", + "what_is_missing": "You haven't said whether this is a one-off or needs to be rerun.", + "what_to_provide": "Say how it will be reused: 'rerun weekly', 'a one-off look', 'document the steps so the team can reproduce it', or 'make the query reusable for every launch'.", + "why_it_matters": "One-off looks and recurring reports are built differently. Saying which one it is decides whether you get a quick answer or a documented, rerunnable pipeline.", }, "task context": { "gap": "task context", diff --git a/inputguard/rules/__init__.py b/inputguard/rules/__init__.py index 203c47f..0a6f434 100644 --- a/inputguard/rules/__init__.py +++ b/inputguard/rules/__init__.py @@ -1,11 +1,11 @@ -"""Built-in rule modules and the registry wiring for the coding and -writing domains. +"""Built-in rule modules and the registry wiring for the coding, writing, +and data-analysis domains. -Importing this package registers both first-party domains — their intent -signals and all built-in rules — through the exact same registry path a -user rule takes. The v0.2 ``run_*_rules`` functions stay exported for -backward compatibility, but the analyzer dispatches through the registry -now. +Importing this package registers all three first-party domains — their +intent signals and all built-in rules — through the exact same registry +path a user rule takes. The v0.2 ``run_*_rules`` functions stay exported +for backward compatibility, but the analyzer dispatches through the +registry now. """ from inputguard.detector import ( @@ -17,6 +17,10 @@ ) from inputguard.registry import REGISTRY from inputguard.rules.coding import CODING_RULES, run_coding_rules +from inputguard.rules.data_analysis import ( + DATA_ANALYSIS_RULES, + DATA_ANALYSIS_SIGNALS, +) from inputguard.rules.debug import DEBUG_RULES, run_debug_rules from inputguard.rules.explanation import EXPLANATION_RULES, run_explanation_rules from inputguard.rules.feature import FEATURE_RULES, run_feature_rules @@ -57,3 +61,13 @@ WRITING_SIGNALS, rules=WRITING_RULES, ) + +# The first-party data-analysis domain (spec §7): the third first-party +# domain, completing the set. The registry name "data-analysis" follows the +# clarity-eval corpus (eval/cases.csv names the domain "data-analysis" and +# the harness passes it verbatim to analyze()). +REGISTRY.register_domain( + "data-analysis", + DATA_ANALYSIS_SIGNALS, + rules=DATA_ANALYSIS_RULES, +) diff --git a/inputguard/rules/data_analysis.py b/inputguard/rules/data_analysis.py new file mode 100644 index 0000000..8e223cd --- /dev/null +++ b/inputguard/rules/data_analysis.py @@ -0,0 +1,418 @@ +"""First-party data-analysis domain: six gap rules for analysis requests. + +The third first-party domain (spec art_bTvdPdJS §7) — the real feature is +the registry; this module exercises it end to end. A data-analysis prompt +("analyze my sales data") fails differently from a coding prompt: the gaps +are the dataset, the question, the deliverable, the tooling, the volume, +and the rerun story — the labeled vocabulary of the clarity-evaluation set +(``eval/cases.csv``, whose ``domain`` column names this domain +``"data-analysis"``; the registry name follows the corpus so the harness +``guard.analyze(text, domain=row["domain"])`` resolves it). + +Design, following the established trigger-and-satisfy shape: + +- Two intents. ``"analysis"`` carries the English trigger terms that mark + an analysis request; ``"reference"`` is the fallback for conceptual + questions about data work ("what is a p-value?") — a clear question is a + clear input, so no rule fires there. Detection is the trigger: the rules + below are satisfy-only, because ``analyze()`` dispatches them only for + the ``analysis`` intent. +- Each rule fires when its gap's satisfying evidence is absent. The + satisfy vocabulary comes from the eval workbook's gap_vocabulary sheet + (what a labeler accepts as "this gap is provided"). +- Gap strings are pinned by the same workbook: ``"output format"`` is + shared with the coding domain's gap string on purpose — the recommender + entry serves both senses. +- Matching goes through the shared word-boundary matcher + (``inputguard.matching``, from the word-boundary PR): standalone words at + the ``(?" phrasing ("find why CSAT dropped") states the goal. +_QUESTION_MARK_RE: Pattern[str] = re.compile(r"\?\s*$") +_FIND_GOAL_RE: Pattern[str] = re.compile( + r"\bfind\s+(?:insights?|why|out|what|which|how|whether|patterns?|drivers?)\b" +) + +DATASET_SATISFIED_TERMS: FrozenSet[str] = frozenset( + { + # provided material + "attached", "pasted", "emailed", "shared", "export", "exports", + # concrete stores and platforms + "database", "databases", "db", "schema", "schemas", "warehouse", + "postgres", "postgresql", "mysql", "bigquery", "snowflake", + "redshift", "databricks", "duckdb", + } +) + +# A named data file (sales.csv, nps_verbatims_2026.xlsx) or an introduced +# dataset/table ("the dataset called churn_2026"). +_DATASET_FILE_RE: Pattern[str] = re.compile( + r"(? " ("80,000 rows", "400m rows", "120,000 users"), +# byte sizes, and stated spans ("90 days of data", "two weeks of beacons"). +_VOLUME_NOUNS = ( + r"(?:rows?|records?|responses?|comments?|entries|events?|users?" + r"|observations?|samples?|items?|transactions?|tickets?|messages?" + r"|clicks?|sessions?|visits?|subscribers?|accounts?|orders?|emails?)" +) +_VOLUME_COUNT_RE: Pattern[str] = re.compile( + r"\b\d[\d.,]*\s*(?:[km]\b)?\s*" + _VOLUME_NOUNS +) +_VOLUME_SIZE_RE: Pattern[str] = re.compile(r"\b\d[\d.,]*\s*(?:kb|mb|gb|tb|bytes)\b") +_VOLUME_INTERVAL_RE: Pattern[str] = re.compile( + r"\b\d[\d.,]*\s*(?:days?|weeks?|months?|years?|quarters?)\s+of\b" +) +_VOLUME_WORD_INTERVAL_RE: Pattern[str] = re.compile( + r"\b(?:one|two|three|four|five|six|seven|eight|nine|ten|eleven|twelve" + r"|dozens|hundreds|thousands|millions)\s+" + r"(?:days?|weeks?|months?|years?|quarters?)\s+of\b" +) + + +# -- matching helpers ---------------------------------------------------------- +# Term lookups go through the shared word-boundary matcher +# (inputguard.matching): standalone words at the (? str: + return re.sub(r"\s+", " ", text.strip().lower()) + + +def _is_satisfied(text: str, terms: FrozenSet[str], *regexes: Pattern[str]) -> bool: + return contains_any(text, terms) or any(regex.search(text) for regex in regexes) + + +# -- the six gap rules --------------------------------------------------------- + + +def _check_missing_dataset_source(text: str) -> Optional[RuleFinding]: + """Fires when no concrete data source is named. + + "my sales data", "this dataset", and "the spreadsheet" do not satisfy — + the labeler requires a file name, a provided export/attachment, or a + named store (postgres, BigQuery, the warehouse, a schema). + """ + text = _normalize(text) + if not _is_satisfied( + text, DATASET_SATISFIED_TERMS, _DATASET_FILE_RE, _DATASET_NAMED_RE + ): + return RuleFinding( + code="missing_dataset_source", + message=( + "Analysis request detected but no concrete data source named — " + "no file, table, export, or database referenced." + ), + severity="high", + gap="dataset/source", + ) + return None + + +def _check_missing_question_goal(text: str) -> Optional[RuleFinding]: + """Fires when no explicit question, goal, or comparison is stated.""" + text = _normalize(text) + if not _is_satisfied( + text, QUESTION_SATISFIED_TERMS, _QUESTION_MARK_RE, _FIND_GOAL_RE + ): + return RuleFinding( + code="missing_question_goal", + message=( + "Analysis request detected but no explicit question, goal, " + "or comparison stated." + ), + severity="high", + gap="question/goal", + ) + return None + + +def _check_missing_deliverable_format(text: str) -> Optional[RuleFinding]: + """Fires when no deliverable shape is named. + + Registry adapter rule id is ``missing_deliverable_format`` — the coding + domain owns the ``missing_output_format`` id; the GAP string + ("output format") is the shared, eval-pinned vocabulary. + """ + text = _normalize(text) + if not _is_satisfied(text, OUTPUT_SATISFIED_TERMS): + return RuleFinding( + code="missing_deliverable_format", + message=( + "Analysis request detected but no deliverable shape " + "specified — no chart, table, summary, or report named." + ), + severity="medium", + gap="output format", + ) + return None + + +def _check_missing_tooling(text: str) -> Optional[RuleFinding]: + """Fires when no tool or library constraint is given.""" + text = _normalize(text) + if not _is_satisfied(text, TOOLING_SATISFIED_TERMS, _TOOLING_R_RE): + return RuleFinding( + code="missing_tooling", + message=( + "Analysis request detected but no tool or library " + "constraint given." + ), + severity="medium", + gap="tooling", + ) + return None + + +def _check_missing_volume(text: str) -> Optional[RuleFinding]: + """Fires when no data scale is stated.""" + text = _normalize(text) + if not _is_satisfied( + text, + frozenset(), + _VOLUME_COUNT_RE, + _VOLUME_SIZE_RE, + _VOLUME_INTERVAL_RE, + _VOLUME_WORD_INTERVAL_RE, + ): + return RuleFinding( + code="missing_volume", + message=( + "Analysis request detected but no data scale stated — no " + "row count, file size, or date range." + ), + severity="low", + gap="volume", + ) + return None + + +def _check_missing_reproducibility(text: str) -> Optional[RuleFinding]: + """Fires when no rerun or reuse expectation is stated.""" + text = _normalize(text) + if not _is_satisfied(text, REPRO_SATISFIED_TERMS): + return RuleFinding( + code="missing_reproducibility", + message=( + "Analysis request detected but no rerun or reuse " + "expectation stated." + ), + severity="low", + gap="reproducibility", + ) + return None + + +# Registry adapters: the check functions above stay the single home of the +# rule logic; these classes expose it through the v0.3 Rule protocol and +# register it through the same path a user rule takes. + + +@register_rule +class MissingDatasetSourceRule: + """Registry adapter for _check_missing_dataset_source.""" + + id = "missing_dataset_source" + domain = "analysis" + severity = "high" + gap = "dataset/source" + + def check(self, text: str) -> Optional[RuleFinding]: + return _check_missing_dataset_source(text) + + +@register_rule +class MissingQuestionGoalRule: + """Registry adapter for _check_missing_question_goal.""" + + id = "missing_question_goal" + domain = "analysis" + severity = "high" + gap = "question/goal" + + def check(self, text: str) -> Optional[RuleFinding]: + return _check_missing_question_goal(text) + + +@register_rule +class MissingDeliverableFormatRule: + """Registry adapter for _check_missing_deliverable_format.""" + + id = "missing_deliverable_format" + domain = "analysis" + severity = "medium" + gap = "output format" + + def check(self, text: str) -> Optional[RuleFinding]: + return _check_missing_deliverable_format(text) + + +@register_rule +class MissingToolingRule: + """Registry adapter for _check_missing_tooling.""" + + id = "missing_tooling" + domain = "analysis" + severity = "medium" + gap = "tooling" + + def check(self, text: str) -> Optional[RuleFinding]: + return _check_missing_tooling(text) + + +@register_rule +class MissingVolumeRule: + """Registry adapter for _check_missing_volume.""" + + id = "missing_volume" + domain = "analysis" + severity = "low" + gap = "volume" + + def check(self, text: str) -> Optional[RuleFinding]: + return _check_missing_volume(text) + + +@register_rule +class MissingReproducibilityRule: + """Registry adapter for _check_missing_reproducibility.""" + + id = "missing_reproducibility" + domain = "analysis" + severity = "low" + gap = "reproducibility" + + def check(self, text: str) -> Optional[RuleFinding]: + return _check_missing_reproducibility(text) + + +DATA_ANALYSIS_RULES = ( + MissingDatasetSourceRule, + MissingQuestionGoalRule, + MissingDeliverableFormatRule, + MissingToolingRule, + MissingVolumeRule, + MissingReproducibilityRule, +) diff --git a/tests/test_data_analysis_domain.py b/tests/test_data_analysis_domain.py new file mode 100644 index 0000000..653e247 --- /dev/null +++ b/tests/test_data_analysis_domain.py @@ -0,0 +1,313 @@ +"""First-party data-analysis domain: rule behavior, vocabulary completeness, +eval-set alignment, and cross-domain leakage guards. + +The six rules follow the pinned eval vocabulary (eval/cases.csv, +gap_vocabulary): dataset/source, question/goal, output format, tooling, +volume, reproducibility. The eval-alignment tests read the versioned CSV +labels directly, so any future label edit keeps these tests honest. +""" + +from __future__ import annotations + +import csv +import json +from pathlib import Path +from typing import Dict, FrozenSet, List + +import pytest + +from inputguard import InputGuard +from inputguard.followups import _FOLLOW_UP_QUESTIONS +from inputguard.recommender import _RECOMMENDATIONS +from inputguard.registry import REGISTRY +from inputguard.rules.data_analysis import ( + DATA_ANALYSIS_RULES, + DATA_ANALYSIS_SIGNALS, + MissingDatasetSourceRule, + MissingDeliverableFormatRule, + MissingQuestionGoalRule, + MissingReproducibilityRule, + MissingToolingRule, + MissingVolumeRule, +) + +# The registry name follows the eval corpus: cases.csv names the domain +# "data-analysis" and the harness passes it verbatim to analyze(). +DOMAIN = "data-analysis" + +# The six built-in data-analysis rule ids, by gap. +_GAP_TO_RULE = { + "dataset/source": MissingDatasetSourceRule, + "question/goal": MissingQuestionGoalRule, + "output format": MissingDeliverableFormatRule, + "tooling": MissingToolingRule, + "volume": MissingVolumeRule, + "reproducibility": MissingReproducibilityRule, +} + +DA_RULE_IDS: FrozenSet[str] = frozenset(rule.id for rule in DATA_ANALYSIS_RULES) +DA_GAPS: FrozenSet[str] = frozenset(_GAP_TO_RULE) + + +def _eval_da_cases() -> List[Dict[str, str]]: + """The versioned data-analysis rows of the eval corpus.""" + path = Path(__file__).resolve().parent.parent / "eval" / "cases.csv" + with open(path, newline="", encoding="utf-8") as handle: + rows = [row for row in csv.DictReader(handle) if row["domain"] == DOMAIN] + assert len(rows) == 15, "the eval corpus should carry 15 data-analysis cases" + return rows + + +_NONE_SENTINELS = {"none", "n/a", ""} + + +def _split_multi(value: str) -> List[str]: + """Parse a semicolon-joined multi-value column, as the harness does. + + ``none``/``n/a``/empty mean no entries (eval/measure_fp.py). + """ + stripped = value.strip() + if stripped.lower() in _NONE_SENTINELS: + return [] + return [part.strip() for part in stripped.split(";") if part.strip()] + + +# --- registration ------------------------------------------------------------- + + +def test_data_analysis_domain_registered(): + assert "data-analysis" in REGISTRY.domain_names() + assert DA_RULE_IDS <= set(REGISTRY.rule_ids()) + + +def test_domain_signals_have_two_intents_with_reference_fallback(): + chain = REGISTRY.get_domain_signals(DOMAIN) + assert [intent for intent, _ in chain] == ["analysis", "reference"] + assert [terms for _, terms in chain if not terms] == [()] + + +def test_data_analysis_signals_shape(): + intents = [intent for intent, _ in DATA_ANALYSIS_SIGNALS] + assert intents[0] == "analysis" + assert len(DATA_ANALYSIS_SIGNALS) == 2 + assert DATA_ANALYSIS_SIGNALS[-1] == ("reference", ()) + + +def test_every_data_analysis_rule_scopes_to_the_analysis_intent(): + for rule in DATA_ANALYSIS_RULES: + assert rule.domain == "analysis" + + +# --- per-rule fire / silence ---------------------------------------------------- + + +@pytest.mark.parametrize( + ("rule_id", "prompt", "gap", "severity"), + [ + # Each prompt satisfies the other five gaps so exactly one rule fires. + ( + "missing_dataset_source", + "Analyze the churn patterns for the board, deliver a summary " + "slide, using pandas in python, about 80,000 rows, and we rerun " + "it weekly.", + "dataset/source", + "high", + ), + ( + "missing_question_goal", + "Analyze the attached export and deliver a summary chart, using " + "pandas, about 80,000 rows, rerun weekly.", + "question/goal", + "high", + ), + ( + "missing_deliverable_format", + "Analyze the attached export to see whether churn improved, " + "using pandas, about 80,000 rows, rerun weekly.", + "output format", + "medium", + ), + ( + "missing_tooling", + "Analyze the attached export to see whether churn improved and " + "deliver a summary chart, about 80,000 rows, rerun weekly.", + "tooling", + "medium", + ), + ( + "missing_volume", + "Analyze the attached export to see whether churn improved, " + "deliver a summary chart, using pandas, rerun weekly.", + "volume", + "low", + ), + ( + "missing_reproducibility", + "Analyze the attached export to see whether churn improved, " + "deliver a summary chart, using pandas, about 80,000 rows.", + "reproducibility", + "low", + ), + ], +) +def test_each_rule_fires_isolated(prompt, rule_id, gap, severity): + result = InputGuard().analyze(prompt, domain=DOMAIN) + assert result.gaps == [gap] + finding = next(f for f in result.findings if f.code == rule_id) + assert finding.severity == severity + assert finding.gap == gap + + +@pytest.mark.parametrize( + ("rule_cls", "satisfying_suffix"), + [ + (MissingDatasetSourceRule, "in the attached postgres export"), + (MissingQuestionGoalRule, "to find why CSAT dropped"), + (MissingDeliverableFormatRule, "and deliver a summary chart"), + (MissingToolingRule, "using pandas in python"), + (MissingVolumeRule, "on about 80,000 rows"), + (MissingReproducibilityRule, "and we rerun it weekly"), + ], +) +def test_each_rule_stays_silent_when_gap_satisfied(rule_cls, satisfying_suffix): + rule = rule_cls() + finding = rule.check("analyze my sales data " + satisfying_suffix) + assert finding is None + + +def test_boundary_safe_matching_on_satisfy_side(): + # A term embedded mid-word must not satisfy: "unshared" is not a shared + # source. The shared matcher bounds every term at word boundaries. + assert MissingDatasetSourceRule().check("analyze the unshared metrics") is not None + # The plain word still satisfies. + assert MissingDatasetSourceRule().check("analyze the export metrics") is None + # Suffix-shaped words at a boundary are genuine inflections under the + # shared matcher's contract: "exporter" counts as "export" (-er ending). + assert MissingDatasetSourceRule().check("analyze the exporter metrics") is None + + +def test_volume_regexes_cover_labeled_scale_forms(): + rule = MissingVolumeRule() + for text in ( + "analyze the attached export, about 80,000 rows, rerun weekly", + "analyze 400m rows of clickstream", + "analyze 90 days of data in the warehouse", + "analyze two weeks of beacons from the warehouse", + ): + assert rule.check(text) is None, text + + +# --- eval-set alignment ----------------------------------------------------------- + + +@pytest.mark.parametrize("row", _eval_da_cases(), ids=lambda row: row["id"]) +def test_eval_alignment(row): + result = InputGuard().analyze(row["text"], domain=DOMAIN) + expected_gaps = set(_split_multi(row["expected_gaps"])) + expected_status = row["expected_status"] + assert set(result.gaps) == expected_gaps, row["id"] + assert result.status == expected_status, row["id"] + + +def test_eval_severities_match_pinned_vocabulary(): + for row in _eval_da_cases(): + if not _split_multi(row["expected_severities"]): + continue + result = InputGuard().analyze(row["text"], domain=DOMAIN) + actual = {f.severity for f in result.findings} + assert actual == set(_split_multi(row["expected_severities"])), row["id"] + + +# --- registry completeness invariant (domain-specific) ----------------------------- + + +def test_every_data_analysis_gap_has_a_recommendation(): + for rule in DATA_ANALYSIS_RULES: + assert rule.gap in _RECOMMENDATIONS, rule.id + entry = _RECOMMENDATIONS[rule.gap] + assert set(entry) == {"gap", "what_is_missing", "what_to_provide", "why_it_matters"} + assert entry["gap"] == rule.gap + + +def test_every_data_analysis_gap_has_follow_up_questions(): + for rule in DATA_ANALYSIS_RULES: + questions = _FOLLOW_UP_QUESTIONS[rule.gap] + assert 1 <= len(questions) <= 2, rule.id + for question in questions: + assert question.endswith("?"), (rule.id, question) + # At least one slot-free template: it renders on any input. + assert any("{" not in q for q in questions), rule.id + + +def test_rule_severities_match_pinned_vocabulary(): + path = Path(__file__).resolve().parent.parent / "eval" / "cases.csv" + # The gap vocabulary's default severities are pinned by the workbook; the + # eval corpus rows carry the same per-gap severities. + with open(path, newline="", encoding="utf-8") as handle: + rows = [row for row in csv.DictReader(handle) if row["domain"] == DOMAIN] + for row in rows: + gaps = _split_multi(row["expected_gaps"]) + severities = _split_multi(row["expected_severities"]) + for gap, severity in zip(gaps, severities): + rule = _GAP_TO_RULE[gap] + assert rule.severity == severity, (row["id"], gap) + + +# --- analyze() integration ----------------------------------------------------------- + + +def test_underspecified_analysis_prompt_end_to_end(): + result = InputGuard().analyze("Analyze my sales data", domain=DOMAIN) + assert result.detected_intent == "analysis" + assert set(result.gaps) == DA_GAPS + assert result.status == "needs_clarification" + assert result.clarity_score == 10 # 100 - (25 + 25 + 15 + 15 + 5 + 5) + # Recommendations cover each gap, in order. + assert [rec["gap"] for rec in result.recommendations] == result.gaps + assert all(set(rec) == {"gap", "what_is_missing", "what_to_provide", "why_it_matters"} for rec in result.recommendations) + # Follow-ups cover every gap. + assert len(result.follow_ups) >= len(result.gaps) + assert all(q.endswith("?") for q in result.follow_ups) + + +def test_to_dict_round_trips_as_json(): + result = InputGuard().analyze("Analyze my sales data", domain=DOMAIN) + payload = json.dumps(result.to_dict()) + assert json.loads(payload)["clarity_score"] == result.clarity_score + + +def test_reference_intent_clears_without_findings(): + result = InputGuard().analyze("What is a p-value?", domain=DOMAIN) + assert result.detected_intent == "reference" + assert result.findings == [] + assert result.status == "ready" + + +# --- no-leakage guards ---------------------------------------------------------------- + + +def test_data_analysis_rules_never_fire_in_coding_analyses(): + result = InputGuard().analyze( + "Build a REST API using FastAPI. Store users in PostgreSQL with " + "fields for name and email. Add email and password login.", + domain="coding", + ) + fired = {f.code for f in result.findings} + assert fired & DA_RULE_IDS == set() + + +def test_coding_rules_never_fire_in_data_analysis_analyses(): + coding_ids = set(REGISTRY.rule_ids()) - DA_RULE_IDS + result = InputGuard().analyze("Analyze my sales data", domain=DOMAIN) + fired = {f.code for f in result.findings} + assert fired <= DA_RULE_IDS + assert fired & coding_ids == set() + + +def test_vague_coding_request_stays_a_coding_result(): + # A vague prompt sent to the coding domain must get coding behavior, + # not data-analysis rules. + result = InputGuard().analyze("Analyze my sales data", domain="coding") + fired = {f.code for f in result.findings} + assert fired & DA_RULE_IDS == set() + assert fired # the coding build rules still do their job diff --git a/tests/test_registry.py b/tests/test_registry.py index 4b74cad..59c5a28 100644 --- a/tests/test_registry.py +++ b/tests/test_registry.py @@ -24,7 +24,9 @@ from inputguard.rules.optimization import run_optimization_rules from inputguard.scorer import calculate_score -# The 19 built-in rule ids, as documented in the README table. +# The 31 built-in rule ids: the 19 coding rules documented in the README +# table plus the six first-party writing rules and the six first-party +# data-analysis rules (spec §7). EXPECTED_BUILTIN_IDS = { "missing_language", "missing_api_structure", @@ -52,6 +54,13 @@ "missing_source_material", "missing_writing_context", "missing_completeness", + # First-party data-analysis domain (6 rules, spec §7). + "missing_dataset_source", + "missing_question_goal", + "missing_deliverable_format", + "missing_tooling", + "missing_volume", + "missing_reproducibility", } _V02_INTENT_RUNNERS = { @@ -92,13 +101,14 @@ def registry_isolation(): REGISTRY._origins.update(origins_before) -# --- the 25 built-ins dogfood the registry path --------------------------- +# --- the 31 built-ins dogfood the registry path --------------------------- def test_all_builtin_rules_registered_through_registry(): - # 19 coding + 6 writing built-ins. The exact-set guard is the point: - # a rule that stops registering (or an unregistered stray) fails here. - assert len(REGISTRY.rule_ids()) == 25 + # 19 coding + 6 writing + 6 data-analysis built-ins. The exact-set guard + # is the point: a rule that stops registering (or an unregistered stray) + # fails here. + assert len(REGISTRY.rule_ids()) == 31 assert set(REGISTRY.rule_ids()) == EXPECTED_BUILTIN_IDS for rule in REGISTRY.rules(): assert isinstance(rule, Rule) @@ -106,9 +116,7 @@ def test_all_builtin_rules_registered_through_registry(): def test_coding_domain_registered_with_priority_signals(): - # Both first-party domains register at import: coding first, writing - # second. - assert REGISTRY.domain_names() == ("coding", "writing") + assert "coding" in REGISTRY.domain_names() chain = REGISTRY.get_domain_signals("coding") assert [intent for intent, _ in chain] == [ "debug", From 4112d8fabddbb229a1980ca4e1328fb03ea7d3ba Mon Sep 17 00:00:00 2001 From: "obvious-autobuild[bot]" <262744130+obvious-autobuild[bot]@users.noreply.github.com> Date: Thu, 17 Sep 2026 21:07:39 +0000 Subject: [PATCH 11/18] fix: degradation coverage for DG-011..014 (detector) (#12) Function-word layer (fr/es/pt stop-word lexicons with English margin) plus uncovered-run layer (3+ consecutive uncovered-script letters) route accented-Latin and mixed English+Han input to the explicit degraded path; DG-012's silent ready and DG-011/013/014's unnoted rule findings are gone. Degradation eval FP 3/14 -> 0/14, note honesty 10/14 -> 14/14, all 121 probe verdicts otherwise byte-identical. --- CHANGELOG.md | 6 + README.md | 11 ++ inputguard/language.py | 244 +++++++++++++++++++++++++++++++++++++++-- tests/test_language.py | 180 ++++++++++++++++++++++++++++++ 4 files changed, 430 insertions(+), 11 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 4879789..b16d681 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -13,6 +13,12 @@ `heuristic_coverage`, `degradation_note` (included in `to_dict()`). - Mixed-script input: rules run whenever the covered share of letters is at least 50%, with a `partial` note between 50–70% coverage. +- Accented-Latin honesty (eval DG-011..014): French, Spanish, and + Portuguese input no longer passes as silently covered — a function-word + layer degrades them like any other uncovered language, and a run of 3+ + consecutive uncovered-script letters (mixed English+Han) degrades even + majority-English input. English prompts (loanwords, URLs, name and + timezone collisions) are pinned unchanged by tests. ## [0.2.0] — 2026-05-29 ### Added diff --git a/README.md b/README.md index 8c4d206..21a4cb9 100644 --- a/README.md +++ b/README.md @@ -122,6 +122,17 @@ Three additive fields on the result carry the probe's verdict: Mixed input is handled by share, not by exclusion: `"make it faster 这个"` is still fully analyzed (English dominates and the rules run); input whose letters fall 50–70% inside coverage gets a `partial` note without a penalty. +Latin script alone is not coverage, though: the rules are English-only, and French, Spanish, or Portuguese text is 100% Latin script yet just as unreadable to them. The probe therefore also recognizes those three languages by their function words — enough distinct stop-word hits, with a margin over the input's English function-word evidence, degrades the input like any other uncovered language. German, Italian, Dutch, and other Latin-script languages are a documented blind spot and still pass as before. And a run of 3+ consecutive letters in an uncovered script degrades even a majority-English input ("Fix this bug 修复这个错误 in the payment flow"); a 2-letter borrow like `"这个"` still rides along. + +```python +result = guard.analyze("Preciso de um aplicativo web com login de usuário e relatórios") + +result.detected_language # 'pt' (function-word guess) +result.heuristic_coverage # 'none' +result.status # 'usable_with_warnings' — never 'ready' +result.degradation_note # names Portuguese (pt) and explains the skip +``` + --- ## How intent detection works diff --git a/inputguard/language.py b/inputguard/language.py index 13be44e..38d4e3f 100644 --- a/inputguard/language.py +++ b/inputguard/language.py @@ -18,6 +18,26 @@ ``ready``. Scripts with coverage (Latin) behave exactly as they did before the probe existed. +Latin script alone is not sufficient coverage, though: the heuristics are +English-only, and French, Spanish, or Portuguese text is 100% Latin script +yet just as unreadable to them as Chinese. Two refinements close that gap, +still with zero dependencies (eval rows DG-011..014): + +- **Function-word layer** — when the dominant script is Latin, the sample's + tokens are matched against small stop-word lexicons for French, Spanish, + and Portuguese. A language is "detected" only when it has at least + ``_MIN_DISTINCT_HITS`` distinct stop-word hits AND those beats the + input's English function-word evidence by ``_MIN_ENGLISH_MARGIN`` — so + genuinely English prompts (rich in English function words) can never + trip it over scattered collisions ("Pour the DES dump into LA storage"). + The lexicon languages are a documented subset: German, Italian, Dutch, + and other Latin-script languages still pass as before. +- **Uncovered-run layer** — a run of at least ``_UNCOVERED_BLOCK_DEGRADES_AT`` + consecutive letters in an uncovered script (e.g. a Han clause inside an + English sentence) downgrades nominal full coverage to the degraded path. + A two-character borrow like "这个" rides along (rules still run on the + English majority); a whole clause the rules cannot read cannot. + The probe is pure and thread-safe: no module-level mutation, only stdlib lookups, matching the ``analyze()`` thread-safety contract. """ @@ -157,6 +177,104 @@ "khmer": "km", } +# --------------------------------------------------------------------------- +# Latin-script language layer (eval rows DG-011, DG-012, DG-014) +# +# The script histogram alone cannot see French, Spanish, or Portuguese: their +# letters are Latin, so a purely script-level probe reports 100% coverage and +# the English-only heuristics run on text they cannot read — the same silent +# ready the probe exists to prevent. The labeling guide's rubric for these +# rows prescribes the fix: a stop-word probe that refuses to let "Latin +# script alone silently pass". +# +# The lexicons are deliberately tiny and conservative — function words +# (articles, pronouns, prepositions, conjunctions, question words) only: +# - every entry is >= 2 letters and is not a standalone English word, so +# scattered collisions (French "pour" = English "pour", "est" = the EST +# timezone) cannot manufacture a detection on their own; +# - matching runs on accent-stripped tokens ("qué" -> "que") and on tokens +# trimmed of edge punctuation ("¿Por" -> "por"), never on interior +# punctuation ("example.com" or "de-serialization" stay single tokens and +# cannot match "com" / "de"); +# - detection additionally requires the margin rule in ``_latin_language_of`` +# (distinct non-English hits must beat English function-word evidence by +# ``_MIN_ENGLISH_MARGIN``), so genuinely English prompts — which are dense +# in English function words — cannot trip the layer at all. +# +# This is an honest documented subset: German, Italian, Dutch, and other +# Latin-script languages still pass as before, because a word list for every +# language is neither shippable nor zero-dependency-honest. The probe reports +# a *guess* ("guessed language"), never a certainty. +# --------------------------------------------------------------------------- +_MIN_DISTINCT_HITS = 3 +_MIN_ENGLISH_MARGIN = 3 +# When the input shows zero English function-word evidence, two distinct +# foreign function words already fire: "Oui merci beaucoup pour ton aide" +# carries no English function words at all, and demanding three would miss +# it. A collision-only English prompt always has English evidence (the/a/ +# and/...), so the relaxed floor cannot trip it. +_MIN_SOLO_HITS = 2 + +# Run length (in consecutive letters) of an uncovered script inside an +# otherwise-covered input that forces the degraded path (eval row DG-013). +# A 1-2 letter borrow ("这个") rides along under the English majority; a run +# of 3+ letters is a word or clause the English heuristics cannot read. +_UNCOVERED_BLOCK_DEGRADES_AT = 3 + +_LATIN_FUNCTION_WORDS: Dict[str, frozenset] = { + "en": frozenset({ + "a", "an", "the", "and", "or", "but", "if", "then", "of", "to", "in", + "on", "at", "by", "for", "with", "from", "into", "over", "under", + "up", "down", "out", "off", "about", "after", "before", "between", + "during", "through", "without", "within", "is", "are", "was", "were", + "be", "been", "being", "am", "do", "does", "did", "have", "has", + "had", "will", "would", "can", "could", "should", "shall", "may", + "might", "must", "i", "you", "he", "she", "it", "we", "they", "me", + "him", "us", "them", "my", "your", "his", "its", "our", "their", + "this", "that", "these", "those", "there", "here", "who", "whom", + "whose", "which", "what", "when", "where", "why", "how", "all", + "any", "some", "each", "every", "both", "few", "many", "much", + "more", "most", "very", "too", "also", "just", "only", "again", + "once", "now", "not", "no", "yes", "so", "such", "than", "as", + "because", "while", "per", "via", "please", + }), + "es": frozenset({ + "el", "los", "las", "una", "unos", "unas", "mi", "mis", "tu", "tus", + "su", "sus", "pero", "para", "por", "que", "como", "donde", + "cuando", "cual", "cuanto", "esto", "esta", "este", "eso", "ese", + "esa", "esos", "esas", "con", "sin", "se", "es", "muy", "mas", + "soy", "eres", "puedo", "puedes", "quiero", "necesito", "gracias", + "hola", "porque", "tambien", "de", + }), + "fr": frozenset({ + "le", "la", "les", "des", "un", "une", "du", "au", "aux", "et", + "est", "dans", "pour", "avec", "sur", "ce", "cet", "cette", "ces", + "je", "tu", "il", "elle", "nous", "vous", "ils", "elles", "mon", + "ton", "son", "mes", "tes", "ses", "leur", "leurs", "qui", "que", + "quoi", "dont", "ou", "mais", "donc", "chez", "tres", "aussi", + "pourquoi", "etre", "avoir", "sans", "entre", "de", + }), + "pt": frozenset({ + "um", "uma", "uns", "umas", "de", "da", "dos", "das", "na", "nas", + "nos", "com", "para", "por", "que", "como", "muito", "mas", "meu", + "minha", "meus", "minhas", "isso", "isto", "estou", "preciso", + "obrigado", "obrigada", "quero", "posso", "onde", "porque", + "tambem", "seu", "sua", "sem", + }), +} + +# Note-facing language names for lexicon-detected Latin-script inputs. +_LATIN_LANGUAGE_NAMES = {"fr": "French", "es": "Spanish", "pt": "Portuguese"} + +# Punctuation that French/Spanish/Portuguese elides into the next word +# ("l'homme", "n'est") — split it so the pieces can match their lexicons. +_APOSTROPHES = str.maketrans({"'": " ", "\u2019": " ", "`": " "}) + +# Trim edge punctuation ("¿Por" -> "Por", "bord." -> "bord") while keeping +# interior punctuation intact ("example.com" stays one token, so a URL can +# never contribute a stray "com" hit). +_EDGE_TRIM = re.compile(r"^[^\w]+|[^\w]+$") + @dataclass(frozen=True) class ScriptProbe: @@ -164,12 +282,19 @@ class ScriptProbe: ``heuristic_coverage`` is the field ``analyze()`` branches on; the other fields feed the degradation note and the result's ``detected_language``. + + ``uncovered_block_script`` / ``uncovered_block_length`` record the longest + run of consecutive uncovered-script letters in the sample (0 when every + classified letter is Latin). They explain *why* a Latin-dominant input + was degraded — the mixed English+Han case — and feed the note's wording. """ detected_language: str dominant_script: Optional[str] covered_share: float heuristic_coverage: str + uncovered_block_script: Optional[str] = None + uncovered_block_length: int = 0 def _sample(text: str, limit: int) -> str: @@ -203,6 +328,54 @@ def _script_of(char: str) -> Optional[str]: return None +def _strip_diacritics(token: str) -> str: + """Accent-fold one token for lexicon matching: "qué" -> "que". + + NFD decomposition splits accented letters into base letter + combining + mark; the marks (category Mn) drop out. Purely lexical — the histogram + itself keeps every letter, accented or not. + """ + decomposed = unicodedata.normalize("NFD", token) + return "".join(ch for ch in decomposed if unicodedata.category(ch) != "Mn") + + +def _latin_language_of(sample: str) -> Optional[str]: + """Guess a non-English Latin-script language from function words, or ``None``. + + A language "fires" only when its distinct stop-word hits reach + :data:`_MIN_DISTINCT_HITS` AND beat the input's distinct English + function-word hits by :data:`_MIN_ENGLISH_MARGIN` — English prompts are + dense in English function words, so scattered collisions with French or + Spanish tokens ("pour", "est", "la") can never fire the layer alone. + Ties between firing languages resolve alphabetically, like the script + histogram's dominant-script tie-break. + """ + distinct: Dict[str, set] = {} + for raw in sample.translate(_APOSTROPHES).lower().split(): + token = _EDGE_TRIM.sub("", raw) + if not token: + continue + stripped = _strip_diacritics(token) + for lang, words in _LATIN_FUNCTION_WORDS.items(): + if stripped in words: + distinct.setdefault(lang, set()).add(stripped) + english = len(distinct.get("en", set())) + # English-dense inputs need the full margin; zero-English inputs fire on + # the relaxed floor (see _MIN_SOLO_HITS). + floor = _MIN_DISTINCT_HITS if english else _MIN_SOLO_HITS + firing = [ + (len(words), lang) + for lang, words in distinct.items() + if lang != "en" + and len(words) >= floor + and (english == 0 or len(words) >= english + _MIN_ENGLISH_MARGIN) + ] + if not firing: + return None + firing.sort(key=lambda item: (-item[0], item[1])) + return firing[0][1] + + def probe_script(text: str) -> ScriptProbe: """Classify an input's script coverage with a :mod:`unicodedata` histogram. @@ -211,16 +384,34 @@ def probe_script(text: str) -> ScriptProbe: letters whose script the probe recognizes; scripts outside :data:`_SCRIPT_KEYWORDS` (Runic, Deseret, ...) leave the covered-share denominator, which is the honest signal that heuristics do not apply. + + On top of the share histogram, two refinements route inputs the share + alone would wrongly call fully covered to the degraded path: Latin-script + text recognized as French/Spanish/Portuguese by function words + (:func:`_latin_language_of`), and Latin-dominant text carrying a run of + uncovered-script letters too long to ride along (``_UNCOVERED_BLOCK_DEGRADES_AT``). """ letter_count = 0 script_counts: Dict[str, int] = {} - for char in _sample(text, _SAMPLE_LIMIT): + uncovered_run = 0 + max_uncovered_run = 0 + max_uncovered_script: Optional[str] = None + sample = _sample(text, _SAMPLE_LIMIT) + for char in sample: if not unicodedata.category(char).startswith("L"): + uncovered_run = 0 # any non-letter breaks a run continue letter_count += 1 script = _script_of(char) if script is not None: script_counts[script] = script_counts.get(script, 0) + 1 + if script == "latin": + uncovered_run = 0 + else: + uncovered_run += 1 + if uncovered_run > max_uncovered_run: + max_uncovered_run = uncovered_run + max_uncovered_script = script classified = sum(script_counts.values()) if classified == 0: @@ -231,25 +422,43 @@ def probe_script(text: str) -> ScriptProbe: dominant_script=None, covered_share=0.0, heuristic_coverage=coverage, + uncovered_block_script=max_uncovered_script if letter_count > 0 else None, + uncovered_block_length=max_uncovered_run if letter_count > 0 else 0, ) covered_share = script_counts.get("latin", 0) / classified - if covered_share >= _FULL_COVERAGE_AT: - coverage = COVERAGE_FULL - elif covered_share >= _NONE_COVERAGE_BELOW: - coverage = COVERAGE_PARTIAL - else: - coverage = COVERAGE_NONE - # Deterministic dominant script: highest count, alphabetical tie-break. dominant = min(script_counts, key=lambda s: (-script_counts[s], s)) - detected_language = _detected_language(dominant, script_counts) + + latin_language = _latin_language_of(sample) if dominant == "latin" else None + if latin_language is not None: + # French/Spanish/Portuguese: 100% Latin script, yet the English-only + # heuristics cannot read a word of it — degraded like any other + # uncovered language, never a silent pass on nominal coverage. + coverage = COVERAGE_NONE + detected_language = latin_language + else: + detected_language = _detected_language(dominant, script_counts) + if covered_share >= _FULL_COVERAGE_AT: + if max_uncovered_run >= _UNCOVERED_BLOCK_DEGRADES_AT: + # Latin-dominant but carrying a substantive uncovered-script + # clause (mixed English+Han, DG-013): the rules cannot read + # that content, so full coverage is a false claim. + coverage = COVERAGE_NONE + else: + coverage = COVERAGE_FULL + elif covered_share >= _NONE_COVERAGE_BELOW: + coverage = COVERAGE_PARTIAL + else: + coverage = COVERAGE_NONE return ScriptProbe( detected_language=detected_language, dominant_script=dominant, covered_share=covered_share, heuristic_coverage=coverage, + uncovered_block_script=max_uncovered_script, + uncovered_block_length=max_uncovered_run, ) @@ -263,11 +472,16 @@ def _detected_language(dominant: str, script_counts: Dict[str, int]) -> str: def degradation_note_for(probe: ScriptProbe) -> str: """The note carried by results on the degraded (uncovered) path.""" - if probe.dominant_script is None: + if probe.detected_language in _LATIN_LANGUAGE_NAMES: + script_desc = ( + f"the {probe.dominant_script} script, but function words read as " + f"{_LATIN_LANGUAGE_NAMES[probe.detected_language]}" + ) + elif probe.dominant_script is None: script_desc = "a script outside the probe's coverage" else: script_desc = f"the {probe.dominant_script} script" - return ( + note = ( f"InputGuard's clarity rules are English-language heuristics; this " f"input appears to use {script_desc} (guessed language: " f"{probe.detected_language}), which they do not cover. Rule analysis " @@ -276,6 +490,14 @@ def degradation_note_for(probe: ScriptProbe) -> str: "input's clarity. Rephrasing the key details in English enables a " "full analysis." ) + if probe.uncovered_block_length >= _UNCOVERED_BLOCK_DEGRADES_AT: + note += ( + " The input also mixes scripts: a run of " + f"{probe.uncovered_block_length} letters in " + f"{probe.uncovered_block_script or 'an uncovered'} script sits " + "outside the heuristics' coverage." + ) + return note def partial_coverage_note(probe: ScriptProbe) -> str: diff --git a/tests/test_language.py b/tests/test_language.py index 997a1ab..3481bc0 100644 --- a/tests/test_language.py +++ b/tests/test_language.py @@ -427,3 +427,183 @@ def run(pair): for text, mode, result in results: baseline = expected_strict if mode == "strict" else expected assert result == baseline[text] + + +# --------------------------------------------------------------------------- +# Accented-Latin and mixed-script coverage (eval rows DG-011..014) +# +# The share histogram alone calls French/Spanish/Portuguese "fully covered" +# (their letters are Latin) and calls a mixed English+Han prompt covered +# whenever the Han minority is small. Both are silent mis-coverage: the +# English-only heuristics cannot read a word of them. These tests pin the +# two refinements that close the gap and the English prompts they must not +# disturb. +# --------------------------------------------------------------------------- + +FRENCH = "Crée une application web avec connexion utilisateur et tableau de bord" +SPANISH = "¿Por qué mi código se ejecuta tan lento y cómo puedo optimizarlo?" +PORTUGUESE = "Preciso de um aplicativo web com login de usuário e relatórios" +MIXED_ENGLISH_HAN = "Fix this bug 修复这个错误 in the payment flow" + + +def test_french_is_not_english_coverage(): + probe = probe_script(FRENCH) + assert probe.dominant_script == "latin" + assert probe.detected_language == "fr" + assert probe.heuristic_coverage == COVERAGE_NONE + + +def test_spanish_is_not_english_coverage(): + probe = probe_script(SPANISH) + assert probe.dominant_script == "latin" + assert probe.detected_language == "es" + assert probe.heuristic_coverage == COVERAGE_NONE + + +def test_portuguese_is_not_english_coverage(): + probe = probe_script(PORTUGUESE) + assert probe.dominant_script == "latin" + assert probe.detected_language == "pt" + assert probe.heuristic_coverage == COVERAGE_NONE + + +def test_french_degrades_even_with_only_two_stop_words(): + # Zero English function-word evidence: two distinct French function + # words already fire (_MIN_SOLO_HITS) — "Oui merci beaucoup pour ton + # aide" must not pass as English just because it is stop-word light. + probe = probe_script("Oui merci beaucoup pour ton aide") + assert probe.detected_language == "fr" + assert probe.heuristic_coverage == COVERAGE_NONE + + +def test_mixed_english_han_clause_is_degraded_by_run_length(): + probe = probe_script(MIXED_ENGLISH_HAN) + assert probe.dominant_script == "latin" + assert probe.detected_language == "en" # majority English, honestly kept + assert probe.heuristic_coverage == COVERAGE_NONE + assert probe.uncovered_block_script == "han" + assert probe.uncovered_block_length == 6 + + +def test_two_letter_han_borrow_stays_full(): + # Boundary guard for the run rule: a 2-letter borrow rides along under + # the English majority (the pinned "make it faster 这个" contract). + probe = probe_script("optimize the query 支持") + assert probe.detected_language == "en" + assert probe.heuristic_coverage == COVERAGE_FULL + + +def test_three_letter_han_block_degrades(): + # A 3-letter uncovered run is a word the English heuristics cannot read. + probe = probe_script("optimize the query 数据库") + assert probe.heuristic_coverage == COVERAGE_NONE + assert probe.uncovered_block_length == 3 + + +def test_accented_english_loanwords_stay_full(): + probe = probe_script("Add a café section for José résumé page") + assert probe.detected_language == "en" + assert probe.heuristic_coverage == COVERAGE_FULL + + +def test_urls_and_hyphenated_words_stay_full(): + # Interior punctuation keeps "example.com" and "de-duplicate" whole — + # neither may contribute a stray "com"/"de" lexicon hit. + probe = probe_script("De-duplicate records from api.example.com and deploy") + assert probe.detected_language == "en" + assert probe.heuristic_coverage == COVERAGE_FULL + + +def test_english_function_words_block_lexicon_fire(): + # Collision-heavy but genuinely English: pour/DES/LA/Mon are scattered + # collisions, and the dense English function-word evidence (the, into, + # before) keeps the margin rule from firing. + probe = probe_script("Pour the DES dump into LA storage before Mon") + assert probe.detected_language == "en" + assert probe.heuristic_coverage == COVERAGE_FULL + + +def test_unlisted_latin_language_stays_full(): + # Documented subset boundary: German is Latin script and not in the + # lexicons — it passes as before rather than pretending to degrade. + probe = probe_script("Ich möchte eine Webanwendung mit Benutzeranmeldung") + assert probe.detected_language == "en" + assert probe.heuristic_coverage == COVERAGE_FULL + + +def test_latin_language_probe_is_deterministic(): + assert probe_script(FRENCH) == probe_script(FRENCH) + assert probe_script(MIXED_ENGLISH_HAN) == probe_script(MIXED_ENGLISH_HAN) + + +def test_latin_language_degrades_end_to_end_like_other_uncovered_languages(): + # The DG-001..010 contract, now for the accented-Latin rows: explicit + # note, no silent ready, rules skipped, undetermined intent, no gaps. + for text, language in ((FRENCH, "fr"), (SPANISH, "es"), (PORTUGUESE, "pt")): + r = InputGuard().analyze(text) + assert r.heuristic_coverage == COVERAGE_NONE, text + assert r.detected_language == language, text + assert r.degradation_note is not None, text + assert r.status == "usable_with_warnings", text + assert r.detected_intent == DEGRADED_INTENT, text + assert r.gaps == [], text + assert r.clarity_score == 100 - DEGRADATION_PENALTY, text + + +def test_spanish_no_longer_returns_silent_ready(): + # DG-012's specific regression: v0.3-before-this-fix answered ready/100 + # with zero findings on Spanish input. + r = InputGuard().analyze(SPANISH) + assert r.status != "ready" + assert r.clarity_score != 100 + assert not r.is_clear() + + +def test_french_degrades_in_strict_mode_too(): + r = InputGuard(mode="strict").analyze(FRENCH) + assert r.status == "needs_clarification" + assert r.degradation_note is not None + + +def test_mixed_english_han_degrades_end_to_end(): + r = InputGuard().analyze(MIXED_ENGLISH_HAN) + assert r.heuristic_coverage == COVERAGE_NONE + assert r.degradation_note is not None + assert r.status == "usable_with_warnings" + assert r.detected_intent == DEGRADED_INTENT + assert r.gaps == [] + + +def test_french_note_names_the_language(): + note = InputGuard().analyze(FRENCH).degradation_note + assert "French" in note + assert "fr" in note + assert "English" in note + + +def test_mixed_note_names_the_uncovered_run(): + note = InputGuard().analyze(MIXED_ENGLISH_HAN).degradation_note + assert "han" in note + assert "6" in note + assert "English" in note + + +def test_all_fourteen_degradation_rows_meet_the_dg_rubric(): + # The acceptance sweep for DG-001..014 (labeling guide rubric: explicit + # degradation note, never a silent ready, no rule firing on vocabulary + # the heuristics cannot read). Labels come straight from eval/cases.csv. + import csv + from pathlib import Path + + cases = Path(__file__).resolve().parent.parent / "eval" / "cases.csv" + with cases.open(newline="", encoding="utf-8") as handle: + dg_rows = [row for row in csv.DictReader(handle) if row["id"].startswith("DG-")] + assert len(dg_rows) == 14 + + guard = InputGuard() + for row in dg_rows: + result = guard.analyze(row["text"], domain=row["domain"]) + assert result.degradation_note, row["id"] + assert result.status != "ready", row["id"] + assert result.heuristic_coverage != COVERAGE_FULL, row["id"] + assert result.gaps == [], row["id"] From 70f41ab6c38e88820e50f134a2ce098fe9fbf25c Mon Sep 17 00:00:00 2001 From: Obvious Date: Thu, 17 Sep 2026 21:25:05 +0000 Subject: [PATCH 12/18] =?UTF-8?q?fix:=20complete=20v0.3.0=20release=20swee?= =?UTF-8?q?p=20=E2=80=94=20version,=20cap,=20degradation,=20docs,=20CI?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - Bump pyproject.toml and inputguard.__version__ to 0.3.0; write the full 0.3.0 CHANGELOG section from the merged release-wave history. - Report the literal "degraded" status on the degraded path in both modes; add the mixed-script (>=5% uncovered remainder) and non-English-Latin (<10% English evidence, filler-resistant denominator) gates so French/ Spanish/Portuguese and mixed English+Han prompts degrade with zero gaps instead of spurious ones; preserve the truncated flag on that path. - Linear dataset-filename extraction (follow-ups + data-analysis rule) with 1K/10K timing tests; benchmark doc records the before/after. - README: rewritten Non-English section, full result-object field table (follow_ups, probe fields, borderline, truncated, score_breakdown), writing and data-analysis rule tables. - eval: docs/false-positive-benchmark.md from measured results. - CI: minimal GitHub Actions workflow running pytest on push/PR (3.9+3.13). Eval: 116/121 matching (all 14 degradation rows, notes on 14/14); the 5 residual mismatches are pre-existing coding boundary/intent rows. Co-authored-by: Kalisetti Nihanth Naidu Human author: Kalisetti Nihanth Naidu (nihanthnaidu007@gmail.com) --- .github/workflows/ci.yml | 21 ++++++ CHANGELOG.md | 79 +++++++++++++++----- README.md | 57 +++++++++++--- docs/false-positive-benchmark.md | 78 +++++++++++++++++++ inputguard/__init__.py | 2 +- inputguard/analyzer.py | 16 ++-- inputguard/followups.py | 35 +++++++-- inputguard/language.py | 7 ++ inputguard/matching.py | 38 +++++++++- inputguard/rules/data_analysis.py | 40 ++++++++-- pyproject.toml | 2 +- tests/test_followups.py | 63 ++++++++++++++++ tests/test_language.py | 120 ++++++++++++++++++++++++++---- 13 files changed, 495 insertions(+), 63 deletions(-) create mode 100644 .github/workflows/ci.yml create mode 100644 docs/false-positive-benchmark.md diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml new file mode 100644 index 0000000..f186827 --- /dev/null +++ b/.github/workflows/ci.yml @@ -0,0 +1,21 @@ +name: CI + +on: + push: + branches: [main] + pull_request: + +jobs: + test: + runs-on: ubuntu-latest + strategy: + matrix: + # 3.9 is the package floor (requires-python), 3.13 is current. + python-version: ["3.9", "3.13"] + steps: + - uses: actions/checkout@v4 + - uses: actions/setup-python@v5 + with: + python-version: ${{ matrix.python-version }} + - run: pip install -e ".[dev]" + - run: pytest -q diff --git a/CHANGELOG.md b/CHANGELOG.md index b16d681..3cda490 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,24 +1,67 @@ # Changelog -## [Unreleased] +## [0.3.0] — 2026-09-17 +The extension release: register your own rules and domains, get follow-up +questions for every gap, calibrated policy control, and honest behavior on +non-English input. Evaluated at **116 of 121** cases on the versioned +clarity-evaluation set in `eval/`; the 5 residual mismatches are documented +in `docs/false-positive-benchmark.md`. + ### Added -- Multilingual degradation: a zero-dependency script probe - (`unicodedata`-based histogram, `inputguard/language.py`) classifies - each input's script before rules run. Scripts without English - heuristic coverage take an explicit degraded path — rules are skipped, - a 20-point confidence penalty applies, and the result carries - `detected_language`, `heuristic_coverage`, and `degradation_note` - instead of silently scoring 100/ready. -- Additive `AnalysisResult` fields: `detected_language`, - `heuristic_coverage`, `degradation_note` (included in `to_dict()`). -- Mixed-script input: rules run whenever the covered share of letters is - at least 50%, with a `partial` note between 50–70% coverage. -- Accented-Latin honesty (eval DG-011..014): French, Spanish, and - Portuguese input no longer passes as silently covered — a function-word - layer degrades them like any other uncovered language, and a run of 3+ - consecutive uncovered-script letters (mixed English+Han) degrades even - majority-English input. English prompts (loanwords, URLs, name and - timezone collisions) are pinned unchanged by tests. +- Typed `Rule` protocol and in-process registry (`#3`): first-party and + third-party rules enter through the same path; unknown severities fail + loudly at score time instead of being swallowed. +- Registry contract enforcement (`#8`): rules must expose a `check(text)` + signature, unique `(intent, id)` pairs, and pass registration guards — + a mis-registered rule aborts instead of silently never firing. +- Per-gap follow-up questions engine (`#4`): every built-in gap carries one + or two templated clarifying questions, surfaced on the result as the + additive `follow_ups` field. `{function}` / `{dataset}` slots fill from + the original input; a gap with no table entry still gets the documented + fallback question — never silence. +- Policy calibration (`#9`): a frozen, validated `Policy` with the v0.2 + constants as defaults — status bands, severity penalties, rule filters, + `min_words`, allowlist patterns, and the 10,000-character input cap with + a visible `truncated` flag — plus a per-result `score_breakdown`. +- First-party writing domain (`#10`): six rules, the `compose` intent, + recommendations, and follow-ups for essay/report/email prompts. +- First-party data-analysis domain (`#11`): six rules (dataset/source, + question/goal, deliverable format, tooling, volume, reproducibility), + the `analysis` intent, recommendations, and follow-ups. +- The 121-case clarity-evaluation set, versioned as `eval/` (`#7`), with a + measured false-positive benchmark in `docs/false-positive-benchmark.md`. +- Multilingual degradation (`#5`, `#12`): a zero-dependency script probe + (unicodedata histogram, `inputguard/language.py`) classifies each input + before rules run. Inputs the English heuristics cannot assess take an + explicit degraded path — rules skipped, a 20-point confidence penalty, + `detected_language`, `heuristic_coverage`, `degradation_note`, and + `detected_intent: undetermined` instead of spurious gaps or a silent + 100/ready. Four inputs degrade: uncovered dominant scripts, and — from + `#12` — Latin-script text recognized as French/Spanish/Portuguese by a + function-word layer with an English margin, and Latin-dominant text + carrying a run of 3+ consecutive uncovered-script letters (mixed + English+Han). English prompts with loanwords, URLs, or name collisions + are pinned unchanged by tests. + +### Changed +- All term matching now happens at word boundaries (`#6`) — detector and + every rule module share one matcher, so "my_error" matches but "error" + inside "terrorist" no longer counts as a mention. +- Degraded results report the literal `degraded` status in both modes. + Mapping the degradation penalty through the ordinary banding returned + `usable_with_warnings` / `needs_clarification`, reading as an ordinary + vagueness verdict about the input; degradation is a language limitation + of the tool and now says so. + +### Fixed +- The dataset-filename regex behind the `{dataset}` follow-up slot (and the + data-analysis dataset rule) backtracked its greedy span against every dot + in a filename-like run — quadratic, measured at 13 ms per 1 K chars and + 1.3 s per 10 K. Both sites now scan maximal filename runs linearly with + the extension checked in Python; 10 K now takes under a millisecond, and + timing tests at 1 K and 10 K pin the growth rate. +- Degraded-path results preserve the `truncated` flag, so a capped input + that also degrades reports both honestly. ## [0.2.0] — 2026-05-29 ### Added diff --git a/README.md b/README.md index 21a4cb9..5a2b13c 100644 --- a/README.md +++ b/README.md @@ -90,6 +90,7 @@ The clarity score is mode-independent. Only the status threshold changes. | `usable_with_warnings` | score 60–84 | never | | `needs_clarification` | score < 60 | score 65–84 | | `blocked` | never | score < 65 | +| `degraded` | never (language limitation — see Non-English input) | never (language limitation) | Use `warning` when you want to surface gaps to the user without blocking. Use `strict` when you want to refuse to forward vague input to the LLM. @@ -97,12 +98,12 @@ Use `warning` when you want to surface gaps to the user without blocking. Use `s ## Non-English input -InputGuard's rules are English-language heuristics. Before any rule runs, a zero-dependency script probe (stdlib `unicodedata` only) classifies the input's script. When the script has no heuristic coverage, InputGuard says so instead of pretending: +InputGuard's rules are English-language heuristics. Before any rule runs, a zero-dependency script probe (stdlib `unicodedata` only) classifies the input's script and language. Inputs the rules cannot assess take the explicit degraded path — an uncovered dominant script, Latin-script text recognized as French/Spanish/Portuguese by its function words, or a Latin-dominant input carrying a run of 3+ consecutive uncovered-script letters. In every case InputGuard says so instead of pretending: ```python result = guard.analyze("建造一个用户登录应用") -result.status # 'usable_with_warnings' — never 'ready' +result.status # 'degraded' — never 'ready', in either mode result.clarity_score # 80 (100 minus the degradation penalty) result.detected_intent # 'undetermined' result.detected_language # 'zh' (coarse, script-derived guess) @@ -110,26 +111,29 @@ result.heuristic_coverage # 'none' result.degradation_note # explains that rules were skipped and why ``` -The rules are **skipped explicitly** on uncovered scripts — running English keyword rules on text they cannot read would produce a silent, unearned verdict. A degraded result is never `ready` in either mode (strict mode returns `needs_clarification`). In v0.2 this input silently scored 100/ready; v0.3 refuses to assert a confidence it does not have. +The rules are **skipped explicitly** — running English keyword rules on text they cannot assess would produce a silent, unearned verdict or spurious gaps invented out of the silence. A degraded result reports the literal `degraded` status in both modes: it is the tool reporting a language limitation of itself, not a judgment of the input's clarity (strict mode's banding would otherwise read as an ordinary critique). In v0.2 this input silently scored 100/ready; v0.3 refuses to assert a confidence it does not have. -Three additive fields on the result carry the probe's verdict: +Four additive fields on the result carry the probe's verdict: | Field | Values | |---|---| -| `detected_language` | coarse script-derived guess (`'en'`, `'zh'`, `'ja'`, `'ko'`, `'ru'`, `'ar'`, ...; `'und'` when unclassifiable) | +| `detected_language` | coarse script-derived guess (`'en'`, `'zh'`, `'ja'`, `'ko'`, `'ru'`, `'ar'`, `'fr'`, `'es'`, `'pt'`, ...; `'und'` when unclassifiable) | | `heuristic_coverage` | `'full'` (≥ 70% of letters covered — rules run exactly as before), `'partial'` (50–70% — rules run, note flags the uncovered remainder), `'none'` (degraded path), `'unknown'` (no letters to classify) | | `degradation_note` | `None`, or an explanation of what was skipped and why | +| `truncated` | `True` when the input was capped to the 10,000-character limit (preserved even on the degraded path) | -Mixed input is handled by share, not by exclusion: `"make it faster 这个"` is still fully analyzed (English dominates and the rules run); input whose letters fall 50–70% inside coverage gets a `partial` note without a penalty. +The degraded path covers three shapes of input the English rules cannot assess: -Latin script alone is not coverage, though: the rules are English-only, and French, Spanish, or Portuguese text is 100% Latin script yet just as unreadable to them. The probe therefore also recognizes those three languages by their function words — enough distinct stop-word hits, with a margin over the input's English function-word evidence, degrades the input like any other uncovered language. German, Italian, Dutch, and other Latin-script languages are a documented blind spot and still pass as before. And a run of 3+ consecutive letters in an uncovered script degrades even a majority-English input ("Fix this bug 修复这个错误 in the payment flow"); a 2-letter borrow like `"这个"` still rides along. +- **Uncovered script** — the dominant script (Cyrillic, Arabic, Han, ...) has no heuristic coverage. +- **Non-English Latin** — French, Spanish, and Portuguese text is 100% Latin script yet just as unreadable to English-only rules, which would find nothing and invent gaps out of the silence. A function-word layer recognizes those three languages — enough distinct stop-word hits with a margin over the input's English function-word evidence — and degrades them like any other uncovered language. German, Italian, Dutch, and other Latin-script languages are a documented blind spot and still pass as before. Accented English (`café`) and short telegraphic prompts (`build todo api`) still run the rules. +- **Mixed scripts** — a run of 3+ consecutive letters in an uncovered script degrades even a majority-English input ("Fix this bug 修复这个错误 in the payment flow"): that clause is content the rules cannot audit. A short borrow like `"make it faster 这个"` still rides along, and input at 50–70% coverage keeps the `partial` path with its note. ```python result = guard.analyze("Preciso de um aplicativo web com login de usuário e relatórios") result.detected_language # 'pt' (function-word guess) result.heuristic_coverage # 'none' -result.status # 'usable_with_warnings' — never 'ready' +result.status # 'degraded' — never 'ready' result.degradation_note # names Portuguese (pt) and explains the skip ``` @@ -217,6 +221,32 @@ Requests like "add search to my existing app", "extend my current API with pagin | `missing_feature_scope` | Feature requested but no definition of what it should specifically do. | high | | `missing_completion_criteria` | No definition of what done looks like for this feature. | low | +### Writing inputs + +Requests like "write a blog post", "draft an email", "proofread my essay". Every gap carries its own follow-up questions, and each gap names one thing at a time — same one-gap-one-question discipline as the coding domains. + +| Rule code | What it catches | Severity | +|---|---|---| +| `missing_audience` | Writing task detected but no audience or reader is specified. | high | +| `missing_purpose` | The goal — what the piece should accomplish — is not stated. | high | +| `missing_structure_format` | No length or organization guidance is provided. | medium | +| `missing_source_material` | Existing material is referenced but not provided — paste or attach the text to work from. | high | +| `missing_writing_context` | The subject or situation is not named. | medium | +| `missing_completeness` | No required content or constraints are specified. | low | + +### Data-analysis inputs + +Requests like "analyze my sales data", "build a dashboard", "report on this spreadsheet". + +| Rule code | What it catches | Severity | +|---|---|---| +| `missing_dataset_source` | Analysis requested but no dataset, file, or source is named. | high | +| `missing_question_goal` | Data present but no question or goal for the analysis is stated. | high | +| `missing_deliverable_format` | No chart, table, summary, or report named as the output. | medium | +| `missing_tooling` | No tool or stack (pandas, SQL, Excel...) specified. | medium | +| `missing_volume` | No sense of data size or scope given. | low | +| `missing_reproducibility` | No refresh/reproducibility expectation stated. | low | + --- ## The result object @@ -225,13 +255,20 @@ Requests like "add search to my existing app", "extend my current API with pagin | Field | Type | Description | |---|---|---| -| `status` | `str` | One of `"ready"`, `"usable_with_warnings"`, `"needs_clarification"`, `"blocked"` | +| `status` | `str` | One of `"ready"`, `"usable_with_warnings"`, `"needs_clarification"`, `"blocked"`, `"degraded"` (language limitation — see Non-English input) | | `clarity_score` | `int` | 0 to 100 | -| `detected_intent` | `str` | Which intent was detected: `build`, `debug`, `optimization`, `explanation`, or `feature` | +| `detected_intent` | `str` | Which intent was detected: `build`, `debug`, `optimization`, `explanation`, `feature`, `compose`, or `analysis` | | `gaps` | `List[str]` | Gap names, in the order rules fired | | `recommendations` | `List[dict]` | One dict per gap (see next section) | +| `follow_ups` | `List[str]` | One or two clarifying questions per gap, ready to send back to the user | | `findings` | `List[RuleFinding]` | Raw rule findings (code, message, severity, gap) | | `interpretation_note` | `Optional[str]` | Set when the input is highly ambiguous (score < 50 or two or more high-severity findings) | +| `detected_language` | `str` | Coarse script-derived language guess (`'en'`, `'zh'`, ...; `'und'` when unclassifiable) | +| `heuristic_coverage` | `str` | `'full'`, `'partial'`, `'none'`, or `'unknown'` — how much of the input the English rules could see | +| `degradation_note` | `Optional[str]` | Set when rule analysis was skipped or limited by language coverage | +| `borderline` | `bool` | `True` when the score sits in the near-miss band just below ready — worth one more pass | +| `truncated` | `bool` | `True` when the input was capped at 10,000 characters (only the prefix was analyzed) | +| `score_breakdown` | `Optional[dict]` | Per-contribution score arithmetic, present when the policy exposes it | Helpers: diff --git a/docs/false-positive-benchmark.md b/docs/false-positive-benchmark.md new file mode 100644 index 0000000..b2e02d4 --- /dev/null +++ b/docs/false-positive-benchmark.md @@ -0,0 +1,78 @@ +# False-positive benchmark + +Measured results of the clarity-evaluation set (`eval/`, 121 labeled rows) +against the shipped analyzer, produced by `eval/measure_fp.py`. This is the +document `eval/README.md` references as the release-wave record; re-run the +script and update the tables when the analyzer or the labels change. + +Last run: v0.3.0 release branch, 2026-09-17. + +## Headline numbers + +| Case type | Rows | Match | FP | FN | FP rate | FN rate | +| -------------- | ---- | ----- | -- | -- | ------- | ------- | +| true_positive | 43 | 42 | 0 | 1 | 0.0% | 2.3% | +| true_negative | 37 | 37 | 0 | 0 | 0.0% | 0.0% | +| boundary | 25 | 21 | 4 | 0 | 16.0% | 0.0% | +| degradation | 14 | 14 | 0 | 0 | 0.0% | 0.0% | +| performance | 2 | 2 | 0 | 0 | 0.0% | 0.0% | +| **overall** | **121** | **116** | **4** | **1** | **3.3%** | **0.8%** | + +Degradation honesty: a degradation note is present on 14 of 14 degradation +rows. Performance wall time (single pass): PF-001 22 ms, PF-002 22 ms — +PF-002 exercises the 10,000-character cap with the `truncated` flag set. + +Zero false positives on true negatives and zero on degradation rows is the +load-bearing number: it says the English keyword rules fire only on English +input, and that non-English input degrades instead of producing invented +gaps. + +## Residual mismatches (5) + +The 5 unmatched rows are all pre-existing behavior on the coding domain, +known at label time and outside the v0.3 feature work: + +- `TP-FEA-02` — false negative: `feature scope` gap not flagged; the row + comes back one status above expected (`usable_with_warnings` vs `ready`). +- `BD-002`, `BD-004`, `BD-006` — boundary rows: the detected intent + switches (build → optimization / debug) and the coding rules of the other + intent fire, adding spurious gaps. +- `BD-014` — boundary row with the same intent-adjacency shape. + +These are candidate labels to re-examine or intent-detector work for a +future release; per `eval/README.md`, labels are never edited to make a +measurement look better. + +## History within the release wave + +Measured with the same script, no label edits in between: + +| State | Match | Note | +| --------------------------------------------- | ------ | ----------------------------------------------- | +| Early v0.3 branch, before writing/analysis rules | 87/121 | 15 data-analysis rows unevaluated (rules absent) | +| After data-analysis rules landed (PR #11) | 102/121 | 14 degradation rows still mismatching | +| Degradation contract completed (this release) | 116/121 | all 14 degradation rows match, notes on 14/14 | + +## Performance: dataset-filename extraction + +The `{dataset}` follow-up slot and the data-analysis dataset rule both +extract filename-like runs. The original single regex backtracked its +greedy span against every dot in a run — quadratic. Measured on dotted +filler (the adversarial shape): + +| Input length | Old regex | Current extractor | +| ------------ | --------- | ----------------- | +| 1,000 chars | 13 ms | < 1 ms | +| 10,000 chars | 1,306 ms | 0.7 ms (plain) / 21 ms (all dots) | +| 100,000 chars | unbounded growth | 130 ms | + +Timing tests at 1 K and 10 K (`tests/test_followups.py`) pin the linear +growth rate. + +## How to reproduce + +```bash +python3 eval/measure_fp.py # full 121-row table +python3 eval/measure_fp.py --json # machine-readable +python3 -m pytest tests/ # full suite including timing guards +``` diff --git a/inputguard/__init__.py b/inputguard/__init__.py index 0ad497f..09ae2b0 100644 --- a/inputguard/__init__.py +++ b/inputguard/__init__.py @@ -3,7 +3,7 @@ from inputguard.registry import REGISTRY, Rule, register_domain, register_rule from inputguard.types import AnalysisResult, RuleFinding -__version__ = "0.2.0" +__version__ = "0.3.0" __all__ = [ # v0.2 public API — unchanged compat contract. diff --git a/inputguard/analyzer.py b/inputguard/analyzer.py index 5b68c44..ba82ec8 100644 --- a/inputguard/analyzer.py +++ b/inputguard/analyzer.py @@ -11,6 +11,7 @@ COVERAGE_PARTIAL, DEGRADATION_PENALTY, DEGRADED_INTENT, + DEGRADED_STATUS, degradation_note_for, partial_coverage_note, probe_script, @@ -93,13 +94,17 @@ def analyze( # explicit degradation path instead of silently passing as ready. probe = probe_script(normalized) if probe.heuristic_coverage == COVERAGE_NONE: - # English-only rules are skipped outright — running them on a - # script they cannot read would produce a silent, unearned - # ready. The penalty keeps the result out of "ready" in both - # modes; the note says honestly what the tool does not know. + # English-only rules are skipped outright — running them on input + # they cannot assess (another script, mixed scripts, or non- + # English Latin wording) would produce a silent, unearned ready + # or spurious gaps invented out of the silence. The status is the + # literal "degraded" in both modes: the result reports a language + # limitation of the tool, not an ordinary vagueness verdict. The + # truncation flag is preserved so a capped input degraded by the + # probe still says so. score = max(0, 100 - DEGRADATION_PENALTY) return AnalysisResult( - status=get_status(score, self.mode, effective_policy), + status=DEGRADED_STATUS, clarity_score=score, detected_intent=DEGRADED_INTENT, gaps=[], @@ -109,6 +114,7 @@ def analyze( detected_language=probe.detected_language, heuristic_coverage=probe.heuristic_coverage, degradation_note=degradation_note_for(probe), + truncated=truncated, ) detected_intent = detect_intent(user_input, domain_signals) diff --git a/inputguard/followups.py b/inputguard/followups.py index 1550371..c96c5a9 100644 --- a/inputguard/followups.py +++ b/inputguard/followups.py @@ -26,6 +26,8 @@ import re from typing import Dict, List, Mapping, Optional, Set, Tuple +from inputguard.matching import filename_with_known_extension + __all__ = ["get_follow_ups"] @@ -150,10 +152,21 @@ {"if", "for", "while", "switch", "catch", "return", "elif", "else", "do", "try"} ) -# A named data file, or a dataset/table introduced with "called"/"named". -_DATASET_FILE_RE = re.compile( - r"(? Optional[str]: def _extract_dataset(text: str) -> Optional[str]: - file_match = _DATASET_FILE_RE.search(text) + file_match = _extract_dataset_file(text) if file_match is not None: - return file_match.group(1) + return file_match named_match = _DATASET_NAMED_RE.search(text) if named_match is not None: return named_match.group(1) return None +def _extract_dataset_file(text: str) -> Optional[str]: + for run_match in _DATASET_FILE_RUN_RE.finditer(text): + filename = filename_with_known_extension( + run_match.group(0), _DATASET_FILE_EXTENSIONS + ) + if filename is not None: + return filename + return None + + def _extract_slots(text: str) -> Dict[str, str]: """Pull fillable slots from the original input (case preserved).""" slots: Dict[str, str] = {} diff --git a/inputguard/language.py b/inputguard/language.py index 38d4e3f..caabe0b 100644 --- a/inputguard/language.py +++ b/inputguard/language.py @@ -57,6 +57,7 @@ "COVERED_SCRIPTS", "DEGRADATION_PENALTY", "DEGRADED_INTENT", + "DEGRADED_STATUS", "ScriptProbe", "degradation_note_for", "partial_coverage_note", @@ -95,6 +96,12 @@ # additive intent value that says so instead of guessing the "build" fallback. DEGRADED_INTENT = "undetermined" +# The literal status a degraded result carries in both modes. Degradation +# reports a language limitation of the tool itself — mapping it through the +# ordinary banding would read as an ordinary vagueness verdict +# (usable_with_warnings / needs_clarification) about the input. +DEGRADED_STATUS = "degraded" + # The probe examines a deterministic stride sample of at most this many # characters, bounding probe cost on arbitrarily long input. _SAMPLE_LIMIT = 4096 diff --git a/inputguard/matching.py b/inputguard/matching.py index 36d3ef8..1bb5529 100644 --- a/inputguard/matching.py +++ b/inputguard/matching.py @@ -40,9 +40,9 @@ import re from functools import lru_cache -from typing import Callable, FrozenSet, Iterable, List, Tuple +from typing import Callable, FrozenSet, Iterable, List, Optional, Tuple -__all__ = ["contains_any", "contains_term"] +__all__ = ["contains_any", "contains_term", "filename_with_known_extension"] # The v0.2 coding matcher's boundary class: a neighboring identifier # character ("my_error") still counts as a mention; a longer word does not. @@ -195,3 +195,37 @@ def contains_any( f"got {terms!r}." ) return _build_finder(frozenset(terms), token_fallback)(_fold(text)) + + +def filename_with_known_extension( + run: str, extensions: FrozenSet[str] +) -> Optional[str]: + """The ``name.ext`` inside a filename-like run whose ``ext`` is known. + + ``run`` is one maximal run of filename characters (identifiers, dots, + dashes), as produced by a linear run regex; ``extensions`` is the set of + extensions the caller accepts. Returns the filename prefix (original case + preserved), or ``None`` when no dot in the run ends in a known extension + at a word boundary — ``"a.csv.dat"`` yields ``"a.csv.dat"`` (rightmost + valid dot wins), ``"file.csv.x"`` yields ``"file.csv"`` (the boundary + sits at the dot), and ``"file.csvx"`` yields ``None`` (no boundary + inside the run). + + Dots are scanned right-to-left, replicating the backtrack order of the + quadratic regexes this replaces (the rightmost valid ``name.ext`` wins), + at linear cost: each dot does O(1) work, and the run regex itself never + re-scans inside a matched run. + """ + lowered = run.lower() + for dot in range(len(run) - 1, 0, -1): + if run[dot] != "." or dot + 1 == len(run): + continue + for ext in extensions: + end = dot + 1 + len(ext) + if end > len(run) or not lowered.startswith(ext, dot + 1, end): + continue + # Word boundary after the extension: end of run, or a run char + # that is not a word char ("." or "-"). + if end == len(run) or run[end] in ".-": + return run[:end] + return None diff --git a/inputguard/rules/data_analysis.py b/inputguard/rules/data_analysis.py index 8e223cd..146e11d 100644 --- a/inputguard/rules/data_analysis.py +++ b/inputguard/rules/data_analysis.py @@ -43,7 +43,7 @@ import re from typing import FrozenSet, Optional, Pattern, Tuple -from inputguard.matching import contains_any +from inputguard.matching import contains_any, filename_with_known_extension from inputguard.registry import register_rule from inputguard.types import RuleFinding @@ -120,9 +120,19 @@ # A named data file (sales.csv, nps_verbatims_2026.xlsx) or an introduced # dataset/table ("the dataset called churn_2026"). -_DATASET_FILE_RE: Pattern[str] = re.compile( - r"(? b # -- the six gap rules --------------------------------------------------------- +def _dataset_file_present(text: str) -> bool: + """Whether any filename-like run ends in a known data extension.""" + for run_match in _DATASET_FILE_RUN_RE.finditer(text): + if ( + filename_with_known_extension( + run_match.group(0), _DATASET_FILE_EXTENSIONS + ) + is not None + ): + return True + return False + + def _check_missing_dataset_source(text: str) -> Optional[RuleFinding]: """Fires when no concrete data source is named. @@ -216,9 +239,12 @@ def _check_missing_dataset_source(text: str) -> Optional[RuleFinding]: named store (postgres, BigQuery, the warehouse, a schema). """ text = _normalize(text) - if not _is_satisfied( - text, DATASET_SATISFIED_TERMS, _DATASET_FILE_RE, _DATASET_NAMED_RE - ): + satisfied = ( + contains_any(text, DATASET_SATISFIED_TERMS) + or _dataset_file_present(text) + or _DATASET_NAMED_RE.search(text) is not None + ) + if not satisfied: return RuleFinding( code="missing_dataset_source", message=( diff --git a/pyproject.toml b/pyproject.toml index 2094fa1..93c94f3 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta" [project] name = "inputguard" -version = "0.2.0" +version = "0.3.0" description = "Catch unclear inputs before they become bad AI outputs." readme = "README.md" requires-python = ">=3.9" diff --git a/tests/test_followups.py b/tests/test_followups.py index f468deb..eaffd2b 100644 --- a/tests/test_followups.py +++ b/tests/test_followups.py @@ -15,6 +15,7 @@ from __future__ import annotations +import time from concurrent.futures import ThreadPoolExecutor import pytest @@ -22,6 +23,8 @@ from inputguard import AnalysisResult, InputGuard from inputguard.followups import ( _FOLLOW_UP_QUESTIONS, + _extract_dataset, + _extract_dataset_file, _extract_slots, _render, get_follow_ups, @@ -336,3 +339,63 @@ def test_parallel_analyze_follow_ups_stay_consistent(): expected = sequential[i] for j in range(i, len(parallel), len(sequential)): assert parallel[j] == expected + + +# --- regex linearization (regression guard against ReDoS-style backtracking) ---- + + +def _wall_ms(fn, arg: str) -> float: + start = time.perf_counter() + fn(arg) + return (time.perf_counter() - start) * 1000.0 + + +# The 10 K wall must stay within 40x of the 1 K wall. The replaced pattern +# (``run\\.(?:ext|...)\\b``) retried every offset inside a filename-shaped run +# and measured 13 ms at 1 K vs 1306 ms at 10 K — ~100x per decade, i.e. +# quadratic. A linear scan grows ~10x per decade; 40x splits the two with +# margin on both sides, so ordinary machine jitter never trips it while any +# super-linear regression blows through by orders of magnitude. +@pytest.mark.parametrize( + "extractor", + [_extract_dataset, _extract_dataset_file], + ids=["extract_dataset", "extract_dataset_file"], +) +def test_dataset_file_extraction_stays_linear_from_1k_to_10k(extractor): + def run(text: str) -> None: + extractor(text) + + small_ms = _wall_ms(run, "a" * 1_000) + large_ms = _wall_ms(run, "a" * 10_000) + assert large_ms < 40 * max(small_ms, 0.05) + # And both stay fast in absolute terms — 1.3 s at 10 K was the old number. + assert large_ms < 50.0 + + +def test_dataset_file_extraction_dotted_filler_stays_linear_from_1k_to_10k(): + def run(text: str) -> None: + _extract_dataset(text) + + small_ms = _wall_ms(run, "a." * 500) + large_ms = _wall_ms(run, "a." * 5_000) + assert large_ms < 40 * max(small_ms, 0.05) + assert large_ms < 50.0 + + +def test_dataset_file_extraction_is_case_insensitive_as_before(): + assert _extract_dataset("export SALES.CSV to me") == "SALES.CSV" + assert _extract_dataset("load Data.PARQUET now") == "Data.PARQUET" + + +def test_dataset_file_extraction_prefers_rightmost_valid_extension(): + # Within one run, the rightmost valid name.ext wins (old backtracking order). + assert _extract_dataset("archive.csv.bak") == "archive.csv" + assert _extract_dataset("archive.tar.dat") == "archive.tar.dat" + # Across runs, the first run holding a valid filename wins (old .search order). + assert _extract_dataset("archive.csv.bak and final.dat") == "archive.csv" + + +def test_dataset_file_extraction_ignores_extension_without_boundary(): + # "csvx" is not "csv" followed by a boundary — no filename here. + assert _extract_dataset("report.csvx junk") is None + diff --git a/tests/test_language.py b/tests/test_language.py index 3481bc0..5881759 100644 --- a/tests/test_language.py +++ b/tests/test_language.py @@ -24,6 +24,7 @@ COVERAGE_UNKNOWN, DEGRADATION_PENALTY, DEGRADED_INTENT, + DEGRADED_STATUS, _sample, degradation_note_for, partial_coverage_note, @@ -115,8 +116,21 @@ def test_digits_and_punctuation_have_no_coverage_signal(): assert probe.heuristic_coverage == COVERAGE_UNKNOWN -def test_majority_english_with_minority_han_is_full(): - # 12 latin letters vs 2 han letters -> covered share well above 0.7. +def test_majority_english_with_minority_han_degrades(): + # Latin dominates but a substantive uncovered-script clause rides along: + # a run of 3+ consecutive uncovered-script letters (the DG-013 shape) is + # content the rules cannot audit, so coverage is not "full". + probe = probe_script("Fix this bug 修复这个错误 in the payment flow") + assert probe.dominant_script == "latin" + assert probe.heuristic_coverage == COVERAGE_NONE + assert probe.uncovered_block_script == "han" + note = degradation_note_for(probe) + assert "mixes scripts" in note + + +def test_short_uncovered_run_rides_along_under_full_coverage(): + # Below the run threshold a stray word in another script ("这个") does + # not invalidate the analysis — the rules still audit the request. probe = probe_script("make it faster 这个") assert probe.dominant_script == "latin" assert probe.heuristic_coverage == COVERAGE_FULL @@ -179,7 +193,7 @@ def test_note_for_dominant_script_mentions_script_and_language(): def test_note_for_unrecognized_script_names_the_limitation(): probe = probe_script("ᚠᚢᚦᚨᚱᚲ") note = degradation_note_for(probe) - assert "outside the probe's coverage" in note + assert "do not cover" in note def test_partial_note_reports_coverage_percentage(): @@ -237,16 +251,18 @@ def test_degraded_result_shape(): assert r.interpretation_note is None -def test_degraded_strict_mode_returns_needs_clarification_not_blocked(): +def test_degraded_strict_mode_reports_degraded_not_blocked(): r = InputGuard(mode="strict").analyze(PROBE_P2_INPUT) - # Score 80 lands in the strict clarify band (65-84), below the ready - # floor in both modes. - assert r.status == "needs_clarification" + # Degradation is the tool reporting a language limitation of itself, not + # a vagueness verdict — the literal "degraded" status in both modes + # (banding would have said "needs_clarification", reading like an + # ordinary critique of the input). + assert r.status == "degraded" -def test_degraded_warning_mode_returns_usable_with_warnings(): +def test_degraded_warning_mode_reports_degraded_not_usable_with_warnings(): r = InputGuard().analyze(PROBE_P2_INPUT) - assert r.status == "usable_with_warnings" + assert r.status == "degraded" def test_degradation_note_is_actionable(): @@ -255,6 +271,82 @@ def test_degradation_note_is_actionable(): assert "han" in r.degradation_note +# --- Latin-script degradation flavors (the DG-011/012/014 contract) ---------- +# These mirror eval rows verbatim: French, Spanish, and Portuguese prompts are +# Latin script, so script coverage alone said "full" and the English keyword +# rules invented spurious gaps out of the silence. The English-evidence gate +# degrades them instead. + + +def test_french_prompt_degrades_with_no_gaps(): + # DG-011, verbatim from eval/cases.csv. + r = InputGuard().analyze( + "Crée une application web avec connexion utilisateur et affiche les " + "données sous forme de tableau." + ) + assert r.status == "degraded" + assert r.heuristic_coverage == COVERAGE_NONE + assert r.detected_intent == DEGRADED_INTENT + assert r.gaps == [] and r.findings == [] + assert r.clarity_score == 100 - DEGRADATION_PENALTY + assert "French" in r.degradation_note + + +def test_spanish_prompt_degrades_with_no_gaps(): + # DG-012, verbatim from eval/cases.csv. + r = InputGuard().analyze( + "¿por qué mi código se ejecuta tan lento y cómo puedo optimizarlo?" + ) + assert r.status == "degraded" + assert r.detected_language == "es" + assert r.gaps == [] and r.findings == [] + + +def test_portuguese_prompt_degrades_with_no_gaps(): + # DG-014, verbatim from eval/cases.csv. "no" (Portuguese "in the") is a + # stray English stop word — one collision out of 13 words stays under the + # evidence threshold. + r = InputGuard().analyze( + "Preciso de um aplicativo web com login de usuário e relatórios no " + "painel." + ) + assert r.status == "degraded" + assert r.gaps == [] and r.findings == [] + + +def test_mixed_english_han_prompt_degrades_instead_of_reporting_gaps(): + # DG-013, verbatim from eval/cases.csv: the Han clause carries the actual + # request; running the rules on the English part alone reported spurious + # gaps invented out of the unaudited remainder. + r = InputGuard().analyze("Fix this bug 修复这个错误 in the payment flow") + assert r.status == "degraded" + assert r.gaps == [] and r.findings == [] + + +def test_accented_english_still_runs_the_rules(): + # Accented English words must not trip the language gate: function/action + # words still carry the evidence even when é splits an accented token. + r = InputGuard().analyze( + "explain how the café module should handle a timeout error" + ) + assert r.status in ("ready", "usable_with_warnings", "needs_clarification") + assert r.heuristic_coverage == COVERAGE_FULL + + +def test_telegraphic_english_still_runs_the_rules(): + # Below five words the gate cannot judge a language; short prompts with + # no stop words stay fully analyzable. + r = InputGuard().analyze("build todo api") + assert r.heuristic_coverage == COVERAGE_FULL + + +def test_degraded_result_preserves_truncation_flag(): + # A capped input that the probe degrades must still report truncation. + r = InputGuard().analyze("修复这个错误 " * 2500) # > 10_000 chars, Chinese + assert r.status == "degraded" + assert r.truncated is True + + def test_degraded_applies_to_every_uncovered_script(): for text in ( "почему моя программа не работает", @@ -347,7 +439,7 @@ def test_to_dict_includes_additive_fields(): # hypothesis is not a dev dependency; zero-dep constraint honored) # --------------------------------------------------------------------------- -_VALID_STATUSES = {"ready", "usable_with_warnings", "needs_clarification", "blocked"} +_VALID_STATUSES = {"ready", "usable_with_warnings", "needs_clarification", "blocked", "degraded"} _UNICODE_SAMPLES = [ "emoji only 🚀🔥🏳️‍🌈", @@ -544,7 +636,7 @@ def test_latin_language_degrades_end_to_end_like_other_uncovered_languages(): assert r.heuristic_coverage == COVERAGE_NONE, text assert r.detected_language == language, text assert r.degradation_note is not None, text - assert r.status == "usable_with_warnings", text + assert r.status == DEGRADED_STATUS, text assert r.detected_intent == DEGRADED_INTENT, text assert r.gaps == [], text assert r.clarity_score == 100 - DEGRADATION_PENALTY, text @@ -561,7 +653,9 @@ def test_spanish_no_longer_returns_silent_ready(): def test_french_degrades_in_strict_mode_too(): r = InputGuard(mode="strict").analyze(FRENCH) - assert r.status == "needs_clarification" + # Literal "degraded" in both modes — banding would read as an ordinary + # critique of the input. + assert r.status == "degraded" assert r.degradation_note is not None @@ -569,7 +663,7 @@ def test_mixed_english_han_degrades_end_to_end(): r = InputGuard().analyze(MIXED_ENGLISH_HAN) assert r.heuristic_coverage == COVERAGE_NONE assert r.degradation_note is not None - assert r.status == "usable_with_warnings" + assert r.status == DEGRADED_STATUS assert r.detected_intent == DEGRADED_INTENT assert r.gaps == [] From 850401f9b62443f70e0cc1365420b52d6e5b0109 Mon Sep 17 00:00:00 2001 From: "obvious-autobuild[bot]" <262744130+obvious-autobuild[bot]@users.noreply.github.com> Date: Thu, 17 Sep 2026 21:42:29 +0000 Subject: [PATCH 13/18] feat: CI matrix, latency budgets, CLI, and packaging metadata for v0.3 (#13) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * test(registry-contract-guards): CI-enforced guards for the three review findings Dedicated suite for the adversarial-API-review gaps (art_hC18m78C), as the named acceptance artifact: - C3/P2: duplicate rule ids inside one register_domain call raise ValueError naming the id; a rejected batch mutates nothing. - B1/N3/P1: a rule whose check cannot be called as check(self, text) (extra param, zero param, keyword-only) is rejected at registration with an actionable message; defaulted and spec signatures accepted and fired. - C2/P4: an exception inside check() aborts analyze() as a RuntimeError carrying the rule id, its registration origin, and the original exception chained -- via the register_domain path (register_rule is covered in test_registry.py). The registry-isolation fixture moves to tests/conftest.py so the guard suite and the policy suite share one snapshot/restore implementation instead of importing across test modules. Co-authored-by: Kalisetti Nihanth Naidu * build(typing): annotate for mypy --strict without behavior change Pre-work for the CI strict-typing gate: missing Iterable[str] on the five rule modules' _contains_any helpers, Dict[str, Any] / Dict[str, str] generics on types.py and recommender.py, a pinned Rule-typed local in registry._instantiate, and the writing signals Tuple[str, ...] type argument. Also removes a literal duplicate "filter by"/"sort by" set entry in the feature scope signals (ruff B033); set semantics make this a no-op at runtime, and rules/__init__ marks its v0.2 signal re-exports as intentional via __all__ (ruff F401). Co-authored-by: Kalisetti Nihanth Naidu * ci(actions-matrix): GitHub Actions CI with lint, typed, tested, benchmark, wheel gates The v0.2 failure mode was "every change ships on trust" — no CI, lint, typing, or coverage tooling at all. This adds the production gate set from the spec (§8) as three jobs: - test: Python 3.9-3.13 matrix running ruff, pytest with a 90% branch-coverage floor (--cov-branch --cov-fail-under=90 --strict), and mypy --strict (finally checking the shipped py.typed). - benchmark: latency budgets as merge gates, not notes. Budgets are derived from measured numbers (build-intent p50 21.94 ms at 10k chars including the C1 InsufficientContextRule double-run, debug 6.59 ms, ready 5.62 ms; 1.6 MB input bounded by the Policy cap at 17.1 ms vs the v0.2 unbounded 3.8 s scan). Hard gate is p50 with CI-noise headroom; p50/p99 print to the job log per spec. Budgets and derivation documented in tests/test_benchmarks.py. - wheel-zero-dep: the zero-dependency promise verified against the built artifact — wheel METADATA carries no Requires-Dist, a bare venv install pulls no third-party package, and the public API works from the installed wheel. Smoke steps run from a scratch directory so the repo's local inputguard/ package can never shadow the installed wheel. pyproject gains the matching tool config ([tool.ruff] E/F/W/B with E501 ignored, [tool.mypy] strict, explicit pytest testpaths) and the expanded [dev] extras the CI jobs install; all dev tooling stays out of runtime dependencies. Co-authored-by: Kalisetti Nihanth Naidu * feat(cli): inputguard analyze with argparse, to_dict JSON, CI-friendly exit codes The developer surface the adoption research flagged (spec §5): a zero-dependency argparse CLI shipped as a console script. - "inputguard analyze TEXT [--stdin] [--domain D] [--mode M] [--format text|json] [--min-score N]"; JSON output is the result's own to_dict() contract, byte-equal to the library's (asserted by test). - Exit codes documented in the module docstring and --help: 0 = analysis ok and at/above the --min-score floor when given; 1 = below the floor (the commit-hook / CI gate); 2 = usage error (missing text, unknown domain) reported cleanly on stderr, never a traceback. - Text rendering follows the spec §5 example shape: status, clarity, intent, domain, missing gaps, ask lines, and the degradation note when the language probe marks the result degraded (probe P2 input shows the note, never a silent ready). Entry point added in pyproject [project.scripts]; stdlib argparse and json only — the zero-dependency promise is untouched. Co-authored-by: Kalisetti Nihanth Naidu * chore(packaging-metadata): 0.3.0, Beta status, 3.13 classifier, project URLs The packaging drift the survey flagged is closed: classifiers extend to Python 3.13 (the runtime the project is developed and tested on), Development Status moves to 4 - Beta per the release plan (1.0 deliberately deferred until the extension API sees real use), [project.urls] points at the repository, changelog, and issue tracker, and inputguard.__version__ syncs with the pyproject version. The changelog's Unreleased v0.3 section is promoted to [0.3.0] with the CI and CLI entries from this PR. Also amends the wheel-gate assertion to what the zero-dependency promise actually means: no UNCONDITIONAL Requires-Dist entries (setuptools emits extra-gated `Requires-Dist: ...; extra == "dev"` lines for the optional tooling — verified against the built wheel, and the bare-venv install proves none of them install at runtime). Co-authored-by: Kalisetti Nihanth Naidu * fix(benchmark-budgets): gate latency only in the untraced benchmark job, budgets from measured CI data The first CI run (35277175054) failed the latency asserts inside the coverage pass: coverage tracing roughly triples per-call cost (debug p50 21.14 ms vs 6.59 ms untraced; ready 20.35 ms vs 5.62 ms), and runner hardware is slower than the dev sandbox. Gating the same budgets under coverage in the test job AND untraced in the benchmark job was a design error — two different measurement conditions cannot share one threshold. - The coverage pass now ignores tests/test_benchmarks.py (with the observed numbers in a comment); latency gates belong exclusively to the dedicated benchmark job, which runs untraced and reflects real-world analysis cost. - Budgets re-derived from measured environments documented in the test module: sandbox untraced (build 21.94 / debug 6.59 / ready 5.62 ms p50) and the CI-traced observation above. New p50 budgets: build 65 ms (~3x), debug 25 ms (~3.8x), ready 25 ms (~4.4x); p95 sustained guards at 2.5x unchanged. Still order-of-magnitude gates: they catch algorithmic regressions (accidental O(n^2), a lost cap) that blow past by 10x+, while absorbing CI hardware variance. Local verification of the exact split: coverage pass 441 passed, 96.29% branch (floor 90); benchmark job untraced 4 passed; ruff + mypy --strict clean. Co-authored-by: Kalisetti Nihanth Naidu * fix: give the extraction linearization bounds CI headroom test (3.10) flaked on the rebased head (run 35277834661): the dotted-filler linearization test's absolute bound asserted a single wall-clock sample under 50 ms, and the coverage-traced 3.10 runner measured 50.61 ms. The meaningful guard is the 40x growth-rate assertion — the regression it protects against measured 1.3 s at 10 K, ~26x the bound. Both timing tests now use a 250 ms absolute bound (5x+ detection margin, no CI-hardware flake); the growth-rate assertion is unchanged. Co-authored-by: Kalisetti Nihanth Naidu --------- Co-authored-by: Obvious Co-authored-by: Kalisetti Nihanth Naidu --- .github/workflows/ci.yml | 138 ++++++++++++- CHANGELOG.md | 17 ++ inputguard/cli.py | 134 +++++++++++++ inputguard/detector.py | 2 +- inputguard/language.py | 6 +- inputguard/recommender.py | 8 +- inputguard/registry.py | 5 +- inputguard/rules/__init__.py | 8 + inputguard/rules/coding.py | 4 +- inputguard/rules/debug.py | 4 +- inputguard/rules/explanation.py | 4 +- inputguard/rules/feature.py | 6 +- inputguard/rules/optimization.py | 4 +- inputguard/rules/writing.py | 2 +- inputguard/types.py | 4 +- pyproject.toml | 39 +++- tests/conftest.py | 33 +++ tests/test_benchmarks.py | 165 +++++++++++++++ tests/test_cli.py | 83 ++++++++ tests/test_followups.py | 15 +- tests/test_policy.py | 3 +- tests/test_registry_contract_guards.py | 268 +++++++++++++++++++++++++ 22 files changed, 917 insertions(+), 35 deletions(-) create mode 100644 inputguard/cli.py create mode 100644 tests/conftest.py create mode 100644 tests/test_benchmarks.py create mode 100644 tests/test_cli.py create mode 100644 tests/test_registry_contract_guards.py diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index f186827..a39597e 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -2,20 +2,148 @@ name: CI on: push: - branches: [main] + branches: [main, "release/**"] pull_request: +permissions: + contents: read + jobs: + # The quality gate: lint, tests with a 90% branch-coverage floor, and + # strict typing across the full supported Python matrix (spec §8). test: runs-on: ubuntu-latest strategy: + fail-fast: false matrix: - # 3.9 is the package floor (requires-python), 3.13 is current. - python-version: ["3.9", "3.13"] + python-version: ["3.9", "3.10", "3.11", "3.12", "3.13"] steps: - uses: actions/checkout@v4 - uses: actions/setup-python@v5 with: python-version: ${{ matrix.python-version }} - - run: pip install -e ".[dev]" - - run: pytest -q + - name: Install with dev tooling + run: | + python -m pip install --upgrade pip + pip install -e ".[dev]" + - name: Ruff + run: ruff check inputguard tests eval + - name: Run tests with 90% branch-coverage floor + # Latency asserts are excluded here and gate only in the dedicated + # benchmark job: coverage tracing roughly triples per-call cost + # (observed on run 35277175054 — debug p50 21.14 ms vs 6.59 ms + # untraced), so gating latency under coverage double-gates the same + # budgets under different conditions. + run: | + pytest -q --cov=inputguard --cov-branch --cov-fail-under=90 --strict \ + --ignore=tests/test_benchmarks.py + - name: Strict typing (checks the shipped py.typed) + run: mypy --strict inputguard + + # Latency budgets are gates, not notes: a regression here is a production + # failure mode (the v0.2 unbounded 3.8 s scan). Measured p50/p99 print to + # the job log; budgets and their derivation live in tests/test_benchmarks.py. + benchmark: + runs-on: ubuntu-latest + steps: + - uses: actions/checkout@v4 + - uses: actions/setup-python@v5 + with: + python-version: "3.13" + - name: Install with dev tooling + run: | + python -m pip install --upgrade pip + pip install -e ".[dev]" + - name: Latency budgets (p50/p99 recorded) + run: pytest -v -s tests/test_benchmarks.py + + # The zero-dependency promise, verified against the built artifact — not + # the source tree: wheel METADATA carries no Requires-Dist, a bare venv + # install pulls no third-party package, and the public API + CLI work from + # the installed wheel alone. The smoke steps run from a scratch directory + # so the repo's local `inputguard/` package can never shadow the wheel. + wheel-zero-dep: + runs-on: ubuntu-latest + steps: + - uses: actions/checkout@v4 + - uses: actions/setup-python@v5 + with: + python-version: "3.13" + - name: Build wheel + run: | + python -m pip install --upgrade pip build + python -m build --wheel + - name: Wheel metadata declares no unconditional runtime dependencies + run: | + python - <<'PY' + import glob + import zipfile + + wheel = glob.glob("dist/*.whl")[0] + names = zipfile.ZipFile(wheel).namelist() + meta_name = next(n for n in names if n.endswith(".dist-info/METADATA")) + meta = zipfile.ZipFile(wheel).read(meta_name).decode() + requires = [ + line + for line in meta.splitlines() + if line.startswith("Requires-Dist:") + ] + unconditional = [ + line for line in requires if 'extra == "' not in line + ] + assert not unconditional, ( + "wheel METADATA declares unconditional runtime dependencies — " + f"the zero-dep promise is broken: {unconditional}" + ) + extra_gated = len(requires) + print( + f"{wheel}: 0 unconditional runtime dependencies " + f"({extra_gated} extra-gated dev-only entries)" + ) + PY + - name: Install wheel into a bare venv + run: | + python -m venv bare + ./bare/bin/pip install --upgrade pip + ./bare/bin/pip install dist/*.whl + - name: Public API works from the installed wheel + run: | + cd "$(mktemp -d)" + "$GITHUB_WORKSPACE/bare/bin/python" - <<'PY' + import json + + import inputguard + + assert inputguard.__version__ + guard = inputguard.InputGuard() + result = guard.analyze("make this faster", domain="coding") + assert isinstance(result.clarity_score, int) + assert 0 <= result.clarity_score <= 100 + payload = result.to_dict() + json.dumps(payload) + print(f"bare-venv analyze OK: {result.status} score={result.clarity_score}") + PY + - name: Bare venv contains no third-party packages + run: | + third_party=$(./bare/bin/pip list --format=freeze \ + | cut -d= -f1 | tr '[:upper:]' '[:lower:]' \ + | grep -v -E '^(inputguard|pip|setuptools)$' || true) + if [ -n "$third_party" ]; then + echo "Third-party packages present after wheel install: $third_party" + exit 1 + fi + echo "bare venv holds only inputguard (+ venv tooling)" + - name: CLI works from the installed wheel + run: | + cd "$(mktemp -d)" + INPUTGUARD="$GITHUB_WORKSPACE/bare/bin/inputguard" + "$INPUTGUARD" analyze "make this faster" --format json + set +e + "$INPUTGUARD" analyze "make this faster" --min-score 85 + code=$? + set -e + if [ "$code" -ne 1 ]; then + echo "expected exit 1 below the min-score floor, got $code" + exit 1 + fi + echo "CLI exit codes OK (0 on analysis, 1 below the min-score floor)" diff --git a/CHANGELOG.md b/CHANGELOG.md index 3cda490..c67387a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -42,6 +42,23 @@ in `docs/false-positive-benchmark.md`. carrying a run of 3+ consecutive uncovered-script letters (mixed English+Han). English prompts with loanwords, URLs, or name collisions are pinned unchanged by tests. +- CI as production gates (`.github/workflows/ci.yml`): a Python 3.9–3.13 + matrix running ruff, pytest with a 90% branch-coverage floor + (`--cov-branch --cov-fail-under=90 --strict`, latency asserts excluded + from the traced pass), and mypy `--strict` (the shipped `py.typed` is + now actually checked); a dedicated untraced latency-benchmark job with + budgets measured on this codebase (10k-char build-intent p50 ≤ 65 ms + including the catch-all re-run path, debug and ready paths ≤ 25 ms; a + 1.6 MB input bounded by the policy cap); and a zero-dependency wheel + gate that builds the wheel, installs it into a bare venv, and asserts + no third-party package is present. +- `inputguard` CLI (`inputguard analyze`) via a console-script entry + point: argparse, `--format json` emitting the same `to_dict()` + contract, and CI-friendly exit codes (0 ok, 1 below `--min-score`, + 2 usage error). +- Packaging metadata: 3.13 classifier, Development Status → 4 - Beta, + project URLs, and the explicit dev extras / tool config + (`[tool.ruff]`, `[tool.mypy]`, `[tool.pytest.ini_options]`). ### Changed - All term matching now happens at word boundaries (`#6`) — detector and diff --git a/inputguard/cli.py b/inputguard/cli.py new file mode 100644 index 0000000..0565439 --- /dev/null +++ b/inputguard/cli.py @@ -0,0 +1,134 @@ +"""The ``inputguard`` command-line interface (spec art_bTvdPdJS §5). + +A zero-dependency argparse front end over the same ``InputGuard.analyze`` +pipeline the library exposes. JSON output is the result's own ``to_dict()`` +contract — the CLI adds no fields and renames none. + +Exit codes (the CI/commit-hook contract): + +- ``0`` — analysis completed; with ``--min-score N``, the score is at or + above the floor. +- ``1`` — analysis completed but the clarity score is below the + ``--min-score`` floor. Use with ``--min-score`` to gate commits/PRs. +- ``2`` — usage error: unknown flags, missing text, or an invalid value + (unknown domain, empty input) reported by the pipeline. + +Examples:: + + inputguard analyze "make this faster" + inputguard analyze "make this faster" --mode strict + inputguard analyze "$(cat prompt.txt)" --format json --min-score 85 + cat prompt.txt | inputguard analyze --stdin --min-score 85 +""" + +from __future__ import annotations + +import argparse +import json +import sys +from typing import Optional, Sequence + +from inputguard import AnalysisResult, InputGuard + +EXIT_OK = 0 +EXIT_BELOW_FLOOR = 1 +EXIT_USAGE = 2 + + +def _build_parser() -> argparse.ArgumentParser: + parser = argparse.ArgumentParser( + prog="inputguard", + description="Pre-flight LLM inputs for clarity before inference.", + ) + subparsers = parser.add_subparsers(dest="command", required=True) + analyze = subparsers.add_parser( + "analyze", + help="Analyze input text and report its clarity verdict.", + ) + analyze.add_argument( + "text", + nargs="?", + help="The input text to analyze; omit when --stdin is given.", + ) + analyze.add_argument( + "--stdin", + action="store_true", + help="Read the input text from stdin instead of the TEXT argument.", + ) + analyze.add_argument( + "--domain", + default="coding", + help="Registered domain to analyze against (default: coding).", + ) + analyze.add_argument( + "--mode", + choices=("warning", "strict"), + default="warning", + help="Status banding mode (default: warning).", + ) + analyze.add_argument( + "--format", + choices=("text", "json"), + default="text", + help="Output format (default: text). json emits the to_dict() contract.", + ) + analyze.add_argument( + "--min-score", + type=int, + default=None, + metavar="N", + help=( + "Exit 1 when the clarity score is below N — a ready-made " + "commit-hook / CI gate." + ), + ) + return parser + + +def _render_text(result: AnalysisResult, domain: str) -> str: + """Human-readable verdict, one fact per line (spec §5 example shape).""" + lines = [ + f"status: {result.status}", + f"clarity: {result.clarity_score}/100", + f"intent: {result.detected_intent}", + f"domain: {domain}", + ] + if result.gaps: + lines.append(f"missing: {', '.join(result.gaps)}") + for question in result.follow_ups: + lines.append(f"ask: {question}") + if result.degradation_note is not None: + lines.append(f"note: {result.degradation_note}") + return "\n".join(lines) + + +def main(argv: Optional[Sequence[str]] = None) -> int: + """Run one analysis; returns the process exit code (see module docstring).""" + parser = _build_parser() + args = parser.parse_args(argv) + + text = sys.stdin.read() if args.stdin else args.text + if not text: + parser.error("provide the text to analyze as TEXT, or pass --stdin") + + guard = InputGuard(mode=args.mode) + try: + result = guard.analyze(text, domain=args.domain) + except ValueError as exc: + # Invalid domain / empty-after-normalization input: a usage error, + # not a crash — the same ValueError contract the library documents. + print(f"inputguard: {exc}", file=sys.stderr) + return EXIT_USAGE + + if args.format == "json": + print(json.dumps(result.to_dict(), indent=2)) + else: + print(_render_text(result, args.domain)) + + if args.min_score is not None and result.clarity_score < args.min_score: + return EXIT_BELOW_FLOOR + return EXIT_OK + + +if __name__ == "__main__": # pragma: no cover — manual invocation convenience + sys.exit(main()) diff --git a/inputguard/detector.py b/inputguard/detector.py index a86fe96..bb62444 100644 --- a/inputguard/detector.py +++ b/inputguard/detector.py @@ -70,7 +70,7 @@ def normalize(text: str) -> str: return re.sub(r"\s+", " ", text.strip().lower()) -def _contains_any(text: str, terms) -> bool: +def _contains_any(text: str, terms: Iterable[str]) -> bool: # v0.3: word-boundary matching shared with the rule modules — "fixture" # is no longer read as the debug signal "fix" (probe P1). return contains_any(text, terms) diff --git a/inputguard/language.py b/inputguard/language.py index caabe0b..ca9ddfe 100644 --- a/inputguard/language.py +++ b/inputguard/language.py @@ -47,7 +47,7 @@ import re import unicodedata from dataclasses import dataclass -from typing import Dict, Optional +from typing import Dict, FrozenSet, Optional, Set __all__ = [ "COVERAGE_FULL", @@ -228,7 +228,7 @@ # of 3+ letters is a word or clause the English heuristics cannot read. _UNCOVERED_BLOCK_DEGRADES_AT = 3 -_LATIN_FUNCTION_WORDS: Dict[str, frozenset] = { +_LATIN_FUNCTION_WORDS: Dict[str, FrozenSet[str]] = { "en": frozenset({ "a", "an", "the", "and", "or", "but", "if", "then", "of", "to", "in", "on", "at", "by", "for", "with", "from", "into", "over", "under", @@ -357,7 +357,7 @@ def _latin_language_of(sample: str) -> Optional[str]: Ties between firing languages resolve alphabetically, like the script histogram's dominant-script tie-break. """ - distinct: Dict[str, set] = {} + distinct: Dict[str, Set[str]] = {} for raw in sample.translate(_APOSTROPHES).lower().split(): token = _EDGE_TRIM.sub("", raw) if not token: diff --git a/inputguard/recommender.py b/inputguard/recommender.py index b36b7a4..a5f8ed3 100644 --- a/inputguard/recommender.py +++ b/inputguard/recommender.py @@ -3,7 +3,7 @@ from typing import Dict, List -_RECOMMENDATIONS: Dict[str, dict] = { +_RECOMMENDATIONS: Dict[str, Dict[str, str]] = { "programming language": { "gap": "programming language", "what_is_missing": "You haven't told it which programming language or technology to use.", @@ -183,7 +183,7 @@ } -def _fallback_recommendation(gap: str) -> dict: +def _fallback_recommendation(gap: str) -> Dict[str, str]: """Generic four-key advice for a gap with no curated entry. The documented fallback (spec art_bTvdPdJS §6): unknown gaps keep their @@ -198,13 +198,13 @@ def _fallback_recommendation(gap: str) -> dict: } -def get_recommendations(gaps: List[str]) -> List[dict]: +def get_recommendations(gaps: List[str]) -> List[Dict[str, str]]: """Build one four-key recommendation per gap, in input order. A gap without a curated entry gets the documented fallback (see :func:`_fallback_recommendation`) — never a silent drop. """ - out: List[dict] = [] + out: List[Dict[str, str]] = [] for gap in gaps: entry = _RECOMMENDATIONS.get(gap) out.append(dict(entry) if entry is not None else _fallback_recommendation(gap)) diff --git a/inputguard/registry.py b/inputguard/registry.py index 6261bf2..dd27943 100644 --- a/inputguard/registry.py +++ b/inputguard/registry.py @@ -121,7 +121,10 @@ def check(self, text: str) -> Optional[RuleFinding]: def _instantiate(rule_cls: type) -> Rule: try: - return rule_cls() + # Bare ``type`` construction is typed Any; the annotation pins the + # protocol so strict mode sees a Rule, not Any. + instance: Rule = rule_cls() + return instance except TypeError as exc: raise TypeError( f"Cannot register rule class {rule_cls.__name__!r}: it must be " diff --git a/inputguard/rules/__init__.py b/inputguard/rules/__init__.py index 0a6f434..c5f350b 100644 --- a/inputguard/rules/__init__.py +++ b/inputguard/rules/__init__.py @@ -36,6 +36,14 @@ "run_optimization_rules", "run_explanation_rules", "run_feature_rules", + # Backward-compat re-exports of the v0.2 signal tables (INTENT_SIGNALS is + # consumed by the coding registration above; the per-intent tables stay + # importable for code that read them from this module in v0.2). + "INTENT_SIGNALS", + "DEBUG_SIGNALS", + "EXPLANATION_SIGNALS", + "FEATURE_SIGNALS", + "OPTIMIZATION_SIGNALS", ] # The coding domain: intent signals in strict priority order (the single diff --git a/inputguard/rules/coding.py b/inputguard/rules/coding.py index 2b3b9ff..75b5369 100644 --- a/inputguard/rules/coding.py +++ b/inputguard/rules/coding.py @@ -1,7 +1,7 @@ from __future__ import annotations import re -from typing import List, Optional, Set +from typing import Iterable, List, Optional, Set from inputguard.matching import contains_any from inputguard.registry import register_rule @@ -111,7 +111,7 @@ def _normalize(text: str) -> str: return re.sub(r"\s+", " ", text.lower()).strip() -def _contains_any(text: str, terms) -> bool: +def _contains_any(text: str, terms: Iterable[str]) -> bool: # v0.3: word-boundary matching via the shared matcher. token_fallback # keeps the v0.2 coding-rule behavior for multiword terms ("def ", # "sign in with"): their words may appear non-adjacent. diff --git a/inputguard/rules/debug.py b/inputguard/rules/debug.py index 7431f15..cd79ec7 100644 --- a/inputguard/rules/debug.py +++ b/inputguard/rules/debug.py @@ -1,7 +1,7 @@ from __future__ import annotations import re -from typing import List, Optional +from typing import Iterable, List, Optional from inputguard.detector import DEBUG_SIGNALS from inputguard.matching import contains_any @@ -43,7 +43,7 @@ def _normalize(text: str) -> str: return re.sub(r"\s+", " ", text.strip().lower()) -def _contains_any(text: str, terms) -> bool: +def _contains_any(text: str, terms: Iterable[str]) -> bool: # v0.3: word-boundary matching via the shared matcher (probe P1 fix). return contains_any(text, terms) diff --git a/inputguard/rules/explanation.py b/inputguard/rules/explanation.py index 69498de..9610b5e 100644 --- a/inputguard/rules/explanation.py +++ b/inputguard/rules/explanation.py @@ -1,7 +1,7 @@ from __future__ import annotations import re -from typing import List, Optional +from typing import Iterable, List, Optional from inputguard.detector import EXPLANATION_SIGNALS from inputguard.matching import contains_any @@ -37,7 +37,7 @@ def _normalize(text: str) -> str: return re.sub(r"\s+", " ", text.strip().lower()) -def _contains_any(text: str, terms) -> bool: +def _contains_any(text: str, terms: Iterable[str]) -> bool: # v0.3: word-boundary matching via the shared matcher (probe P1 fix). return contains_any(text, terms) diff --git a/inputguard/rules/feature.py b/inputguard/rules/feature.py index 51b6948..891653d 100644 --- a/inputguard/rules/feature.py +++ b/inputguard/rules/feature.py @@ -1,7 +1,7 @@ from __future__ import annotations import re -from typing import List, Optional +from typing import Iterable, List, Optional from inputguard.detector import FEATURE_SIGNALS from inputguard.matching import contains_any @@ -27,7 +27,7 @@ "oauth", "jwt", "session", "api key", "role-based", "rbac", "upload", "download", "preview", "thumbnail", "dashboard", "chart", "graph", "table", "export", - "search by", "filter by", "sort by", "group by", + "search by", "group by", } COMPLETION_CRITERIA_SIGNALS = { @@ -44,7 +44,7 @@ def _normalize(text: str) -> str: return re.sub(r"\s+", " ", text.strip().lower()) -def _contains_any(text: str, terms) -> bool: +def _contains_any(text: str, terms: Iterable[str]) -> bool: # v0.3: word-boundary matching via the shared matcher (probe P1 fix). return contains_any(text, terms) diff --git a/inputguard/rules/optimization.py b/inputguard/rules/optimization.py index fdfc587..e6eda25 100644 --- a/inputguard/rules/optimization.py +++ b/inputguard/rules/optimization.py @@ -1,7 +1,7 @@ from __future__ import annotations import re -from typing import List, Optional +from typing import Iterable, List, Optional from inputguard.detector import OPTIMIZATION_SIGNALS from inputguard.matching import contains_any @@ -43,7 +43,7 @@ def _normalize(text: str) -> str: return re.sub(r"\s+", " ", text.strip().lower()) -def _contains_any(text: str, terms) -> bool: +def _contains_any(text: str, terms: Iterable[str]) -> bool: # v0.3: word-boundary matching via the shared matcher (probe P1 fix). return contains_any(text, terms) diff --git a/inputguard/rules/writing.py b/inputguard/rules/writing.py index a0b8007..87a992e 100644 --- a/inputguard/rules/writing.py +++ b/inputguard/rules/writing.py @@ -397,7 +397,7 @@ def check(self, text: str) -> Optional[RuleFinding]: # The single fallback intent: every writing-domain input is a composition # task. Exactly one empty-terms intent is what the registry requires. -WRITING_SIGNALS: Tuple[Tuple[str, tuple], ...] = ( +WRITING_SIGNALS: Tuple[Tuple[str, Tuple[str, ...]], ...] = ( ("compose", ()), ) diff --git a/inputguard/types.py b/inputguard/types.py index 73bde72..85316c5 100644 --- a/inputguard/types.py +++ b/inputguard/types.py @@ -27,7 +27,7 @@ class AnalysisResult: clarity_score: int detected_intent: str gaps: List[str] = field(default_factory=list) - recommendations: List[dict] = field(default_factory=list) + recommendations: List[Dict[str, Any]] = field(default_factory=list) findings: List[RuleFinding] = field(default_factory=list) interpretation_note: Optional[str] = None # v0.3, additive: templated clarifying questions, one or two per gap, @@ -48,7 +48,7 @@ class AnalysisResult: truncated: bool = False score_breakdown: Optional[Dict[str, Any]] = None - def to_dict(self) -> dict: + def to_dict(self) -> Dict[str, Any]: return { "status": self.status, "clarity_score": self.clarity_score, diff --git a/pyproject.toml b/pyproject.toml index 93c94f3..8a3e547 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -14,7 +14,7 @@ keywords = [ "agents", "rag", "input-guard", "prompt-quality" ] classifiers = [ - "Development Status :: 3 - Alpha", + "Development Status :: 4 - Beta", "Intended Audience :: Developers", "Operating System :: OS Independent", "Programming Language :: Python :: 3", @@ -22,13 +22,48 @@ classifiers = [ "Programming Language :: Python :: 3.10", "Programming Language :: Python :: 3.11", "Programming Language :: Python :: 3.12", + "Programming Language :: Python :: 3.13", ] +[project.urls] +Homepage = "https://github.com/nihanthnaidu007/Input_Guard" +Repository = "https://github.com/nihanthnaidu007/Input_Guard" +Changelog = "https://github.com/nihanthnaidu007/Input_Guard/blob/main/CHANGELOG.md" +Issues = "https://github.com/nihanthnaidu007/Input_Guard/issues" + [project.optional-dependencies] -dev = ["build>=1.2", "pytest>=8", "twine>=5"] +dev = [ + "build>=1.2", + "mypy>=1.8", + "pytest>=8", + "pytest-cov>=5", + "ruff>=0.5", + "twine>=5", +] + +[project.scripts] +inputguard = "inputguard.cli:main" [tool.setuptools.packages.find] include = ["inputguard*"] [tool.setuptools.package-data] inputguard = ["py.typed"] + +[tool.pytest.ini_options] +testpaths = ["tests"] + +[tool.ruff] +line-length = 100 + +[tool.ruff.lint] +select = ["E", "F", "W", "B"] +# E501: the codebase's prose-bearing table lines are exempt; the CI run +# treats ruff as the gate with this exact policy. +ignore = ["E501"] + +[tool.mypy] +strict = true +files = ["inputguard"] +# No python_version pin: modern mypy rejects 3.9 as a target, and the CI +# matrix already type-checks under each supported interpreter. diff --git a/tests/conftest.py b/tests/conftest.py new file mode 100644 index 0000000..c8459ea --- /dev/null +++ b/tests/conftest.py @@ -0,0 +1,33 @@ +"""Shared test fixtures. + +The registry-isolation fixture lives here so every module that mutates the +module-level ``REGISTRY`` (the contract-guard suite, the policy suite) uses +one canonical snapshot/restore implementation instead of reaching into +private registry state ad hoc. +""" + +from __future__ import annotations + +import pytest + +from inputguard.registry import REGISTRY + + +@pytest.fixture +def registry_isolation(): + """Snapshot the registry around tests that mutate it. + + Registration is permanent by design (there is no unregister), so tests + that register probe rules/domains restore the snapshot afterwards to keep + the process-global registry unpolluted for the rest of the suite. + """ + rules_before = dict(REGISTRY._rules) + domains_before = dict(REGISTRY._domains) + origins_before = dict(REGISTRY._origins) + yield REGISTRY + REGISTRY._rules.clear() + REGISTRY._rules.update(rules_before) + REGISTRY._domains.clear() + REGISTRY._domains.update(domains_before) + REGISTRY._origins.clear() + REGISTRY._origins.update(origins_before) diff --git a/tests/test_benchmarks.py b/tests/test_benchmarks.py new file mode 100644 index 0000000..e300346 --- /dev/null +++ b/tests/test_benchmarks.py @@ -0,0 +1,165 @@ +"""Latency budgets for the documented analysis path (spec art_bTvdPdJS §8). + +Every budget below is derived from measured numbers on this codebase, not +guesses. Two measurement environments inform them: + +- **Sandbox, untraced** (Python 3.13, 10,000-character inputs, 60 samples + after warmup, ``time.perf_counter``): + + build-intent (vague, C1 path): p50 21.94 ms / p95 26.09 / p99 27.30 + debug-intent (with findings): p50 6.59 ms / p95 8.06 / p99 8.07 + ready path (score 100): p50 5.62 ms / p95 5.88 / p99 6.03 + 1.6 MB input under the cap: 17.1 ms (truncated=True) + +- **CI runners, coverage-traced** (the first benchmark push observed, run + 35277175054): debug p50 21.14 ms, ready p50 20.35 ms — coverage tracing + roughly triples per-call cost, and runner hardware is slower than the + sandbox. That observation is why these tests are excluded from the + coverage pass (the ``test`` job) and gate only in the dedicated + benchmark job, untraced. + +What the measured path includes — deliberately: + +- **C1 double-run (adversarial review art_hC18m78C):** the build path runs + ``InsufficientContextRule.check``, which re-executes the six built-in build + checks plus the intent-detail check inside itself because the Rule protocol + is stateless — a ~2x multiplier on the busiest path. The build input below + is intentionally vague so this rule is always part of the measured cost; + any latency budget for build-intent inputs must account for it. +- **Word-boundary matching (PR #6):** ``contains_any`` matches at word + boundaries instead of raw substring scans; its measured delta is part of + every number above (it replaced the v0.2 substring matcher everywhere). +while still catching the failure class this guards against — algorithmic +regressions (accidental O(n^2), a lost length cap) blow past these budgets +by an order of magnitude, as the v0.2 unbounded 3.8 s scan would. The hard +gate is p50 (a stable statistic under CI noise), with a p95 +sustained-regression guard at 2.5x the p50 budget. Single-sample maxima +are NOT asserted — CI runners are noisy and a one-off slow sample must not +fail a merge. p50/p99 are printed so the CI benchmark job records them per +spec ("p50/p99 recorded in CI"). +""" + +from __future__ import annotations + +import time + +from inputguard import InputGuard + +# 10k characters: the default Policy.max_chars cap — the documented worst +# case for a single analyze() call. +_CAP = 10_000 +_SAMPLES = 60 +_WARMUP = 5 + +# Measured p50 21.94 ms untraced (including the C1 double-run); CI-traced +# observation stayed under 40 ms. Budget ~3x -> 65 ms; p95 guard 2.5x. +BUILD_P50_BUDGET_MS = 65.0 + +# Measured p50 6.59 ms untraced; 21.14 ms coverage-traced on CI runners. +# Budget ~3.8x untraced -> 25 ms; p95 guard 2.5x. +DEBUG_P50_BUDGET_MS = 25.0 + +# Measured p50 5.62 ms untraced; 20.35 ms coverage-traced on CI runners. +# Budget ~4.4x untraced -> 25 ms; p95 guard 2.5x. +READY_P50_BUDGET_MS = 25.0 + + +def _pad(text: str, filler: str, n: int = _CAP) -> str: + """Extend ``text`` to exactly ``n`` characters with neutral filler prose. + + The filler carries no rule triggers (no error/build/feature vocabulary); + only the seed text decides which intent and findings fire. + """ + reps = (n - len(text)) // len(filler) + 1 + return (text + " " + (filler * reps))[:n] + + +# Build-intent, deliberately vague: InsufficientContextRule is in the build +# ruleset, so this path pays the C1 re-run of seven checks inside check(). +_BUILD_VAGUE = _pad( + "build me an app", + "please and then the thing should be there somehow ", +) + +# Debug-intent with findings: a trigger with unresolved satisfies. +_DEBUG_GAPS = _pad( + "my python code is throwing an error when the users log in, " + "the app is just wrong", + "the function receives the arguments and the error happens again " + "when running it ", +) + +# Ready path: a fully-specified debug request (eval corpus TN-DBG-01 shape) +# padded to the cap — every trigger-and-satisfy pair resolves, score 100. +_READY = _pad( + "getting a TypeError in my Python function, expected a list but got None, " + "the function should return an empty list instead of crashing, " + "stack trace shows the failure happens on the iteration step", + "the function receives the arguments described above and returns the " + "value it should return; the behavior matches the description given ", +) + + +def _measure(input_text: str) -> tuple[float, float, float]: + """Return (p50, p95, p99) in milliseconds over _SAMPLES analyzed calls.""" + guard = InputGuard() + for _ in range(_WARMUP): + guard.analyze(input_text) + samples: list[float] = [] + for _ in range(_SAMPLES): + start = time.perf_counter() + guard.analyze(input_text) + samples.append((time.perf_counter() - start) * 1000.0) + samples.sort() + p50 = samples[len(samples) // 2] + p95 = samples[int(len(samples) * 0.95)] + p99 = samples[min(len(samples) - 1, int(len(samples) * 0.99))] + return p50, p95, p99 + + +def test_build_intent_p50_within_budget() -> None: + """C1 path: vague 10k build input, InsufficientContextRule re-running the + build checks inside check() — the documented worst-case busy path.""" + p50, p95, p99 = _measure(_BUILD_VAGUE) + print( + f"\n[benchmark] build-intent 10k: p50={p50:.2f}ms p95={p95:.2f}ms " + f"p99={p99:.2f}ms (budget p50<={BUILD_P50_BUDGET_MS}ms)" + ) + assert p50 <= BUILD_P50_BUDGET_MS, f"build p50 {p50:.2f}ms > {BUILD_P50_BUDGET_MS}ms" + assert p95 <= BUILD_P50_BUDGET_MS * 2.5, f"build p95 {p95:.2f}ms > sustained guard" + + +def test_debug_intent_p50_within_budget() -> None: + p50, p95, p99 = _measure(_DEBUG_GAPS) + print( + f"\n[benchmark] debug-intent 10k: p50={p50:.2f}ms p95={p95:.2f}ms " + f"p99={p99:.2f}ms (budget p50<={DEBUG_P50_BUDGET_MS}ms)" + ) + assert p50 <= DEBUG_P50_BUDGET_MS, f"debug p50 {p50:.2f}ms > {DEBUG_P50_BUDGET_MS}ms" + assert p95 <= DEBUG_P50_BUDGET_MS * 2.5, f"debug p95 {p95:.2f}ms > sustained guard" + + +def test_ready_path_p50_within_budget() -> None: + p50, p95, p99 = _measure(_READY) + print( + f"\n[benchmark] ready-path 10k: p50={p50:.2f}ms p95={p95:.2f}ms " + f"p99={p99:.2f}ms (budget p50<={READY_P50_BUDGET_MS}ms)" + ) + assert p50 <= READY_P50_BUDGET_MS, f"ready p50 {p50:.2f}ms > {READY_P50_BUDGET_MS}ms" + assert p95 <= READY_P50_BUDGET_MS * 2.5, f"ready p95 {p95:.2f}ms > sustained guard" + + +def test_unbounded_input_is_bounded_by_cap() -> None: + """The v0.2 failure mode (1.6 MB input taking ~3.8 s of linear scans) must + stay dead: the Policy cap bounds the scan and the result says so.""" + big = _pad(_BUILD_VAGUE, "more of the same neutral prose ", n=1_600_000) + assert len(big) == 1_600_000 + guard = InputGuard() + start = time.perf_counter() + result = guard.analyze(big) + elapsed_ms = (time.perf_counter() - start) * 1000.0 + print(f"\n[benchmark] 1.6MB input with cap: {elapsed_ms:.1f}ms, truncated={result.truncated}") + assert result.truncated is True + # 1s ceiling: 100x headroom over the measured capped path (~10 ms) while + # still catching any regression to the v0.2 unbounded 3.8 s behavior. + assert elapsed_ms < 1000.0, f"1.6MB input took {elapsed_ms:.1f}ms — the cap is not bounding the scan" diff --git a/tests/test_cli.py b/tests/test_cli.py new file mode 100644 index 0000000..91824c8 --- /dev/null +++ b/tests/test_cli.py @@ -0,0 +1,83 @@ +"""Tests for the ``inputguard`` CLI (spec art_bTvdPdJS §5). + +Covers the documented exit-code contract (0 = ok / 1 = below the +``--min-score`` floor / 2 = usage error), the text rendering shape, and that +``--format json`` emits exactly the result's ``to_dict()`` contract — the +CLI adds no fields and renames none. +""" + +from __future__ import annotations + +import io +import json + +import pytest + +from inputguard import InputGuard +from inputguard.cli import EXIT_BELOW_FLOOR, EXIT_OK, EXIT_USAGE, main + +VAGUE = "make this faster" # optimization intent, needs clarification +READY = ( + "getting a TypeError in my Python function, expected a list but got None" +) # eval corpus TN-DBG-01 shape: score 100 + + +def test_text_output_contains_verdict_lines(capsys: pytest.CaptureFixture[str]) -> None: + assert main(["analyze", VAGUE]) == EXIT_OK + out = capsys.readouterr().out + assert "status: needs_clarification" in out + assert "clarity: " in out + assert "intent: optimization" in out + assert "domain: coding" in out + assert "missing: " in out + assert "ask: " in out # the follow-up questions surface on the CLI too + + +def test_json_output_is_exactly_to_dict(capsys: pytest.CaptureFixture[str]) -> None: + assert main(["analyze", VAGUE, "--format", "json"]) == EXIT_OK + parsed = json.loads(capsys.readouterr().out) + expected = InputGuard().analyze(VAGUE).to_dict() + assert parsed == expected + + +def test_min_score_below_floor_exits_one() -> None: + assert main(["analyze", VAGUE, "--min-score", "85"]) == EXIT_BELOW_FLOOR + + +def test_min_score_at_or_above_floor_exits_zero() -> None: + assert main(["analyze", READY, "--min-score", "85"]) == EXIT_OK + + +def test_missing_text_is_a_usage_error() -> None: + with pytest.raises(SystemExit) as exc_info: + main(["analyze"]) + assert exc_info.value.code == EXIT_USAGE + + +def test_unknown_domain_is_a_usage_error_not_a_traceback( + capsys: pytest.CaptureFixture[str], +) -> None: + assert main(["analyze", VAGUE, "--domain", "legal"]) == EXIT_USAGE + err = capsys.readouterr().err + assert "inputguard:" in err + assert "legal" in err + + +def test_stdin_input(monkeypatch: pytest.MonkeyPatch) -> None: + monkeypatch.setattr("sys.stdin", io.StringIO(VAGUE)) + assert main(["analyze", "--stdin"]) == EXIT_OK + + +def test_strict_mode_flag_is_accepted() -> None: + assert main(["analyze", READY, "--mode", "strict", "--min-score", "85"]) == EXIT_OK + + +def test_degraded_note_appears_in_text_output(capsys: pytest.CaptureFixture[str]) -> None: + # A Chinese build request (spec probe P2): the language probe must mark + # the result degraded, and the CLI must show that note instead of a + # silent ready. + degraded_input = "建造一个用户登录应用" + assert main(["analyze", degraded_input]) == EXIT_OK + out = capsys.readouterr().out + assert "note: " in out + assert "status: ready" not in out diff --git a/tests/test_followups.py b/tests/test_followups.py index eaffd2b..1d7d6c1 100644 --- a/tests/test_followups.py +++ b/tests/test_followups.py @@ -73,7 +73,7 @@ def test_every_template_renders_with_unknown_slot_failing_loudly(): # Rendering every template against empty slots exercises the render path: # a typo'd slot name raises (loud table bug); a known-but-unfilled slot # skips the template (documented behavior); plain templates pass through. - for gap, questions in _FOLLOW_UP_QUESTIONS.items(): + for _gap, questions in _FOLLOW_UP_QUESTIONS.items(): for question in questions: rendered = _render(question, slots={}) if "{" in question: @@ -368,8 +368,13 @@ def run(text: str) -> None: small_ms = _wall_ms(run, "a" * 1_000) large_ms = _wall_ms(run, "a" * 10_000) assert large_ms < 40 * max(small_ms, 0.05) - # And both stay fast in absolute terms — 1.3 s at 10 K was the old number. - assert large_ms < 50.0 + # And both stay fast in absolute terms — 1.3 s at 10 K was the old + # number. The bound is a single wall-clock sample, so it carries CI + # headroom: coverage-traced runs on slow runners measured up to ~51 ms + # (a bare 50 ms bound flaked there, run 35277834661), while the + # quadratic regression this guards against is 1.3 s — 5x+ margin even + # at 250 ms. + assert large_ms < 250.0 def test_dataset_file_extraction_dotted_filler_stays_linear_from_1k_to_10k(): @@ -379,7 +384,9 @@ def run(text: str) -> None: small_ms = _wall_ms(run, "a." * 500) large_ms = _wall_ms(run, "a." * 5_000) assert large_ms < 40 * max(small_ms, 0.05) - assert large_ms < 50.0 + # Same CI headroom rationale as above: single sample, traced runners + # measured ~51 ms, regression is 1.3 s. + assert large_ms < 250.0 def test_dataset_file_extraction_is_case_insensitive_as_before(): diff --git a/tests/test_policy.py b/tests/test_policy.py index 0fc1622..23db685 100644 --- a/tests/test_policy.py +++ b/tests/test_policy.py @@ -17,7 +17,8 @@ from inputguard.policy import SEVERITIES from inputguard.registry import KNOWN_SEVERITIES, register_rule from inputguard.scorer import require_known_severity -from test_registry import _test_rule, registry_isolation # noqa: F401 — pytest fixture +from test_registry import _test_rule # noqa: F401 — shared rule builder +# registry_isolation resolves as a pytest fixture from tests/conftest.py. # Observed v0.2 / release-branch behavior (runtime probes), reused as fixtures. VAGUE_INPUT = "do something now" # build intent, one high finding: insufficient_context diff --git a/tests/test_registry_contract_guards.py b/tests/test_registry_contract_guards.py new file mode 100644 index 0000000..1ec0cac --- /dev/null +++ b/tests/test_registry_contract_guards.py @@ -0,0 +1,268 @@ +"""Registry contract-guard tests (adversarial review art_hC18m78C). + +The three tests the review found missing, as a dedicated CI-enforced suite: + +1. Duplicate rule ids within one ``register_domain`` call raise ``ValueError`` + (review C3/probe P2 — the high-severity second rule used to vanish + silently) and nothing mutates on rejection. +2. A rule whose ``check`` signature cannot be called as ``check(self, text)`` + is rejected *at registration* with an actionable message (review B1/N3/ + probe P1 — a mismatched rule used to pass registration and detonate mid- + ``analyze()`` with a ``TypeError`` on arbitrary user input). +3. A rule that raises inside ``check()`` surfaces as a ``RuntimeError`` + carrying the rule's id and registration origin, with the original + exception chained (review C2/P4 — exceptions used to escape raw, with no + attribution and no recovery path). + +These complement the remediation tests in ``test_registry.py`` (which covers +the ``register_rule`` path); this module is the contract-guards gate the CI +matrix runs on every push, and also pins the ``register_domain`` road. +""" + +from __future__ import annotations + +import pytest + +from inputguard import InputGuard, RuleFinding +from inputguard.registry import REGISTRY, register_domain, register_rule + + +# --- 1 · same-call duplicate rule ids raise, nothing mutates ---------------- + + +class _ProbeRule: + """Probe rule with a configurable id, firing on the word ``widget``.""" + + domain = "probe thing" + severity = "high" + gap = "probe gap" + + def __init__(self, rule_id: str) -> None: + self.id = rule_id + + def check(self, text: str): + if "widget" in text: + return RuleFinding( + code=self.id, + message="probe rule fired", + severity=self.severity, + gap=self.gap, + ) + return None + + +_PROBE_SIGNALS = {"probe thing": ("widget",), "probe review": ()} + + +def test_same_call_duplicate_rule_ids_raise_and_mutate_nothing(registry_isolation): + """Review C3: duplicate ids inside one register_domain call must raise. + + The original hole: the pre-check compared only against *already + registered* rules, so a same-call duplicate was silently skipped in the + mutation loop — the second rule disappeared without an error. The error + must name the duplicated id, and a rejected batch must leave the registry + untouched ("nothing mutates unless every check passes"). + """ + with pytest.raises(ValueError, match="register_domain call.*'dup_x'") as exc_info: + register_domain( + "probe_dup", + _PROBE_SIGNALS, + rules=[_ProbeRule("dup_x"), _ProbeRule("dup_x")], + ) + + # The message names the offending id so the author can fix it in one look. + assert "dup_x" in str(exc_info.value) + + # No partial mutation: neither the domain nor either rule registered. + assert "probe_dup" not in REGISTRY.domain_names() + assert "dup_x" not in REGISTRY.rule_ids() + + # A corrected call through the same path works — the guard blocks + # duplicates, not the batch API itself. + register_domain("probe_dup_fixed", _PROBE_SIGNALS, rules=[_ProbeRule("solo_rule")]) + result = InputGuard().analyze("widget please", domain="probe_dup_fixed") + assert "solo_rule" in [f.code for f in result.findings] + + +def test_duplicate_hidden_among_distinct_ids_raises(registry_isolation): + """A duplicate buried in a batch of otherwise-distinct ids still raises, + and the message names exactly the duplicated id.""" + with pytest.raises(ValueError, match="'twin_b'") as exc_info: + register_domain( + "probe_dup_mixed", + _PROBE_SIGNALS, + rules=[ + _ProbeRule("unique_a"), + _ProbeRule("twin_b"), + _ProbeRule("twin_b"), + _ProbeRule("unique_c"), + ], + ) + # The distinct ids are not blamed. + assert "unique_a" not in str(exc_info.value) + assert "unique_c" not in str(exc_info.value) + assert "probe_dup_mixed" not in REGISTRY.domain_names() + assert "unique_a" not in REGISTRY.rule_ids() + assert "twin_b" not in REGISTRY.rule_ids() + + +# --- 2 · wrong check() signature is a registration error -------------------- + + +def test_extra_parameter_check_rejected_at_registration(registry_isolation): + """Review B1: the retired check(self, text, intent) shape must be a + registration error, not a mid-analyze TypeError. The message is + actionable: it names the rule, the one-positional-argument dispatch, and + the expected check(self, text) signature.""" + probe_signals = {"probe sig": ("gadget",), "probe sig review": ()} + + class ExtraArgRule: + id = "probe_extra_arg" + domain = "probe sig" + severity = "low" + gap = None + + def check(self, text, intent): # noqa: ARG001 — the retired shape + return None + + with pytest.raises(TypeError) as exc_info: + register_domain("probe_sig", probe_signals, rules=[ExtraArgRule()]) + + message = str(exc_info.value) + assert "probe_extra_arg" in message + assert "check(self, text)" in message + assert "probe_extra_arg" not in REGISTRY.rule_ids() + assert "probe_sig" not in REGISTRY.domain_names() + + +def test_zero_parameter_check_rejected_at_registration(registry_isolation): + """check() with no parameters cannot bind the normalized text — reject at + registration.""" + + class NoArgRule: + id = "probe_noarg" + domain = "debug" + severity = "low" + gap = None + + def check(self): + return None + + with pytest.raises(TypeError, match="probe_noarg") as exc_info: + register_rule(NoArgRule()) + assert "check(self, text)" in str(exc_info.value) + assert "probe_noarg" not in REGISTRY.rule_ids() + + +def test_keyword_only_check_rejected_at_registration(registry_isolation): + """check(self, *, text) cannot be called positionally — the analyzer + dispatches rule.check(normalized), so keyword-only rejects too.""" + + class KeywordOnlyRule: + id = "probe_kwonly" + domain = "debug" + severity = "low" + gap = None + + def check(self, *, text): + return None + + with pytest.raises(TypeError, match="probe_kwonly"): + register_rule(KeywordOnlyRule()) + assert "probe_kwonly" not in REGISTRY.rule_ids() + + +def test_defaulted_check_is_accepted_and_fires(registry_isolation): + """check(self, text="...") binds one positional argument, so it satisfies + the dispatch contract — the gate rejects what cannot be called, not + signatures it merely dislikes.""" + + class DefaultedRule: + id = "probe_defaulted" + domain = "debug" + severity = "low" + gap = None + + def check(self, text=""): + if "fix the bug" in text: + return RuleFinding( + code=self.id, + message="Defaulted-signature rule fired.", + severity=self.severity, + gap=self.gap, + ) + return None + + register_rule(DefaultedRule()) + result = InputGuard().analyze("fix the bug in my app") + assert "probe_defaulted" in [f.code for f in result.findings] + + +def test_spec_signature_rule_is_accepted_and_fires(registry_isolation): + """Positive control for the arity gate: check(self, text) — the pinned + spec's signature (art_bTvdPdJS §1) — registers cleanly and the analyzer + dispatches it. The gate rejects deviations, never the documented shape.""" + + class SpecRule: + id = "probe_spec_signature" + domain = "debug" + severity = "medium" + gap = "error description" + + def check(self, text: str): + if "fix the bug" in text: + return RuleFinding( + code=self.id, + message="Spec-signature rule fired.", + severity=self.severity, + gap=self.gap, + ) + return None + + register_rule(SpecRule()) + result = InputGuard().analyze("fix the bug in my app") + assert "probe_spec_signature" in [f.code for f in result.findings] + + +# --- 3 · dispatch exceptions carry rule attribution -------------------------- + + +def test_raising_rule_surfaces_id_origin_and_cause_via_register_domain( + registry_isolation, +): + """Review C2/P4: an exception inside check() aborts analyze() with the + rule's id, its registration origin, and the original traceback chained. + Tested through the register_domain path — test_registry.py covers the + register_rule path — so both registration roads get the same loud + attribution.""" + + class ExplodingRule: + id = "probe_boom_domain_rule" + domain = "probe boom thing" + severity = "high" + gap = None + + def check(self, text: str): + raise ZeroDivisionError("division by zero in probe rule") + + register_domain( + "probe_boom", + {"probe boom thing": ("widget",), "probe boom review": ()}, + rules=[ExplodingRule()], + ) + + with pytest.raises(RuntimeError) as exc_info: + InputGuard().analyze("widget please", domain="probe_boom") + + message = str(exc_info.value) + # Rule id, exception type, and registration origin all named. + assert "probe_boom_domain_rule" in message + assert "ZeroDivisionError" in message + assert "registered at" in message + # The original exception is chained, not swallowed. + assert isinstance(exc_info.value.__cause__, ZeroDivisionError) + assert "division by zero in probe rule" in str(exc_info.value.__cause__) + # The origin names the file that registered the rule — this test module, + # not inputguard internals (registration happened via register_domain + # above, so the recorded call site is the register_domain line here). + assert __file__ in message From 0f4c043d763cfc96a3bf0f9d325d39af5b6ab2ac Mon Sep 17 00:00:00 2001 From: Obvious Date: Thu, 17 Sep 2026 21:59:45 +0000 Subject: [PATCH 14/18] docs(readme-tested-examples): rewrite README for v0.3 with every example executed by tests The v0.2 README drifted from recommender.py within two releases. The rewritten README documents the v0.3 surface (extension API, Policy, follow-ups, degradation, CLI) and every example now runs: a new tests/test_readme_examples.py executes all python blocks in document order in a subprocess, diffs the embedded to_dict() JSON against live analyzer output, and checks exports, rule tables, gap vocabulary, and CLI output against the live registry. Example values carry real assertions sourced from runtime probes. Co-authored-by: Kalisetti Nihanth Naidu --- README.md | 511 ++++++++++++++++++++++++++-------- tests/test_readme_examples.py | 164 +++++++++++ 2 files changed, 558 insertions(+), 117 deletions(-) create mode 100644 tests/test_readme_examples.py diff --git a/README.md b/README.md index 5a2b13c..cdb1e6e 100644 --- a/README.md +++ b/README.md @@ -3,10 +3,14 @@ Catch unclear inputs before they become bad AI outputs. -InputGuard is a pre-flight input clarity layer. It sits between a user's input and an LLM call. It detects vague, incomplete, or unspecific inputs before they reach the AI — saving the correction cycle that wastes time and tokens when the AI guesses wrong. +InputGuard is a pre-flight input clarity layer. It sits between a user's input and an LLM call. It detects vague, incomplete, or unspecific inputs before they reach the AI — saving the correction cycle that wastes time and tokens when the AI guesses wrong. When it finds gaps, it returns the questions that close them, ready to send back to the user. Zero LLM calls. Zero external dependencies. Pure local Python. +- **Deterministic rules you can extend** — register your own rules and analysis domains through the same typed API the built-ins use. +- **Calibration as data** — status bands, severity penalties, rule filters, and the input cap live in one frozen `Policy` object. +- **Honest by construction** — non-English input is reported as a limitation of the tool, never silently scored `ready`. + --- ## The problem it solves @@ -33,7 +37,7 @@ A non-technical user asks to "fix my code." The AI guesses at the problem, picks pip install inputguard ``` -Python 3.9+. No external dependencies. +Python 3.9+. No external dependencies — `pip install inputguard` imports nothing outside the standard library. --- @@ -45,21 +49,56 @@ from inputguard import InputGuard guard = InputGuard() result = guard.analyze("fix my code") -print(result.detected_intent) # 'debug' -print(result.status) # 'needs_clarification' -print(result.clarity_score) # 35 -print(result.gaps) # ['error description', 'expected vs actual behavior', 'code context'] -print(result.is_clear()) # False +assert result.detected_intent == "debug" +assert result.status == "needs_clarification" +assert result.clarity_score == 35 +assert result.gaps == ["error description", "expected vs actual behavior", "code context"] +assert result.is_clear() is False +``` + +Every gap comes with plain-English advice and the questions that close it: -for rec in result.recommendations: - print(rec["gap"]) - print(rec["what_is_missing"]) - print(rec["what_to_provide"]) - print(rec["why_it_matters"]) - print() +```python +for question in result.follow_ups: + print("-", question) + +assert len(result.follow_ups) == 3 +assert result.follow_ups[0].startswith("What is the exact error message") ``` -`analyze()` takes an optional `domain` argument, which defaults to `"coding"`. Phase 1 supports `"coding"` only. Intent detection is automatic — no extra parameters needed. +`analyze()` takes an optional `domain` argument, which defaults to `"coding"`. Three domains ship built in — `coding`, `writing`, and `data-analysis` — and you can register your own (see [Custom domains](#custom-domains)). + +--- + +## The public API + +Nine names; everything else is internal: + +```python +from inputguard import ( + AnalysisResult, + InputGuard, + Policy, + REGISTRY, + Rule, + RuleFinding, + __version__, + register_domain, + register_rule, +) +``` + +| Name | What it is | +|---|---| +| `InputGuard` | The engine: `InputGuard(mode="warning", policy=None)` | +| `AnalysisResult` | The frozen result of `analyze()` | +| `RuleFinding` | One rule's raw verdict: `code`, `message`, `severity`, `gap` | +| `Policy` | Frozen calibration: bands, penalties, filters, input cap | +| `Rule` | The protocol a custom rule implements | +| `register_rule` | Register one rule (decorator or call) | +| `register_domain` | Register an analysis domain: intent signals + rules | +| `REGISTRY` | The process-global registry (read during `analyze()`) | +| `__version__` | The installed version string | --- @@ -74,12 +113,12 @@ warn_guard = InputGuard(mode="warning") # Strict mode — blocks inputs that fall below the clarity threshold strict_guard = InputGuard(mode="strict") -result_warn = warn_guard.analyze("build a REST API") +result_warn = warn_guard.analyze("build a REST API") result_strict = strict_guard.analyze("build a REST API") -print(result_warn.status) # 'needs_clarification' -print(result_strict.status) # 'blocked' -print(result_warn.clarity_score == result_strict.clarity_score) # True +assert result_warn.status == "needs_clarification" +assert result_strict.status == "blocked" +assert result_warn.clarity_score == result_strict.clarity_score ``` The clarity score is mode-independent. Only the status threshold changes. @@ -96,6 +135,25 @@ Use `warning` when you want to surface gaps to the user without blocking. Use `s --- +## Follow-up questions + +Every gap maps to one or two templated clarifying questions — the sentence the user can answer verbatim. Questions dedupe and order with the gaps: + +```python +result = guard.analyze("make this faster") + +assert result.gaps == [ + "optimization target", + "performance baseline", + "optimization constraint", +] +assert result.follow_ups[0] == "Which function or module should get faster?" +``` + +A gap with no built-in question table entry — including gaps from your own custom rules — gets a documented fallback question, never silence. The same contract holds for recommendations. + +--- + ## Non-English input InputGuard's rules are English-language heuristics. Before any rule runs, a zero-dependency script probe (stdlib `unicodedata` only) classifies the input's script and language. Inputs the rules cannot assess take the explicit degraded path — an uncovered dominant script, Latin-script text recognized as French/Spanish/Portuguese by its function words, or a Latin-dominant input carrying a run of 3+ consecutive uncovered-script letters. In every case InputGuard says so instead of pretending: @@ -103,12 +161,12 @@ InputGuard's rules are English-language heuristics. Before any rule runs, a zero ```python result = guard.analyze("建造一个用户登录应用") -result.status # 'degraded' — never 'ready', in either mode -result.clarity_score # 80 (100 minus the degradation penalty) -result.detected_intent # 'undetermined' -result.detected_language # 'zh' (coarse, script-derived guess) -result.heuristic_coverage # 'none' -result.degradation_note # explains that rules were skipped and why +assert result.status == "degraded" +assert result.clarity_score == 80 # 100 minus the degradation penalty +assert result.detected_intent == "undetermined" +assert result.detected_language == "zh" +assert result.heuristic_coverage == "none" +assert result.degradation_note is not None ``` The rules are **skipped explicitly** — running English keyword rules on text they cannot assess would produce a silent, unearned verdict or spurious gaps invented out of the silence. A degraded result reports the literal `degraded` status in both modes: it is the tool reporting a language limitation of itself, not a judgment of the input's clarity (strict mode's banding would otherwise read as an ordinary critique). In v0.2 this input silently scored 100/ready; v0.3 refuses to assert a confidence it does not have. @@ -131,17 +189,168 @@ The degraded path covers three shapes of input the English rules cannot assess: ```python result = guard.analyze("Preciso de um aplicativo web com login de usuário e relatórios") -result.detected_language # 'pt' (function-word guess) -result.heuristic_coverage # 'none' -result.status # 'degraded' — never 'ready' -result.degradation_note # names Portuguese (pt) and explains the skip +assert result.detected_language == "pt" +assert result.heuristic_coverage == "none" +assert result.status == "degraded" +assert "Portuguese" in result.degradation_note +``` + +--- + +## Custom rules + +The extension contract is four members and one method. Built-in rules register through the exact same path — the API is exercised by all 31 built-in rules before anyone writes their own. + +```python +from inputguard import RuleFinding, register_domain + +class CheckRollbackPlan: + id = "missing_rollback_plan" # unique across the registry + domain = "deploy" # the INTENT name this rule runs for + severity = "high" # "low" | "medium" | "high" — validated + gap = "rollback plan" # groups findings for scoring dedup + + def check(self, text: str) -> RuleFinding | None: + # text arrives normalized: lowercased, whitespace-collapsed. + if "rollback" not in text and "roll back" not in text: + return RuleFinding( + code=self.id, + message="No rollback plan described.", + severity=self.severity, + gap=self.gap, + ) + return None +``` + +The rule above ships with the custom domain in the next section — `register_domain` registers rules through the same path `register_rule` uses, once the intents they name exist. The contracts a rule author accepts: + +- **`check` receives normalized text** (lowercased, whitespace-collapsed) and returns at most one `RuleFinding`, or `None` when the rule does not fire. +- **A `check` exception is never swallowed** — it aborts the `analyze()` call in flight, naming the rule and where it was registered. +- **Registration is permanent** for the process lifetime (there is no unregister), and `REGISTRY` is a process-global singleton shared by everything that imports inputguard. +- **Duplicate rule ids raise `ValueError`** at registration, as do unknown severities and rules naming an intent no registered domain declares. +- **Every gap deserves advice**: a gap with no recommendation or follow-up table entry gets a documented generic fallback — never an empty list and never silence. + +--- + +## Custom domains + +A domain is a named analysis scope: intent signals in priority order, plus its rules. The single intent with empty terms is the fallback for inputs no other intent matches. + +```python +register_domain( + "devops", + { + "deploy": ("deploy", "deployment", "release", "rollout", "ship", "shipping"), + "general": (), # fallback intent — inputs no deploy signal matches + }, + [CheckRollbackPlan], +) + +result = guard.analyze("deploy the new checkout service to production", domain="devops") + +assert result.detected_intent == "deploy" +assert result.gaps == ["rollback plan"] +assert result.clarity_score == 75 +assert result.status == "usable_with_warnings" +assert result.follow_ups == ["Can you add the rollback plan this request is missing?"] +``` + +Intent names are globally unique across domains — registering a domain whose intent another domain already declares raises `ValueError`, because rules dispatch on intent name alone. The same guard rejects duplicate domain names, so re-registering `devops` above raises too. + +--- + +## Policy: calibrate the guard + +Every scoring constant is data on a frozen, validated `Policy` — with the v0.2 values pinned as defaults, so existing behavior cannot drift. Two layers stay separate: per-rule severity decides *what fires*; the policy's bands decide *what happens*. + +| Field | Default | Controls | +|---|---|---| +| `ready_at` | `85` | score ≥ `ready_at` is `ready`, both modes | +| `usable_at` | `60` | warning-mode floor for `usable_with_warnings` | +| `strict_clarify_at` | `65` | strict-mode floor for `needs_clarification`; below it, strict blocks | +| `penalty_low` / `penalty_medium` / `penalty_high` | `5` / `15` / `25` | points per distinct gap, by highest severity | +| `min_words` | `3` | shorter input is never flagged as vague | +| `max_chars` | `10_000` | input cap — truncation is reported, never silent | +| `borderline_at` | `74` | near-miss band just below `ready_at` — the "worth one more pass" signal | +| `disabled_rules` | `frozenset()` | rule ids to skip entirely | +| `allow_patterns` | `()` | regex patterns; matching input is never flagged | + +```python +from inputguard import Policy + +# Disable a rule whose advice your product handles elsewhere — +# the disabled low-severity penalty comes back (55 -> 60) and the +# status flips out of needs_clarification. +tuned = InputGuard( + mode="warning", + policy=Policy(disabled_rules=frozenset({"missing_optimization_constraint"})), +) +result = tuned.analyze("make this faster") +assert result.clarity_score == 60 +assert result.status == "usable_with_warnings" + +# Allowlist internal input shapes — matching input is never flagged. +allowlisted = InputGuard(policy=Policy(allow_patterns=("^re:",))) +result = allowlisted.analyze("re: invoice numbering scheme question") +assert result.status == "ready" +assert result.gaps == [] +``` + +The cap bounds every analysis; a longer input is truncated and the result says so: + +```python +result = InputGuard().analyze("word " * 3000) +assert result.truncated is True +``` + +Policy is validated at construction — mis-ordered bands raise `ValueError` instead of silently distorting statuses — and a per-call `policy=` argument overrides the guard's for one `analyze()` call: + +```python +strict = InputGuard(mode="strict") + +default_bands = strict.analyze("make this faster") +assert default_bands.status == "blocked" # 55 < 65 + +loosened = strict.analyze("make this faster", policy=Policy(usable_at=50, strict_clarify_at=50)) +assert loosened.status == "needs_clarification" # 55 >= 50 ``` --- +## CLI + +A zero-dependency `argparse` front end over the same pipeline, installed as a console script: + +```bash +pip install inputguard +inputguard analyze "make this faster" --mode strict +``` + +```text +status: blocked +clarity: 55/100 +intent: optimization +domain: coding +missing: optimization target, performance baseline, optimization constraint +ask: Which function or module should get faster? +ask: How slow is it today, and what latency would be acceptable? +ask: What must not change while it gets faster — an interface, readability, behavior others depend on? +``` + +`--format json` emits the same `to_dict()` contract the library documents, and `--min-score` turns the CLI into a commit-hook / CI gate: + +```bash +inputguard analyze "$(cat prompt.txt)" --format json --min-score 85 +cat prompt.txt | inputguard analyze --stdin --min-score 85 +``` + +Exit codes: `0` analysis completed (score at or above `--min-score` when given), `1` score below the `--min-score` floor, `2` usage error (unknown flags, missing text, invalid domain or input). + +--- + ## How intent detection works -InputGuard automatically detects what kind of coding input it is receiving. No extra parameters needed. The same `.analyze()` call handles all five intent types. +Each domain detects intent over its own registered signals — insertion order is the priority chain, and the fallback intent catches what no signal matches. The coding domain detects five intents automatically; no extra parameters needed. | Intent | What it covers | Example input | |---|---|---| @@ -150,18 +359,20 @@ InputGuard automatically detects what kind of coding input it is receiving. No e | `optimization` | Performance, speed, refactoring | "make this function faster" | | `explanation` | Understanding code or concepts | "explain what this decorator does" | | `feature` | Adding to existing code | "add search to my existing React app" | +| `compose` *(writing)* | Essays, emails, documents, posts | "write a blog post" | +| `analysis` *(data analysis)* | Analyzing datasets, building reports | "analyze my sales data" | ```python -guard = InputGuard() - -guard.analyze("build a REST API").detected_intent # 'build' -guard.analyze("fix my code").detected_intent # 'debug' -guard.analyze("make this faster").detected_intent # 'optimization' -guard.analyze("explain how async works").detected_intent # 'explanation' -guard.analyze("add search to my existing app").detected_intent # 'feature' +assert guard.analyze("build a REST API").detected_intent == "build" +assert guard.analyze("fix my code").detected_intent == "debug" +assert guard.analyze("make this faster").detected_intent == "optimization" +assert guard.analyze("explain how async works").detected_intent == "explanation" +assert guard.analyze("add search to my existing app").detected_intent == "feature" +assert guard.analyze("write a blog post", domain="writing").detected_intent == "compose" +assert guard.analyze("analyze my sales data", domain="data-analysis").detected_intent == "analysis" ``` -When an input is ambiguous, debug always wins. "Fix this slow function" is a debug request, not optimization. Priority order is: debug → optimization → explanation → feature → build. +When a coding input is ambiguous, debug always wins. "Fix this slow function" is a debug request, not optimization. Priority order is: debug → optimization → explanation → feature → build. --- @@ -223,7 +434,7 @@ Requests like "add search to my existing app", "extend my current API with pagin ### Writing inputs -Requests like "write a blog post", "draft an email", "proofread my essay". Every gap carries its own follow-up questions, and each gap names one thing at a time — same one-gap-one-question discipline as the coding domains. +Requests like "write a blog post", "draft an email", "proofread my essay" — analyzed with `domain="writing"`. Every gap carries its own follow-up questions, and each gap names one thing at a time — same one-gap-one-question discipline as the coding domains. | Rule code | What it catches | Severity | |---|---|---| @@ -236,7 +447,7 @@ Requests like "write a blog post", "draft an email", "proofread my essay". Every ### Data-analysis inputs -Requests like "analyze my sales data", "build a dashboard", "report on this spreadsheet". +Requests like "analyze my sales data", "build a dashboard", "report on this spreadsheet" — analyzed with `domain="data-analysis"`. | Rule code | What it catches | Severity | |---|---|---| @@ -257,9 +468,9 @@ Requests like "analyze my sales data", "build a dashboard", "report on this spre |---|---|---| | `status` | `str` | One of `"ready"`, `"usable_with_warnings"`, `"needs_clarification"`, `"blocked"`, `"degraded"` (language limitation — see Non-English input) | | `clarity_score` | `int` | 0 to 100 | -| `detected_intent` | `str` | Which intent was detected: `build`, `debug`, `optimization`, `explanation`, `feature`, `compose`, or `analysis` | +| `detected_intent` | `str` | The detected intent: `build`, `debug`, `optimization`, `explanation`, `feature`, `compose`, `analysis`, or `undetermined` (degraded inputs) | | `gaps` | `List[str]` | Gap names, in the order rules fired | -| `recommendations` | `List[dict]` | One dict per gap (see next section) | +| `recommendations` | `List[dict]` | One dict per gap (see The gap vocabulary) | | `follow_ups` | `List[str]` | One or two clarifying questions per gap, ready to send back to the user | | `findings` | `List[RuleFinding]` | Raw rule findings (code, message, severity, gap) | | `interpretation_note` | `Optional[str]` | Set when the input is highly ambiguous (score < 50 or two or more high-severity findings) | @@ -268,7 +479,7 @@ Requests like "analyze my sales data", "build a dashboard", "report on this spre | `degradation_note` | `Optional[str]` | Set when rule analysis was skipped or limited by language coverage | | `borderline` | `bool` | `True` when the score sits in the near-miss band just below ready — worth one more pass | | `truncated` | `bool` | `True` when the input was capped at 10,000 characters (only the prefix was analyzed) | -| `score_breakdown` | `Optional[dict]` | Per-contribution score arithmetic, present when the policy exposes it | +| `score_breakdown` | `Optional[dict]` | Per-contribution score arithmetic (base, per-gap penalties, final) | Helpers: @@ -290,132 +501,190 @@ Example `result.to_dict()` for `guard.analyze("fix my code")`: "recommendations": [ { "gap": "error description", - "what_is_missing": "You haven't included the actual error message or exception.", - "what_to_provide": "Copy and paste the exact error message. For example: 'I'm getting TypeError: cannot read property of undefined on line 23'.", - "why_it_matters": "The exact wording tells the AI exactly what went wrong. Without it, the AI guesses and often fixes the wrong thing." + "what_is_missing": "You haven't shared the actual error message, exception, or output you're seeing.", + "what_to_provide": "Include the exact error message or exception you are seeing. Copy and paste it exactly as it appears. For example: 'I'm getting TypeError: cannot read property of undefined on line 23' or 'it throws a 500 Internal Server Error with message: connection refused'. The exact wording tells the AI exactly what went wrong.", + "why_it_matters": "Without the exact error, the AI has to guess what failure mode you're hitting. The wrong guess sends you down a fix path that doesn't apply to your actual problem." }, { "gap": "expected vs actual behavior", - "what_is_missing": "You haven't described what should happen vs what actually happens.", - "what_to_provide": "Describe both. For example: 'it should return a list of users but instead returns None every time'.", - "why_it_matters": "Without this, the AI is guessing what the problem is. It may fix something that was not broken." + "what_is_missing": "You haven't described what you expected to happen and what is actually happening.", + "what_to_provide": "Describe two things: what you expected to happen, and what actually happened. For example: 'I expected the function to return a list of users, but it returns an empty list every time' or 'the button should submit the form but nothing happens when I click it'. Without this, the AI is guessing what the problem is.", + "why_it_matters": "A bug is the gap between what you wanted and what happened. Without both sides, the AI cannot tell what counts as a fix." }, { "gap": "code context", - "what_is_missing": "You haven't pointed to the specific part of your code with the problem.", - "what_to_provide": "Name the language and the function. For example: 'this is a Python function called get_users()'.", - "why_it_matters": "The more specific you are, the more targeted the fix will be." + "what_is_missing": "You haven't pointed to a language, file, function, or snippet for the AI to look at.", + "what_to_provide": "Tell it which language you are using and point to the specific part of your code that has the problem. For example: 'this is a Python function called get_users()' or 'this is in my React component UserList.jsx on line 45'. The more specific you are, the more targeted the fix will be.", + "why_it_matters": "Without a code reference, the AI suggests generic fixes that may not apply to your actual code. Pointing to the exact location lets it propose a precise change." } ], + "follow_ups": [ + "What is the exact error message or exception you're seeing (copy it verbatim if you can)?", + "What did you expect to happen, and what actually happens instead?", + "Which file, function, or part of your code does the problem live in?" + ], "findings": [ - {"code": "missing_error_message", "message": "Debug request detected but no error message or exception described.", "severity": "high", "gap": "error description"}, - {"code": "missing_expected_vs_actual", "message": "No description of expected vs actual behavior provided.", "severity": "high", "gap": "expected vs actual behavior"}, - {"code": "missing_debug_code_context", "message": "No code context provided.", "severity": "medium", "gap": "code context"} + { + "code": "missing_error_message", + "message": "Debug request detected but no error message or exception described.", + "severity": "high", + "gap": "error description" + }, + { + "code": "missing_expected_vs_actual", + "message": "No description of expected vs actual behavior provided.", + "severity": "high", + "gap": "expected vs actual behavior" + }, + { + "code": "missing_debug_code_context", + "message": "No code context provided — no language, function name, or snippet referenced.", + "severity": "medium", + "gap": "code context" + } ], - "interpretation_note": "This input is ambiguous in multiple ways. Addressing each gap below before sending will prevent the AI from making assumptions that lead to the wrong output." + "interpretation_note": "This input is ambiguous in multiple ways. Addressing each gap below before sending will prevent the AI from making assumptions that lead to the wrong output.", + "detected_language": "en", + "heuristic_coverage": "full", + "degradation_note": null, + "borderline": false, + "truncated": false, + "score_breakdown": { + "base": 100, + "penalties": [ + { + "code": "missing_error_message", + "severity": "high", + "points": -25 + }, + { + "code": "missing_expected_vs_actual", + "severity": "high", + "points": -25 + }, + { + "code": "missing_debug_code_context", + "severity": "medium", + "points": -15 + } + ], + "final": 35 + } } ``` --- -## Recommendations +## The gap vocabulary + +Every built-in gap string, by intent. These strings are de-facto API — consumers switch on them — and every one has a recommendation and at least one follow-up question. + +- **Build:** `programming language`, `api structure`, `data model`, `integration specifics`, `authentication type`, `output format`, `task context` +- **Debug:** `error description`, `expected vs actual behavior`, `code context` +- **Optimization:** `optimization target`, `performance baseline`, `optimization constraint` +- **Explanation:** `code reference`, `explanation depth` +- **Feature:** `existing stack`, `feature scope`, `completion criteria` +- **Compose:** `audience`, `purpose`, `structure/format`, `source material`, `context`, `completeness` +- **Analysis:** `dataset/source`, `question/goal`, `output format`, `tooling`, `volume`, `reproducibility` Every entry in `result.recommendations` is a plain dict with four keys, all written for non-technical users: ```python +result = guard.analyze("fix my code") for rec in result.recommendations: + assert set(rec) == {"gap", "what_is_missing", "what_to_provide", "why_it_matters"} print(rec["gap"]) # which gap this addresses print(rec["what_is_missing"]) # plain English — what the user forgot print(rec["what_to_provide"]) # concrete example they can copy print(rec["why_it_matters"]) # what goes wrong if they skip it ``` -Gap names by intent type: - -- **Build:** `programming language`, `api structure`, `data model`, `integration specifics`, `authentication type`, `output format`, `task context` -- **Debug:** `error description`, `expected vs actual behavior`, `code context` -- **Optimization:** `optimization target`, `performance baseline`, `optimization constraint` -- **Explanation:** `code reference`, `explanation depth` -- **Feature:** `existing stack`, `feature scope`, `completion criteria` - --- ## Real-world examples ```python from inputguard import InputGuard + guard = InputGuard() # Build — vague -guard.analyze("build me an app") -# detected_intent: 'build' -# score: 60 -# status: usable_with_warnings -# gaps: ['programming language', 'output format'] +result = guard.analyze("build me an app") +assert result.detected_intent == "build" +assert result.clarity_score == 60 +assert result.status == "usable_with_warnings" +assert result.gaps == ["programming language", "output format"] # Debug — vague -guard.analyze("fix my code") -# detected_intent: 'debug' -# score: 35 -# status: needs_clarification -# gaps: ['error description', 'expected vs actual behavior', 'code context'] +result = guard.analyze("fix my code") +assert result.clarity_score == 35 +assert result.status == "needs_clarification" +assert result.gaps == ["error description", "expected vs actual behavior", "code context"] # Optimization — vague -guard.analyze("make this faster") -# detected_intent: 'optimization' -# score: 55 -# status: needs_clarification -# gaps: ['optimization target', 'performance baseline', 'optimization constraint'] +result = guard.analyze("make this faster") +assert result.clarity_score == 55 +assert result.status == "needs_clarification" +assert result.gaps == ["optimization target", "performance baseline", "optimization constraint"] # Build — fully specified -guard.analyze( +result = guard.analyze( "Build a REST API using FastAPI. " "Store users in PostgreSQL with fields: id, name, email. " "Expose GET /users and POST /users endpoints. " "Add JWT authentication." ) -# detected_intent: 'build' -# score: 100 -# status: ready -# gaps: [] +assert result.detected_intent == "build" +assert result.clarity_score == 100 +assert result.status == "ready" +assert result.gaps == [] # Debug — fully specified -guard.analyze( +result = guard.analyze( "Fix this Python function get_users() — it should return a list " "of user dicts but instead returns None. " "The error says: TypeError: NoneType is not iterable on line 45." ) -# detected_intent: 'debug' -# score: 100 -# status: ready -# gaps: [] +assert result.detected_intent == "debug" +assert result.clarity_score == 100 +assert result.status == "ready" +assert result.gaps == [] ``` --- ## Integration pattern -Drop it in front of your existing LLM call. Two minimal patterns: +Drop it in front of your existing LLM call: ```python from inputguard import InputGuard guard = InputGuard(mode="warning") +def call_your_llm(user_input: str) -> str: + return "..." # your real LLM call + def handle_user_input(user_input: str): result = guard.analyze(user_input) - if result.status in ("needs_clarification", "blocked"): - # Return feedback to the user before calling the LLM + if result.status in ("needs_clarification", "blocked", "degraded"): + # Surface the gaps and the questions that close them, + # instead of calling the LLM. `degraded` inputs arrive here + # in both modes; `blocked` only in strict mode. return { "status": result.status, "detected_intent": result.detected_intent, "gaps": result.gaps, + "follow_ups": result.follow_ups, "recommendations": result.recommendations, } # Input is clear enough — proceed to LLM return call_your_llm(user_input) + +payload = handle_user_input("fix my code") +assert payload["status"] == "needs_clarification" ``` For a hard gate, use `mode="strict"` and check `result.is_clear()`: @@ -428,13 +697,17 @@ def handle_user_input(user_input: str): if not result.is_clear(): return { "status": result.status, - "detected_intent": result.detected_intent, "gaps": result.gaps, - "recommendations": result.recommendations, + "follow_ups": result.follow_ups, } return call_your_llm(user_input) + +payload = handle_user_input("make this faster") +assert payload["status"] == "blocked" ``` +Framework-ready versions of this pattern — LangChain and LiteLLM, consuming `to_dict()` output — live in [docs/recipes/](docs/recipes/). The recipes are copy-paste patterns, not adapter packages: inputguard stays zero-dependency. + --- ## Package layout @@ -442,29 +715,27 @@ def handle_user_input(user_input: str): ``` inputguard/ ├── inputguard/ -│ ├── __init__.py -│ ├── analyzer.py -│ ├── detector.py -│ ├── recommender.py -│ ├── scorer.py -│ ├── types.py +│ ├── __init__.py # public exports +│ ├── analyzer.py # the analyze() pipeline +│ ├── cli.py # the inputguard console script +│ ├── detector.py # intent detection (coding signal chain) +│ ├── followups.py # per-gap clarifying questions +│ ├── language.py # script probe / multilingual degradation +│ ├── matching.py # shared word-boundary term matcher +│ ├── policy.py # Policy — calibration as data +│ ├── recommender.py # per-gap recommendations +│ ├── registry.py # Rule protocol, register_rule, register_domain +│ ├── scorer.py # scoring and status banding +│ ├── types.py # AnalysisResult, RuleFinding │ ├── py.typed -│ └── rules/ -│ ├── __init__.py -│ ├── coding.py -│ ├── debug.py -│ ├── optimization.py -│ ├── explanation.py -│ └── feature.py +│ └── rules/ # the 31 built-in rules +│ ├── coding.py # build, debug, optimization, explanation, feature +│ ├── writing.py # compose +│ └── data_analysis.py # analysis +├── eval/ # versioned 121-case clarity-evaluation set +├── docs/ # benchmark + framework recipes ├── tests/ -│ ├── test_coding.py -│ ├── test_detector.py -│ ├── test_debug.py -│ ├── test_optimization.py -│ ├── test_explanation.py -│ └── test_feature.py -├── pyproject.toml -└── README.md +└── pyproject.toml ``` --- @@ -473,9 +744,15 @@ inputguard/ ```bash pip install -e ".[dev]" -python -m pytest tests/ -v +python -m pytest tests/ +ruff check . +mypy # strict — the shipped py.typed is checked +python -m pytest tests/ --cov=inputguard --cov-branch --cov-fail-under=90 +python3 eval/measure_fp.py # clarity-evaluation benchmark (see docs/) ``` +CI runs the same gates on Python 3.9–3.13, plus a latency-benchmark job and a zero-dependency wheel check. + --- ## Publishing @@ -492,4 +769,4 @@ python -m twine upload dist/* MIT — see [LICENSE](LICENSE) for the full text. -Copyright © 2026 Nihanth Kalisetti. \ No newline at end of file +Copyright © 2026 Nihanth Kalisetti. diff --git a/tests/test_readme_examples.py b/tests/test_readme_examples.py new file mode 100644 index 0000000..2d813bd --- /dev/null +++ b/tests/test_readme_examples.py @@ -0,0 +1,164 @@ +"""README anti-drift gate: every example in README.md is executed. + +The v0.2 README shipped ``to_dict()`` examples that drifted from the code +within two releases (survey art_CnghyDxp §5.6) — the first docs a new user +read were wrong. This module makes that failure mode a test failure: + +1. Every ```python fenced block in README.md runs, in document order, in + one fresh subprocess. The README's code blocks carry real ``assert`` + statements for their documented values, so a stale value fails here. + The subprocess also isolates the extension examples' process-global + registry registrations from the modules this suite runs afterwards. +2. The embedded ```json block must equal the analyzer's live ``to_dict()`` + output for the same input — regenerated docs, never hand-copied. +3. The rule tables, gap-name lists, export list, and CLI output block must + match the live package: rule ids, severities, and gap strings per + intent, straight from the registry. +""" + +from __future__ import annotations + +import json +import re +import subprocess +import sys +from pathlib import Path +from typing import List + +import inputguard +from inputguard import REGISTRY + +README_PATH = Path(__file__).resolve().parent.parent / "README.md" + +_FENCE = re.compile(r"```([a-z]*)\n(.*?)```", re.DOTALL) + +# "What gets checked" section header -> the intent whose rules it documents. +_SECTION_INTENTS = [ + ("### Build inputs", "build"), + ("### Debug inputs", "debug"), + ("### Optimization inputs", "optimization"), + ("### Explanation inputs", "explanation"), + ("### Feature inputs", "feature"), + ("### Writing inputs", "compose"), + ("### Data-analysis inputs", "analysis"), +] + +# Bold label in the gap-vocabulary bullets -> intent id. +_GAP_LABELS = { + "Build": "build", + "Debug": "debug", + "Optimization": "optimization", + "Explanation": "explanation", + "Feature": "feature", + "Compose": "compose", + "Analysis": "analysis", +} + + +def _blocks(lang: str) -> List[str]: + """The README's fenced blocks in one language, in document order.""" + text = README_PATH.read_text(encoding="utf-8") + found = [match.group(2) for match in _FENCE.finditer(text) if match.group(1) == lang] + assert found, f"no {lang!r} blocks found in README.md — extraction rotted" + return found + + +def test_python_blocks_run_in_document_order() -> None: + blocks = _blocks("python") + # If this count drops, block extraction rotted — fix the extraction, not the docs. + assert len(blocks) >= 10 + script = "\n\n".join(blocks) + proc = subprocess.run( + [sys.executable, "-c", script], + capture_output=True, + text=True, + timeout=120, + ) + assert proc.returncode == 0, ( + "A README example failed — README and code have drifted.\n" + f"--- stdout ---\n{proc.stdout}\n--- stderr ---\n{proc.stderr}" + ) + + +def test_to_dict_json_block_matches_live_output() -> None: + blocks = _blocks("json") + assert len(blocks) == 1, "README should embed exactly one to_dict() JSON example" + documented = json.loads(blocks[0]) + live = inputguard.InputGuard().analyze("fix my code").to_dict() + assert documented == live + + +def test_public_exports_block_matches_package() -> None: + import_blocks = [b for b in _blocks("python") if "from inputguard import (" in b] + assert len(import_blocks) == 1, "README should show the public exports exactly once" + names = re.findall(r"^\s{4}(\w+),?$", import_blocks[0], re.MULTILINE) + assert set(names) == set(inputguard.__all__) + for name in names: + assert hasattr(inputguard, name), f"README export {name!r} does not exist" + + +def test_rule_tables_match_registry() -> None: + text = README_PATH.read_text(encoding="utf-8") + lines = text.splitlines() + for header, intent in _SECTION_INTENTS: + at = lines.index(header) # a vanished section raises — that is the failure + documented = {} + for line in lines[at + 1 :]: + if line.startswith("#"): + break + if line.startswith("| `"): + cells = [cell.strip() for cell in line.strip().strip("|").split("|")] + documented[cells[0].strip("`")] = cells[-1] + expected = {r.id: r.severity for r in REGISTRY.rules() if r.domain == intent} + assert documented == expected, ( + f"README rule table for {intent!r} drifted from the registry" + ) + registry_intents = {r.domain for r in REGISTRY.rules()} + assert registry_intents == {intent for _, intent in _SECTION_INTENTS}, ( + "a registry intent has no README rule-table section" + ) + + +def test_gap_name_lists_match_registry() -> None: + text = README_PATH.read_text(encoding="utf-8") + section = text.split("## The gap vocabulary", 1)[1].split("\n## ", 1)[0] + documented = {} + for line in section.splitlines(): + match = re.match(r"- \*\*(\w+):?\*\*\s*(.*)", line) + if not match: + continue + intent = _GAP_LABELS[match.group(1)] + documented[intent] = sorted(re.findall(r"`([^`]+)`", match.group(2))) + assert set(documented) == set(_GAP_LABELS.values()), "a gap list is missing from README" + for intent, gaps in documented.items(): + expected = sorted( + {r.gap for r in REGISTRY.rules() if r.domain == intent and r.gap is not None} + ) + assert gaps == expected, f"README gap list for {intent!r} drifted from the registry" + + +def test_cli_examples_behave_as_documented() -> None: + proc = subprocess.run( + [sys.executable, "-m", "inputguard.cli", "analyze", "make this faster", "--mode", "strict"], + capture_output=True, + text=True, + timeout=60, + ) + assert proc.returncode == 0 + text_blocks = _blocks("text") + assert len(text_blocks) == 1, "README should embed exactly one CLI output block" + assert proc.stdout.strip() == text_blocks[0].strip(), ( + "README CLI output drifted from the real CLI" + ) + + gated = subprocess.run( + [sys.executable, "-m", "inputguard.cli", "analyze", "build a REST API", "--min-score", "85"], + capture_output=True, + text=True, + timeout=60, + ) + assert gated.returncode == 1, "--min-score below the floor must exit 1" + + bash_blocks = _blocks("bash") + assert any("pip install inputguard" in block for block in bash_blocks) + assert any("inputguard analyze" in block for block in bash_blocks) From 5c696fd079c42426409bd27aaf92cfae3bebe95d Mon Sep 17 00:00:00 2001 From: Obvious Date: Thu, 17 Sep 2026 22:01:29 +0000 Subject: [PATCH 15/18] docs(integration-recipes): LangChain and LiteLLM recipes consuming to_dict() MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Copy-paste integration patterns over the to_dict() serialization boundary — preflight-and-gate and clarify-then-complete. The recipes are documentation only: framework packages stay out of core and out of the dependency tree. tests/test_recipes.py executes every recipe block in document order, stubbing the single framework idiom each wiring block uses, so the inputguard-side contract cannot drift silently. Co-authored-by: Kalisetti Nihanth Naidu --- docs/recipes/langchain.md | 82 ++++++++++++++++++++++++++++++++++++ docs/recipes/litellm.md | 87 ++++++++++++++++++++++++++++++++++++++ tests/test_recipes.py | 88 +++++++++++++++++++++++++++++++++++++++ 3 files changed, 257 insertions(+) create mode 100644 docs/recipes/langchain.md create mode 100644 docs/recipes/litellm.md create mode 100644 tests/test_recipes.py diff --git a/docs/recipes/langchain.md b/docs/recipes/langchain.md new file mode 100644 index 0000000..18640bb --- /dev/null +++ b/docs/recipes/langchain.md @@ -0,0 +1,82 @@ +# Recipe: InputGuard as a LangChain pre-flight gate + +InputGuard is framework-agnostic: it takes a string and returns a plain `to_dict()` payload, so wiring it into a LangChain chain is one `RunnableLambda` in front of your model call. There is no adapter package and no shared dependency — `inputguard` stays zero-dependency and imports nothing from `langchain-core`. + +> **Prerequisites:** `pip install inputguard` (and your own LangChain setup — inputguard installs none of it). This recipe targets langchain-core's runnable protocol; the same shape works for LCEL chains, agents, and LangGraph nodes. + +## 1 · The pre-flight payload (pure inputguard) + +This step is framework-free — a plain function returning the `to_dict()` contract your chain branches on. It is executed by CI (`tests/test_recipes.py`), including its assertions: + +```python +from inputguard import InputGuard + +guard = InputGuard() # warning mode: measure and flag, do not block + +def preflight(question: str) -> dict: + """Analyze input and return the to_dict() contract the chain routes on.""" + payload = guard.analyze(question).to_dict() + + if payload["status"] in ("needs_clarification", "blocked", "degraded"): + return { + "route": "clarify", + "question": question, + "gaps": payload["gaps"], + "follow_ups": payload["follow_ups"], + } + return {"route": "model", "question": question} + + +routed = preflight("fix my code") +assert routed["route"] == "clarify" +assert routed["gaps"] == ["error description", "expected vs actual behavior", "code context"] +assert routed["follow_ups"][0].startswith("What is the exact error message") + +specified = preflight( + "Fix this Python function get_users() — it should return a list " + "of user dicts but instead returns None. " + "The error says: TypeError: NoneType is not iterable on line 45." +) +assert specified["route"] == "model" +``` + +`degraded` rides the clarify route too: it is InputGuard reporting its own language limitation, not a judgment of the input — but the safest behavior is still to ask the user rather than run English-only analysis downstream. + +## 2 · Slot it into a chain + +The clarify path never touches the model — in the demo below, `call_your_model` is a `RuntimeError` to prove it: + +```python +from langchain_core.runnables import RunnableLambda + +def clarify_message(payload: dict) -> dict: + return { + **payload, + "message": "Before I can help, could you answer: " + " ".join(payload["follow_ups"]), + } + +def call_your_model(question: str) -> str: + # Replace with your real model invocation, e.g. ChatOpenAI(...).invoke(...) + raise RuntimeError("the clarify path must never reach the model") + +chain = RunnableLambda(preflight) | RunnableLambda( + lambda payload: clarify_message(payload) + if payload["route"] == "clarify" + else call_your_model(payload["question"]) +) + +result = chain.invoke("fix my code") +assert result["route"] == "clarify" +assert result["message"].startswith("Before I can help") +``` + +## Notes + +- **Warning vs strict:** in warning mode only `needs_clarification` and `degraded` arrive at the clarify branch. In strict mode `blocked` joins them — same route, more refusal. Pick via `InputGuard(mode=...)`; the payload contract does not change. +- **Where the score lives:** `payload["clarity_score"]` and `payload["score_breakdown"]` are on every result if you want to log the arithmetic or threshold on your side. +- **Follow-ups are ready to send:** `follow_ups` are plain sentences, deduped and gap-ordered — drop them into your clarification message verbatim. +- **Streaming/agents:** call `preflight()` before entering the agent loop; the payload is a plain dict, safe to attach to any run state. + +--- + +🔗 [Obvious Project](https://app.obvious.ai/p/inputguard-package-upgrade-plan-ZWAga4Fk) · 🧵 [Obvious Thread](https://app.obvious.ai/p/inputguard-package-upgrade-plan-ZWAga4Fk?thread=th_2ZBoxRug) diff --git a/docs/recipes/litellm.md b/docs/recipes/litellm.md new file mode 100644 index 0000000..bc55657 --- /dev/null +++ b/docs/recipes/litellm.md @@ -0,0 +1,87 @@ +# Recipe: InputGuard as a LiteLLM gate + +LiteLLM routes one OpenAI-shaped call to a hundred providers. InputGuard sits in front of that call: analyze the prompt first, and only forward to `litellm.completion(...)` when the payload says the input is worth the tokens. There is no adapter package and no shared dependency — `inputguard` stays zero-dependency and imports nothing from `litellm`. + +> **Prerequisites:** `pip install inputguard` (and your own LiteLLM setup — inputguard installs none of it). + +## 1 · The gate decision (pure inputguard) + +This step is framework-free — a strict-mode guard turns vague input into a refusal without a model call. It is executed by CI (`tests/test_recipes.py`), including its assertions: + +```python +from inputguard import InputGuard + +guard = InputGuard(mode="strict") # hard gate: refuse to forward vague input + +def guard_prompt(question: str) -> dict: + """Return {"ok": False, ...} to refuse, or {"ok": True, "prompt": ...} to forward.""" + payload = guard.analyze(question).to_dict() + + if payload["status"] != "ready": + return { + "ok": False, + "status": payload["status"], + "gaps": payload["gaps"], + "follow_ups": payload["follow_ups"], + } + return {"ok": True, "prompt": question} + + +blocked = guard_prompt("make this faster") +assert blocked["ok"] is False +assert blocked["status"] == "blocked" +assert blocked["gaps"] == [ + "optimization target", + "performance baseline", + "optimization constraint", +] +assert blocked["follow_ups"][0] == "Which function or module should get faster?" + +specified = guard_prompt( + "Fix this Python function get_users() — it should return a list " + "of user dicts but instead returns None. " + "The error says: TypeError: NoneType is not iterable on line 45." +) +assert specified["ok"] is True +``` + +A `degraded` payload also fails the gate (`"ready"` is required): non-English input gets an honest note back instead of a silent pass. + +## 2 · Wire it in front of the completion call + +The clarify path returns before `litellm.completion` is ever reached — in the CI run the stub raises if it is: + +```python +import litellm + +def answer(question: str) -> str: + decision = guard_prompt(question) + + if not decision["ok"]: + return ( + "I need more detail before I can help.\n" + "Missing: " + ", ".join(decision["gaps"]) + "\n" + "Could you answer: " + " ".join(decision["follow_ups"]) + ) + + response = litellm.completion( + model="gpt-4o-mini", # LiteLLM routes to 100+ providers behind this one call + messages=[{"role": "user", "content": decision["prompt"]}], + ) + return response.choices[0].message.content + + +reply = answer("make this faster") +assert reply.startswith("I need more detail") +``` + +## Notes + +- **Proxy deployments:** run the same check at the proxy edge — `guard_prompt` is a pure function, so a LiteLLM pre-call hook or a tiny FastAPI route in front of the proxy can reuse it verbatim. +- **Audit trail:** `payload["score_breakdown"]` gives you the per-gap arithmetic for logs; `payload["detected_intent"]` and `payload["detected_language"]` are useful routing metadata. +- **Softer rollout:** start with `InputGuard()` in warning mode and log `"ok": False` cases without refusing; flip to `mode="strict"` once you trust the measured false-positive rate (see [docs/false-positive-benchmark.md](../false-positive-benchmark.md)). +- **Cost math:** every refused call is a completion you never paid for; the gate costs one local, dependency-free `analyze()` pass. + +--- + +🔗 [Obvious Project](https://app.obvious.ai/p/inputguard-package-upgrade-plan-ZWAga4Fk) · 🧵 [Obvious Thread](https://app.obvious.ai/p/inputguard-package-upgrade-plan-ZWAga4Fk?thread=th_2ZBoxRug) diff --git a/tests/test_recipes.py b/tests/test_recipes.py new file mode 100644 index 0000000..8776fdb --- /dev/null +++ b/tests/test_recipes.py @@ -0,0 +1,88 @@ +"""docs/recipes example gate: the recipe code blocks execute. + +Each recipe is split into a framework-free block (pure inputguard, runs +verbatim) and a wiring block whose only framework usage is the runnable/ +completion idiom. The tests run every block in order inside a fresh +subprocess, with minimal stubs for the framework imports — proving the +inputguard-side contract the recipe relies on, without adding framework +dependencies to this repository. +""" + +from __future__ import annotations + +import re +import subprocess +import sys +from pathlib import Path + +RECIPES_DIR = Path(__file__).resolve().parent.parent / "docs" / "recipes" +_FENCE = re.compile(r"```python\n(.*?)```", re.DOTALL) + +# Minimal langchain-core runnable protocol: RunnableLambda wraps a callable, +# `|` composes pipelines left-to-right, .invoke() runs the chain. +_LANGCHAIN_STUB = """ +import types as _types + +class _Pipeline: + def __init__(self, steps): self._steps = steps + def __or__(self, other): return _Pipeline(self._steps + [other]) + def invoke(self, value): + for step in self._steps: + value = step.invoke(value) + return value + +class RunnableLambda: + def __init__(self, fn): self._fn = fn + def invoke(self, value): return self._fn(value) + def __or__(self, other): return _Pipeline([self, other]) + +_langchain_core = _types.ModuleType("langchain_core") +_runnables = _types.ModuleType("langchain_core.runnables") +_runnables.RunnableLambda = RunnableLambda +_langchain_core.runnables = _runnables +sys.modules["langchain_core"] = _langchain_core +sys.modules["langchain_core.runnables"] = _runnables +""" + +# The clarify path must return before the completion call — the stub raises +# if the recipe wiring ever reaches the model on the demo input. +_LITELLM_STUB = """ +import types as _types + +def _completion(*args, **kwargs): + raise AssertionError("litellm.completion called on the clarify path") + +litellm = _types.ModuleType("litellm") +litellm.completion = _completion +sys.modules["litellm"] = litellm +""" + +_RUNNERS = { + "langchain.md": _LANGCHAIN_STUB, + "litellm.md": _LITELLM_STUB, +} + + +def test_both_recipe_pages_exist_and_use_to_dict() -> None: + for name in _RUNNERS: + page = RECIPES_DIR / name + assert page.is_file(), f"missing recipe page docs/recipes/{name}" + assert "to_dict()" in page.read_text(encoding="utf-8") + + +def test_recipe_blocks_execute_in_document_order() -> None: + for name, stub_setup in _RUNNERS.items(): + page = RECIPES_DIR / name + blocks = _FENCE.findall(page.read_text(encoding="utf-8")) + assert len(blocks) >= 2, f"{name}: expected preflight + wiring blocks" + script = "import sys\n" + stub_setup + "\n" + "\n\n".join(blocks) + proc = subprocess.run( + [sys.executable, "-c", script], + capture_output=True, + text=True, + timeout=120, + ) + assert proc.returncode == 0, ( + f"A docs/recipes/{name} example failed.\n" + f"--- stdout ---\n{proc.stdout}\n--- stderr ---\n{proc.stderr}" + ) From 8efb9f759adb9cbad73432e6605a8aa511ada9af Mon Sep 17 00:00:00 2001 From: Obvious Date: Thu, 17 Sep 2026 22:02:47 +0000 Subject: [PATCH 16/18] docs(fp-benchmark): refresh measured rates and state the boundary target honestly MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Re-ran eval/measure_fp.py on the release branch: overall 116/121 match, FP 3.3% / FN 0.8%, zero false positives on true negatives and degradation rows, notes on 14/14 degradation rows, PF rows 19-20 ms. Adds the labeling-guide target contrast the doc lacked: fixture-style boundary FP measured 4/6 (v0.2 baseline 6/6, target 0/6) — driven by the matcher's deliberate inflection tolerance, documented as an eval-driven trade-off for a future release, not a label edit. Co-authored-by: Kalisetti Nihanth Naidu --- docs/false-positive-benchmark.md | 53 ++++++++++++++++++++++++-------- 1 file changed, 41 insertions(+), 12 deletions(-) diff --git a/docs/false-positive-benchmark.md b/docs/false-positive-benchmark.md index b2e02d4..f65d773 100644 --- a/docs/false-positive-benchmark.md +++ b/docs/false-positive-benchmark.md @@ -19,7 +19,7 @@ Last run: v0.3.0 release branch, 2026-09-17. | **overall** | **121** | **116** | **4** | **1** | **3.3%** | **0.8%** | Degradation honesty: a degradation note is present on 14 of 14 degradation -rows. Performance wall time (single pass): PF-001 22 ms, PF-002 22 ms — +rows. Performance wall time (single pass): PF-001 19 ms, PF-002 20 ms — PF-002 exercises the 10,000-character cap with the `truncated` flag set. Zero false positives on true negatives and zero on degradation rows is the @@ -27,21 +27,50 @@ load-bearing number: it says the English keyword rules fire only on English input, and that non-English input degrades instead of producing invented gaps. +## Fixture-style boundary target: 4 of 6, against v0.2's 6 of 6 + +The labeling guide (`art_XPvHhPeZ`) sets the release target: the six +fixture-style boundary rows (BD-002, BD-003, BD-004, BD-006, BD-013, +BD-014) must come back clean, with BD-012 — the genuine-debug-request +no-regression guard — still passing. The v0.2 baseline flagged all six +(substring matching read "fixture" as "fix"). + +Measured at this release: **2 of 6 pass cleanly (BD-003, BD-013), and +BD-012 passes**, but 4 of 6 still flag (BD-002, BD-004, BD-006, BD-014). +The 0/6 target is therefore **not met** — the honest read is "improved from +6/6 to 4/6 flagged, short of the 0/6 target." + +The mechanism is deliberate, verified in `inputguard/matching.py`: the +shared matcher kills the embedded-word false positives ("fixture" is no +longer "fix"), but it still absorbs common inflectional endings +(`-s`, `-ed`, `-ing`, `-er`, `-ly`, ...) so natural word forms keep firing +— "debugged" still matches "debug", "slowly" still matches "slow". The +class-of-hit the v0.2 coding matcher preserved was read as recall on +real inputs; these four labels were written against strict +boundary-only semantics and sit exactly on that trade-off. Closing them +means re-deciding that recall trade-off on the labeled set — an +eval-driven change for a future release, never a label edit. + +For the spec verification row "False positive eliminated", this run +records: v0.2 baseline 6/6 flagged; v0.3 measured 4/6 flagged with +BD-012 (no-regression guard) passing and zero false positives on all 37 +true negatives. + ## Residual mismatches (5) -The 5 unmatched rows are all pre-existing behavior on the coding domain, -known at label time and outside the v0.3 feature work: +The 5 unmatched rows, per the measured run at the top of this document: -- `TP-FEA-02` — false negative: `feature scope` gap not flagged; the row - comes back one status above expected (`usable_with_warnings` vs `ready`). -- `BD-002`, `BD-004`, `BD-006` — boundary rows: the detected intent - switches (build → optimization / debug) and the coding rules of the other - intent fire, adding spurious gaps. -- `BD-014` — boundary row with the same intent-adjacency shape. +- `TP-FEA-02` — false negative: the `feature scope` gap is not flagged; the + row lands one status above expected (`usable_with_warnings` vs `ready`). +- `BD-002`, `BD-004`, `BD-006`, `BD-014` — fixture-style boundary rows + still flagging through the matcher's deliberate inflection tolerance + (see the target section above): the detected intent flips + (build → optimization / debug) and the other intent's coding rules fire, + adding spurious gaps. -These are candidate labels to re-examine or intent-detector work for a -future release; per `eval/README.md`, labels are never edited to make a -measurement look better. +These are candidate re-evaluations of the inflection-recall trade-off, or +intent-detector work, for a future release; per `eval/README.md`, labels +are never edited to make a measurement look better. ## History within the release wave From 4cb106b4ac31e382659077d96b949bd5ec0fb27c Mon Sep 17 00:00:00 2001 From: Obvious Date: Thu, 17 Sep 2026 22:04:03 +0000 Subject: [PATCH 17/18] chore(release-0.3.0): complete the 0.3.0 changelog for the docs wave MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Version 0.3.0 and the Development Status :: 4 - Beta classifier were already in pyproject.toml and inputguard.__version__ — verified, no bump needed. This commit completes the 0.3.0 entry with the docs wave: tested README rewrite, integration recipes, and the measured false-positive benchmark with the boundary-target contrast. Co-authored-by: Kalisetti Nihanth Naidu --- CHANGELOG.md | 22 ++++++++++++++++++++++ 1 file changed, 22 insertions(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index c67387a..55cf6c3 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -60,6 +60,28 @@ in `docs/false-positive-benchmark.md`. project URLs, and the explicit dev extras / tool config (`[tool.ruff]`, `[tool.mypy]`, `[tool.pytest.ini_options]`). +### Docs +- README rewritten for the v0.3 surface — public exports (including the + extension API), Policy customization, custom rule/domain registration, + follow-up questions, multilingual degradation, and CLI usage — with the + v0.2 example-drift failure mode closed structurally: every README + example is executed by `tests/test_readme_examples.py`, the embedded + `to_dict()` JSON is regenerated from live analyzer output, and the + export list, rule tables, gap vocabulary, and CLI output are checked + against the registry. +- Framework integration recipes: `docs/recipes/langchain.md` and + `docs/recipes/litellm.md`, copy-paste patterns over the `to_dict()` + serialization boundary. Framework packages stay out of core; the recipe + blocks are example-tested with framework stubs in + `tests/test_recipes.py`. +- False-positive benchmark documentation refreshed from a measured run of + `eval/measure_fp.py` on this release state: 116/121 overall, FP 3.3% / + FN 0.8%, zero false positives on true negatives and degradation rows. + The labeling-guide boundary target is stated honestly: 4 of 6 + fixture-style rows still flag (v0.2 baseline 6/6, target 0/6) through + the matcher's deliberate inflection tolerance — documented as an + eval-driven trade-off for a future release, not a label edit. + ### Changed - All term matching now happens at word boundaries (`#6`) — detector and every rule module share one matcher, so "my_error" matches but "error" From 35c984c9881f7a0a817ebd99d46762752f84e8ce Mon Sep 17 00:00:00 2001 From: Obvious Date: Thu, 17 Sep 2026 22:06:21 +0000 Subject: [PATCH 18/18] fix(readme-examples): 3.9-compatible annotation in the custom-rule example MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The custom-rule example annotated check() with PEP 604 union syntax (-> RuleFinding | None), which is evaluated at definition time and raises TypeError on Python 3.9 — caught by the CI matrix's 3.9 job running the README anti-drift gate. Use typing.Optional instead; an AST scan of every executed README block confirms no other runtime-evaluated 3.9-incompatible construct remains. Co-authored-by: Kalisetti Nihanth Naidu --- README.md | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/README.md b/README.md index cdb1e6e..b3254a6 100644 --- a/README.md +++ b/README.md @@ -202,6 +202,8 @@ assert "Portuguese" in result.degradation_note The extension contract is four members and one method. Built-in rules register through the exact same path — the API is exercised by all 31 built-in rules before anyone writes their own. ```python +from typing import Optional + from inputguard import RuleFinding, register_domain class CheckRollbackPlan: @@ -210,7 +212,7 @@ class CheckRollbackPlan: severity = "high" # "low" | "medium" | "high" — validated gap = "rollback plan" # groups findings for scoring dedup - def check(self, text: str) -> RuleFinding | None: + def check(self, text: str) -> Optional[RuleFinding]: # text arrives normalized: lowercased, whitespace-collapsed. if "rollback" not in text and "roll back" not in text: return RuleFinding(