Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
21 changes: 21 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
name: CI

on:
push:
branches: [main]
pull_request:

jobs:
test:
runs-on: ubuntu-latest
strategy:
matrix:
# 3.9 is the package floor (requires-python), 3.13 is current.
python-version: ["3.9", "3.13"]
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: ${{ matrix.python-version }}
- run: pip install -e ".[dev]"
- run: pytest -q
63 changes: 63 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,68 @@
# Changelog

## [0.3.0] — 2026-09-17
The extension release: register your own rules and domains, get follow-up
questions for every gap, calibrated policy control, and honest behavior on
non-English input. Evaluated at **116 of 121** cases on the versioned
clarity-evaluation set in `eval/`; the 5 residual mismatches are documented
in `docs/false-positive-benchmark.md`.

### Added
- Typed `Rule` protocol and in-process registry (`#3`): first-party and
third-party rules enter through the same path; unknown severities fail
loudly at score time instead of being swallowed.
- Registry contract enforcement (`#8`): rules must expose a `check(text)`
signature, unique `(intent, id)` pairs, and pass registration guards —
a mis-registered rule aborts instead of silently never firing.
- Per-gap follow-up questions engine (`#4`): every built-in gap carries one
or two templated clarifying questions, surfaced on the result as the
additive `follow_ups` field. `{function}` / `{dataset}` slots fill from
the original input; a gap with no table entry still gets the documented
fallback question — never silence.
- Policy calibration (`#9`): a frozen, validated `Policy` with the v0.2
constants as defaults — status bands, severity penalties, rule filters,
`min_words`, allowlist patterns, and the 10,000-character input cap with
a visible `truncated` flag — plus a per-result `score_breakdown`.
- First-party writing domain (`#10`): six rules, the `compose` intent,
recommendations, and follow-ups for essay/report/email prompts.
- First-party data-analysis domain (`#11`): six rules (dataset/source,
question/goal, deliverable format, tooling, volume, reproducibility),
the `analysis` intent, recommendations, and follow-ups.
- The 121-case clarity-evaluation set, versioned as `eval/` (`#7`), with a
measured false-positive benchmark in `docs/false-positive-benchmark.md`.
- Multilingual degradation (`#5`, `#12`): a zero-dependency script probe
(unicodedata histogram, `inputguard/language.py`) classifies each input
before rules run. Inputs the English heuristics cannot assess take an
explicit degraded path — rules skipped, a 20-point confidence penalty,
`detected_language`, `heuristic_coverage`, `degradation_note`, and
`detected_intent: undetermined` instead of spurious gaps or a silent
100/ready. Four inputs degrade: uncovered dominant scripts, and — from
`#12` — Latin-script text recognized as French/Spanish/Portuguese by a
function-word layer with an English margin, and Latin-dominant text
carrying a run of 3+ consecutive uncovered-script letters (mixed
English+Han). English prompts with loanwords, URLs, or name collisions
are pinned unchanged by tests.

### Changed
- All term matching now happens at word boundaries (`#6`) — detector and
every rule module share one matcher, so "my_error" matches but "error"
inside "terrorist" no longer counts as a mention.
- Degraded results report the literal `degraded` status in both modes.
Mapping the degradation penalty through the ordinary banding returned
`usable_with_warnings` / `needs_clarification`, reading as an ordinary
vagueness verdict about the input; degradation is a language limitation
of the tool and now says so.

### Fixed
- The dataset-filename regex behind the `{dataset}` follow-up slot (and the
data-analysis dataset rule) backtracked its greedy span against every dot
in a filename-like run — quadratic, measured at 13 ms per 1 K chars and
1.3 s per 10 K. Both sites now scan maximal filename runs linearly with
the extension checked in Python; 10 K now takes under a millisecond, and
timing tests at 1 K and 10 K pin the growth rate.
- Degraded-path results preserve the `truncated` flag, so a capped input
that also degrades reports both honestly.

## [0.2.0] — 2026-05-29
### Added
- Auto intent detection. `.analyze()` now detects whether the input
Expand Down
81 changes: 79 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -90,11 +90,55 @@ The clarity score is mode-independent. Only the status threshold changes.
| `usable_with_warnings` | score 60–84 | never |
| `needs_clarification` | score < 60 | score 65–84 |
| `blocked` | never | score < 65 |
| `degraded` | never (language limitation — see Non-English input) | never (language limitation) |

Use `warning` when you want to surface gaps to the user without blocking. Use `strict` when you want to refuse to forward vague input to the LLM.

---

## Non-English input

InputGuard's rules are English-language heuristics. Before any rule runs, a zero-dependency script probe (stdlib `unicodedata` only) classifies the input's script and language. Inputs the rules cannot assess take the explicit degraded path — an uncovered dominant script, Latin-script text recognized as French/Spanish/Portuguese by its function words, or a Latin-dominant input carrying a run of 3+ consecutive uncovered-script letters. In every case InputGuard says so instead of pretending:

```python
result = guard.analyze("建造一个用户登录应用")

result.status # 'degraded' — never 'ready', in either mode
result.clarity_score # 80 (100 minus the degradation penalty)
result.detected_intent # 'undetermined'
result.detected_language # 'zh' (coarse, script-derived guess)
result.heuristic_coverage # 'none'
result.degradation_note # explains that rules were skipped and why
```

The rules are **skipped explicitly** — running English keyword rules on text they cannot assess would produce a silent, unearned verdict or spurious gaps invented out of the silence. A degraded result reports the literal `degraded` status in both modes: it is the tool reporting a language limitation of itself, not a judgment of the input's clarity (strict mode's banding would otherwise read as an ordinary critique). In v0.2 this input silently scored 100/ready; v0.3 refuses to assert a confidence it does not have.

Four additive fields on the result carry the probe's verdict:

| Field | Values |
|---|---|
| `detected_language` | coarse script-derived guess (`'en'`, `'zh'`, `'ja'`, `'ko'`, `'ru'`, `'ar'`, `'fr'`, `'es'`, `'pt'`, ...; `'und'` when unclassifiable) |
| `heuristic_coverage` | `'full'` (≥ 70% of letters covered — rules run exactly as before), `'partial'` (50–70% — rules run, note flags the uncovered remainder), `'none'` (degraded path), `'unknown'` (no letters to classify) |
| `degradation_note` | `None`, or an explanation of what was skipped and why |
| `truncated` | `True` when the input was capped to the 10,000-character limit (preserved even on the degraded path) |

The degraded path covers three shapes of input the English rules cannot assess:

- **Uncovered script** — the dominant script (Cyrillic, Arabic, Han, ...) has no heuristic coverage.
- **Non-English Latin** — French, Spanish, and Portuguese text is 100% Latin script yet just as unreadable to English-only rules, which would find nothing and invent gaps out of the silence. A function-word layer recognizes those three languages — enough distinct stop-word hits with a margin over the input's English function-word evidence — and degrades them like any other uncovered language. German, Italian, Dutch, and other Latin-script languages are a documented blind spot and still pass as before. Accented English (`café`) and short telegraphic prompts (`build todo api`) still run the rules.
- **Mixed scripts** — a run of 3+ consecutive letters in an uncovered script degrades even a majority-English input ("Fix this bug 修复这个错误 in the payment flow"): that clause is content the rules cannot audit. A short borrow like `"make it faster 这个"` still rides along, and input at 50–70% coverage keeps the `partial` path with its note.

```python
result = guard.analyze("Preciso de um aplicativo web com login de usuário e relatórios")

result.detected_language # 'pt' (function-word guess)
result.heuristic_coverage # 'none'
result.status # 'degraded' — never 'ready'
result.degradation_note # names Portuguese (pt) and explains the skip
```

---

## How intent detection works

InputGuard automatically detects what kind of coding input it is receiving. No extra parameters needed. The same `.analyze()` call handles all five intent types.
Expand Down Expand Up @@ -177,6 +221,32 @@ Requests like "add search to my existing app", "extend my current API with pagin
| `missing_feature_scope` | Feature requested but no definition of what it should specifically do. | high |
| `missing_completion_criteria` | No definition of what done looks like for this feature. | low |

### Writing inputs

Requests like "write a blog post", "draft an email", "proofread my essay". Every gap carries its own follow-up questions, and each gap names one thing at a time — same one-gap-one-question discipline as the coding domains.

| Rule code | What it catches | Severity |
|---|---|---|
| `missing_audience` | Writing task detected but no audience or reader is specified. | high |
| `missing_purpose` | The goal — what the piece should accomplish — is not stated. | high |
| `missing_structure_format` | No length or organization guidance is provided. | medium |
| `missing_source_material` | Existing material is referenced but not provided — paste or attach the text to work from. | high |
| `missing_writing_context` | The subject or situation is not named. | medium |
| `missing_completeness` | No required content or constraints are specified. | low |

### Data-analysis inputs

Requests like "analyze my sales data", "build a dashboard", "report on this spreadsheet".

| Rule code | What it catches | Severity |
|---|---|---|
| `missing_dataset_source` | Analysis requested but no dataset, file, or source is named. | high |
| `missing_question_goal` | Data present but no question or goal for the analysis is stated. | high |
| `missing_deliverable_format` | No chart, table, summary, or report named as the output. | medium |
| `missing_tooling` | No tool or stack (pandas, SQL, Excel...) specified. | medium |
| `missing_volume` | No sense of data size or scope given. | low |
| `missing_reproducibility` | No refresh/reproducibility expectation stated. | low |

---

## The result object
Expand All @@ -185,13 +255,20 @@ Requests like "add search to my existing app", "extend my current API with pagin

| Field | Type | Description |
|---|---|---|
| `status` | `str` | One of `"ready"`, `"usable_with_warnings"`, `"needs_clarification"`, `"blocked"` |
| `status` | `str` | One of `"ready"`, `"usable_with_warnings"`, `"needs_clarification"`, `"blocked"`, `"degraded"` (language limitation — see Non-English input) |
| `clarity_score` | `int` | 0 to 100 |
| `detected_intent` | `str` | Which intent was detected: `build`, `debug`, `optimization`, `explanation`, or `feature` |
| `detected_intent` | `str` | Which intent was detected: `build`, `debug`, `optimization`, `explanation`, `feature`, `compose`, or `analysis` |
| `gaps` | `List[str]` | Gap names, in the order rules fired |
| `recommendations` | `List[dict]` | One dict per gap (see next section) |
| `follow_ups` | `List[str]` | One or two clarifying questions per gap, ready to send back to the user |
| `findings` | `List[RuleFinding]` | Raw rule findings (code, message, severity, gap) |
| `interpretation_note` | `Optional[str]` | Set when the input is highly ambiguous (score < 50 or two or more high-severity findings) |
| `detected_language` | `str` | Coarse script-derived language guess (`'en'`, `'zh'`, ...; `'und'` when unclassifiable) |
| `heuristic_coverage` | `str` | `'full'`, `'partial'`, `'none'`, or `'unknown'` — how much of the input the English rules could see |
| `degradation_note` | `Optional[str]` | Set when rule analysis was skipped or limited by language coverage |
| `borderline` | `bool` | `True` when the score sits in the near-miss band just below ready — worth one more pass |
| `truncated` | `bool` | `True` when the input was capped at 10,000 characters (only the prefix was analyzed) |
| `score_breakdown` | `Optional[dict]` | Per-contribution score arithmetic, present when the policy exposes it |

Helpers:

Expand Down
78 changes: 78 additions & 0 deletions docs/false-positive-benchmark.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,78 @@
# False-positive benchmark

Measured results of the clarity-evaluation set (`eval/`, 121 labeled rows)
against the shipped analyzer, produced by `eval/measure_fp.py`. This is the
document `eval/README.md` references as the release-wave record; re-run the
script and update the tables when the analyzer or the labels change.

Last run: v0.3.0 release branch, 2026-09-17.

## Headline numbers

| Case type | Rows | Match | FP | FN | FP rate | FN rate |
| -------------- | ---- | ----- | -- | -- | ------- | ------- |
| true_positive | 43 | 42 | 0 | 1 | 0.0% | 2.3% |
| true_negative | 37 | 37 | 0 | 0 | 0.0% | 0.0% |
| boundary | 25 | 21 | 4 | 0 | 16.0% | 0.0% |
| degradation | 14 | 14 | 0 | 0 | 0.0% | 0.0% |
| performance | 2 | 2 | 0 | 0 | 0.0% | 0.0% |
| **overall** | **121** | **116** | **4** | **1** | **3.3%** | **0.8%** |

Degradation honesty: a degradation note is present on 14 of 14 degradation
rows. Performance wall time (single pass): PF-001 22 ms, PF-002 22 ms —
PF-002 exercises the 10,000-character cap with the `truncated` flag set.

Zero false positives on true negatives and zero on degradation rows is the
load-bearing number: it says the English keyword rules fire only on English
input, and that non-English input degrades instead of producing invented
gaps.

## Residual mismatches (5)

The 5 unmatched rows are all pre-existing behavior on the coding domain,
known at label time and outside the v0.3 feature work:

- `TP-FEA-02` — false negative: `feature scope` gap not flagged; the row
comes back one status above expected (`usable_with_warnings` vs `ready`).
- `BD-002`, `BD-004`, `BD-006` — boundary rows: the detected intent
switches (build → optimization / debug) and the coding rules of the other
intent fire, adding spurious gaps.
- `BD-014` — boundary row with the same intent-adjacency shape.

These are candidate labels to re-examine or intent-detector work for a
future release; per `eval/README.md`, labels are never edited to make a
measurement look better.

## History within the release wave

Measured with the same script, no label edits in between:

| State | Match | Note |
| --------------------------------------------- | ------ | ----------------------------------------------- |
| Early v0.3 branch, before writing/analysis rules | 87/121 | 15 data-analysis rows unevaluated (rules absent) |
| After data-analysis rules landed (PR #11) | 102/121 | 14 degradation rows still mismatching |
| Degradation contract completed (this release) | 116/121 | all 14 degradation rows match, notes on 14/14 |

## Performance: dataset-filename extraction

The `{dataset}` follow-up slot and the data-analysis dataset rule both
extract filename-like runs. The original single regex backtracked its
greedy span against every dot in a run — quadratic. Measured on dotted
filler (the adversarial shape):

| Input length | Old regex | Current extractor |
| ------------ | --------- | ----------------- |
| 1,000 chars | 13 ms | < 1 ms |
| 10,000 chars | 1,306 ms | 0.7 ms (plain) / 21 ms (all dots) |
| 100,000 chars | unbounded growth | 130 ms |

Timing tests at 1 K and 10 K (`tests/test_followups.py`) pin the linear
growth rate.

## How to reproduce

```bash
python3 eval/measure_fp.py # full 121-row table
python3 eval/measure_fp.py --json # machine-readable
python3 -m pytest tests/ # full suite including timing guards
```
Loading
Loading