Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
19 commits
Select commit Hold shift + click to select a range
94508cb
chore: release branch for InputGuard v0.3 production upgrade
ObviousApp Sep 17, 2026
efa8c63
feat: typed Rule protocol, in-process registry, and loud scorer failu…
obvious-autobuild[bot] Sep 17, 2026
38a9e56
feat: per-gap follow-up questions engine with additive follow_ups res…
obvious-autobuild[bot] Sep 17, 2026
7f15ed0
feat: multilingual degradation — script probe and explicit degraded p…
obvious-autobuild[bot] Sep 17, 2026
c52f485
feat(eval-prompt-set): version the 121-case clarity-evaluation set as…
obvious-autobuild[bot] Sep 17, 2026
c2ccfd4
fix: enforce registry contract — check(text) signature, unique intent…
obvious-autobuild[bot] Sep 17, 2026
68f2c30
fix: match terms at word boundaries across detector and all rule modu…
obvious-autobuild[bot] Sep 17, 2026
bee08f0
feat(policy): frozen validated Policy with pinned v0.2 defaults, band…
obvious-autobuild[bot] Sep 17, 2026
1d5c8c9
feat: first-party writing domain — six rules, compose intent, complet…
obvious-autobuild[bot] Sep 17, 2026
8cfd558
feat: first-party data-analysis domain — six rules, analysis intent, …
obvious-autobuild[bot] Sep 17, 2026
4112d8f
fix: degradation coverage for DG-011..014 (detector) (#12)
obvious-autobuild[bot] Sep 17, 2026
70f41ab
fix: complete v0.3.0 release sweep — version, cap, degradation, docs, CI
ObviousApp Sep 17, 2026
850401f
feat: CI matrix, latency budgets, CLI, and packaging metadata for v0.…
obvious-autobuild[bot] Sep 17, 2026
61d21f5
merge: reconcile main (squash of 70f41ab) into release/v0.3.0 — relea…
ObviousApp Sep 17, 2026
0f4c043
docs(readme-tested-examples): rewrite README for v0.3 with every exam…
ObviousApp Sep 17, 2026
5c696fd
docs(integration-recipes): LangChain and LiteLLM recipes consuming to…
ObviousApp Sep 17, 2026
8efb9f7
docs(fp-benchmark): refresh measured rates and state the boundary tar…
ObviousApp Sep 17, 2026
4cb106b
chore(release-0.3.0): complete the 0.3.0 changelog for the docs wave
ObviousApp Sep 17, 2026
35c984c
fix(readme-examples): 3.9-compatible annotation in the custom-rule ex…
ObviousApp Sep 17, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
138 changes: 133 additions & 5 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -2,20 +2,148 @@ name: CI

on:
push:
branches: [main]
branches: [main, "release/**"]
pull_request:

permissions:
contents: read

jobs:
# The quality gate: lint, tests with a 90% branch-coverage floor, and
# strict typing across the full supported Python matrix (spec §8).
test:
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
# 3.9 is the package floor (requires-python), 3.13 is current.
python-version: ["3.9", "3.13"]
python-version: ["3.9", "3.10", "3.11", "3.12", "3.13"]
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: ${{ matrix.python-version }}
- run: pip install -e ".[dev]"
- run: pytest -q
- name: Install with dev tooling
run: |
python -m pip install --upgrade pip
pip install -e ".[dev]"
- name: Ruff
run: ruff check inputguard tests eval
- name: Run tests with 90% branch-coverage floor
# Latency asserts are excluded here and gate only in the dedicated
# benchmark job: coverage tracing roughly triples per-call cost
# (observed on run 35277175054 — debug p50 21.14 ms vs 6.59 ms
# untraced), so gating latency under coverage double-gates the same
# budgets under different conditions.
run: |
pytest -q --cov=inputguard --cov-branch --cov-fail-under=90 --strict \
--ignore=tests/test_benchmarks.py
- name: Strict typing (checks the shipped py.typed)
run: mypy --strict inputguard

# Latency budgets are gates, not notes: a regression here is a production
# failure mode (the v0.2 unbounded 3.8 s scan). Measured p50/p99 print to
# the job log; budgets and their derivation live in tests/test_benchmarks.py.
benchmark:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.13"
- name: Install with dev tooling
run: |
python -m pip install --upgrade pip
pip install -e ".[dev]"
- name: Latency budgets (p50/p99 recorded)
run: pytest -v -s tests/test_benchmarks.py

# The zero-dependency promise, verified against the built artifact — not
# the source tree: wheel METADATA carries no Requires-Dist, a bare venv
# install pulls no third-party package, and the public API + CLI work from
# the installed wheel alone. The smoke steps run from a scratch directory
# so the repo's local `inputguard/` package can never shadow the wheel.
wheel-zero-dep:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.13"
- name: Build wheel
run: |
python -m pip install --upgrade pip build
python -m build --wheel
- name: Wheel metadata declares no unconditional runtime dependencies
run: |
python - <<'PY'
import glob
import zipfile

wheel = glob.glob("dist/*.whl")[0]
names = zipfile.ZipFile(wheel).namelist()
meta_name = next(n for n in names if n.endswith(".dist-info/METADATA"))
meta = zipfile.ZipFile(wheel).read(meta_name).decode()
requires = [
line
for line in meta.splitlines()
if line.startswith("Requires-Dist:")
]
unconditional = [
line for line in requires if 'extra == "' not in line
]
assert not unconditional, (
"wheel METADATA declares unconditional runtime dependencies — "
f"the zero-dep promise is broken: {unconditional}"
)
extra_gated = len(requires)
print(
f"{wheel}: 0 unconditional runtime dependencies "
f"({extra_gated} extra-gated dev-only entries)"
)
PY
- name: Install wheel into a bare venv
run: |
python -m venv bare
./bare/bin/pip install --upgrade pip
./bare/bin/pip install dist/*.whl
- name: Public API works from the installed wheel
run: |
cd "$(mktemp -d)"
"$GITHUB_WORKSPACE/bare/bin/python" - <<'PY'
import json

import inputguard

assert inputguard.__version__
guard = inputguard.InputGuard()
result = guard.analyze("make this faster", domain="coding")
assert isinstance(result.clarity_score, int)
assert 0 <= result.clarity_score <= 100
payload = result.to_dict()
json.dumps(payload)
print(f"bare-venv analyze OK: {result.status} score={result.clarity_score}")
PY
- name: Bare venv contains no third-party packages
run: |
third_party=$(./bare/bin/pip list --format=freeze \
| cut -d= -f1 | tr '[:upper:]' '[:lower:]' \
| grep -v -E '^(inputguard|pip|setuptools)$' || true)
if [ -n "$third_party" ]; then
echo "Third-party packages present after wheel install: $third_party"
exit 1
fi
echo "bare venv holds only inputguard (+ venv tooling)"
- name: CLI works from the installed wheel
run: |
cd "$(mktemp -d)"
INPUTGUARD="$GITHUB_WORKSPACE/bare/bin/inputguard"
"$INPUTGUARD" analyze "make this faster" --format json
set +e
"$INPUTGUARD" analyze "make this faster" --min-score 85
code=$?
set -e
if [ "$code" -ne 1 ]; then
echo "expected exit 1 below the min-score floor, got $code"
exit 1
fi
echo "CLI exit codes OK (0 on analysis, 1 below the min-score floor)"
39 changes: 39 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,6 +42,45 @@ in `docs/false-positive-benchmark.md`.
carrying a run of 3+ consecutive uncovered-script letters (mixed
English+Han). English prompts with loanwords, URLs, or name collisions
are pinned unchanged by tests.
- CI as production gates (`.github/workflows/ci.yml`): a Python 3.9–3.13
matrix running ruff, pytest with a 90% branch-coverage floor
(`--cov-branch --cov-fail-under=90 --strict`, latency asserts excluded
from the traced pass), and mypy `--strict` (the shipped `py.typed` is
now actually checked); a dedicated untraced latency-benchmark job with
budgets measured on this codebase (10k-char build-intent p50 ≤ 65 ms
including the catch-all re-run path, debug and ready paths ≤ 25 ms; a
1.6 MB input bounded by the policy cap); and a zero-dependency wheel
gate that builds the wheel, installs it into a bare venv, and asserts
no third-party package is present.
- `inputguard` CLI (`inputguard analyze`) via a console-script entry
point: argparse, `--format json` emitting the same `to_dict()`
contract, and CI-friendly exit codes (0 ok, 1 below `--min-score`,
2 usage error).
- Packaging metadata: 3.13 classifier, Development Status → 4 - Beta,
project URLs, and the explicit dev extras / tool config
(`[tool.ruff]`, `[tool.mypy]`, `[tool.pytest.ini_options]`).

### Docs
- README rewritten for the v0.3 surface — public exports (including the
extension API), Policy customization, custom rule/domain registration,
follow-up questions, multilingual degradation, and CLI usage — with the
v0.2 example-drift failure mode closed structurally: every README
example is executed by `tests/test_readme_examples.py`, the embedded
`to_dict()` JSON is regenerated from live analyzer output, and the
export list, rule tables, gap vocabulary, and CLI output are checked
against the registry.
- Framework integration recipes: `docs/recipes/langchain.md` and
`docs/recipes/litellm.md`, copy-paste patterns over the `to_dict()`
serialization boundary. Framework packages stay out of core; the recipe
blocks are example-tested with framework stubs in
`tests/test_recipes.py`.
- False-positive benchmark documentation refreshed from a measured run of
`eval/measure_fp.py` on this release state: 116/121 overall, FP 3.3% /
FN 0.8%, zero false positives on true negatives and degradation rows.
The labeling-guide boundary target is stated honestly: 4 of 6
fixture-style rows still flag (v0.2 baseline 6/6, target 0/6) through
the matcher's deliberate inflection tolerance — documented as an
eval-driven trade-off for a future release, not a label edit.

### Changed
- All term matching now happens at word boundaries (`#6`) — detector and
Expand Down
Loading
Loading