Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
31 changes: 24 additions & 7 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -87,7 +87,7 @@ proper S3b design pass.
Two are S2b additions:

- `--webmail` (env `ABUSEKIT_WEBMAIL_CONFIG`, default `config/webmail.yaml`) — the public list of
consumer webmail provider domains `webmail_recipient_share`/`webmail_sends_1h` match against.
consumer webmail provider domains `email.webmail_recipient_share`/`email.webmail_sends_1h` match against.
- `--brands-extra` (env `ABUSEKIT_BRANDS_EXTRA_CONFIG`, **no default**) — an optional path to a
private, `config/brands.yaml`-shaped brand list, merged (`feature.MergeBrandSets`) alongside the
shipped public `--brands` list. Empty (the default) merges in nothing. This is how an operator
Expand Down Expand Up @@ -125,7 +125,7 @@ positive subject was first flagged at).
Two corpus shapes (design §4.6), both JSONL:

- **Label-snapshot**: `{id, input:{features,text,context}, label, split, source, meta}` — one row
per already-resolved decision point. Schema: `eval/schema/corpus-v1.schema.json`.
per already-resolved decision point. Schema: `eval/schema/corpus-v2.schema.json`.
- **Event-replay pair**: an events file (`{subject,type,at,data,links?}`, internal/event's own wire
shape) plus a labels file (`{subject,label,category?,source,decision_at:{<slice>:<RFC3339>}}`).
Each subject's feature vector is rebuilt from events STRICTLY BEFORE its decision_at, using
Expand Down Expand Up @@ -159,10 +159,10 @@ go build -o abusekit ./cmd/abusekit
--slice full --floors eval/floors.yaml

# A raw scoring pipe for an external framework — stdin/stdout JSONL, no persistence.
echo '{"id":"x","input":{"features":{"subject_age_h":0.1}}}' | \
echo '{"id":"x","input":{"features":{"core.subject_age_h":0.1}}}' | \
./abusekit score --jsonl --rule new_account_velocity --scorer local

# Export labelled corpus_examples rows (design §4.9) as a corpus-v1 file
# Export labelled corpus_examples rows (design §4.9) as a corpus-v2 file
# `abusekit eval` can read straight back in. KNOWN GAP: nothing yet marks
# a row `gated` (design's "a label enters the gate corpus only after a
# second source agrees") — every row is exported regardless, with its
Expand Down Expand Up @@ -222,10 +222,27 @@ after each event and scheduled rescore:

```sh
abusekit eval --golden --brands-extra eval/fixtures/test_brands.yaml \
--golden-check eval/golden/reference-flat.jsonl
--golden-check eval/golden/reference-ns.jsonl
```

On ARM64, select `eval/golden/reference-flat-arm64.jsonl`; on AMD64 without FMA,
select `eval/golden/reference-flat-amd64-no-fma.jsonl`. The replay pins
On ARM64, select `eval/golden/reference-ns-arm64.jsonl`; on AMD64 without FMA,
select `eval/golden/reference-ns-amd64-no-fma.jsonl`. The replay pins
float64 bits, rule hashes, scorer versions, and rescore times without a database
or vendor calls. See [the baseline contract](eval/golden/README.md).

### Feature key migration (`ns-v1`)

Feature keys now belong to `core.*`, `email.*`, or `brand.*`. For example,
`key_total` is now `core.credential_total`. Flat-key score/corpus inputs fail
with `feature_renamed` and the replacement name; no aliases are accepted.
Corpus-v2 rows and cassette headers require `feature_key_space: ns-v1`.
Export new corpora after running `abusekit migrate`; migration rewrites stored
feature keys atomically and rejects unknown flat keys or collisions. Existing
reason text is retained as version 1; newly rendered reasons use namespaced keys
and `reason_version: 2`. Keep a database backup when migrating: an older binary
cannot interpret the new feature keys. No hosted deployment is part of P1.

Scoring values and rescore timing are unchanged. Fixed registry order preserves
legacy floating-point arithmetic; the immutable P0 and new P1 replay references
prove equality on all three numeric profiles. Hashes, scorer versions, and file
SHAs deliberately change once. See [the rename contract](docs/design/2026-09-29-generic-feature-packs.md#52-the-one-time-rename-bit-exact).
18 changes: 12 additions & 6 deletions cmd/abusekit/corpus_cmd.go
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@ import (
"encoding/json"
"flag"
"fmt"
"github.com/tokencanopy/abusekit/internal/feature/registry"
"os"
"strings"

Expand All @@ -15,13 +16,14 @@ import (
)

// corpusExportRow is `abusekit corpus export`'s output shape: the
// label-snapshot corpus-v1 schema (design §4.6, eval/schema/
// corpus-v1.schema.json) — this is the SAME shape eval.LoadSnapshotCorpus
// label-snapshot corpus-v2 schema (design §4.6, eval/schema/
// corpus-v2.schema.json) — this is the SAME shape eval.LoadSnapshotCorpus
// reads, so `abusekit corpus export ... > corpus.jsonl` output is always
// a valid `abusekit eval --dataset corpus.jsonl` input.
type corpusExportRow struct {
ID string `json:"id"`
Input struct {
FeatureKeySpace string `json:"feature_key_space"`
ID string `json:"id"`
Input struct {
Features map[string]float64 `json:"features"`
Text map[string][]string `json:"text,omitempty"`
Context string `json:"context,omitempty"`
Expand Down Expand Up @@ -74,7 +76,7 @@ func parseCorpusExportFlags(args []string) (corpusFlags, error) {

// runCorpus dispatches `abusekit corpus <subcommand>`. `export` is the
// only one S4 implements — design §4.9's `abusekit corpus export --split
// all|train|test --schema corpus-v1.json > corpus.jsonl`.
// all|train|test --schema corpus-v2.json > corpus.jsonl`.
func runCorpus(args []string) error {
if len(args) == 0 || args[0] != "export" {
return exitCode2(fmt.Errorf("usage: abusekit corpus export [flags] (only \"export\" is implemented)"))
Expand All @@ -84,7 +86,7 @@ func runCorpus(args []string) error {

// runCorpusExport reads every corpus_examples row for --tenant (design
// §4.9's `corpus_examples`, written by internal/serve's POST /v1/labels
// handler) and writes it to stdout as one corpus-v1 JSONL line per row.
// handler) and writes it to stdout as one corpus-v2 JSONL line per row.
//
// Known, deliberate gap (S4 scope decision — not fixed here): design
// §4.9 says "A label enters the gate corpus only after a second source
Expand Down Expand Up @@ -133,7 +135,11 @@ func runCorpusExport(args []string) error {
}

func corpusRowToExportRow(row store.CorpusExportRow) (corpusExportRow, error) {
if row.FeatureKeySpace != registry.KeySpace {
return corpusExportRow{}, fmt.Errorf("feature_key_space: cannot export %q as %s", row.FeatureKeySpace, registry.KeySpace)
}
var out corpusExportRow
out.FeatureKeySpace = registry.KeySpace
out.ID = fmt.Sprintf("corpus_%d", row.ID)
out.Label = row.Label
out.Split = row.Split
Expand Down
8 changes: 4 additions & 4 deletions cmd/abusekit/eval_cmd_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -79,7 +79,7 @@ floors:
// (runEval, not eval.Run directly — see eval/floors_test.go's
// TestGate_WeightRegressionFailsFloors for the equivalent in-process
// check) against the real eval/floors.yaml must fail the gate (exit 1).
// resource_velocity_1h is one of 9 (of 18) weights the PR body's own
// core.resource_velocity_1h is one of 9 (of 18) weights the PR body's own
// weight-zeroing sweep found DOES break the gate; the other 9 pass when
// zeroed, each with a documented reason in the PR body (redundant with
// an already-gated feature, genuinely small/secondary by design, or
Expand All @@ -90,7 +90,7 @@ func TestRunEval_MutatedWeightsBreaksGate(t *testing.T) {
if err != nil {
t.Fatalf("read shipped weights: %v", err)
}
mutated := strings.Replace(string(shipped), "resource_velocity_1h: 0.35", "resource_velocity_1h: 0.0", 1)
mutated := strings.Replace(string(shipped), "core.resource_velocity_1h: 0.35", "core.resource_velocity_1h: 0.0", 1)
if mutated == string(shipped) {
t.Fatalf("mutation did not match any line in config/local_weights.yaml — has it been reformatted?")
}
Expand All @@ -102,7 +102,7 @@ func TestRunEval_MutatedWeightsBreaksGate(t *testing.T) {
args := syntheticCorpusArgs(t, "--weights", weightsPath, "--floors", filepath.Join(root, "eval", "floors.yaml"), "--out", filepath.Join(t.TempDir(), "run.json"))
err = runEval(args)
if err == nil {
t.Fatalf("zeroing resource_velocity_1h did not break the gate")
t.Fatalf("zeroing core.resource_velocity_1h did not break the gate")
}
var ec *exitError
if !errors.As(err, &ec) || ec.code != 1 {
Expand Down Expand Up @@ -183,7 +183,7 @@ func TestScoreOneRow_MatchesEvalScoreOne(t *testing.T) {
scorer, _ := cfg.ScorerFor(rule)

row := scoreRow{ID: "x"}
row.Input.Features = map[string]float64{"subject_age_h": 0.1}
row.Input.Features = map[string]float64{"core.subject_age_h": 0.1}
got := scoreOneRow(context.Background(), row, rule, scorer, cfg.Tiers, "v1")
want, _, err := eval.ScoreOne(context.Background(), rule, scorer, eval.Options{Tiers: cfg.Tiers, PromptVersion: "v1"}, eval.Point{Features: row.Input.Features})
if err != nil {
Expand Down
8 changes: 4 additions & 4 deletions cmd/abusekit/golden_cmd_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -22,9 +22,9 @@ func goldenArgs(t *testing.T, extra ...string) []string {

func TestGoldenLastBitWeightMutationFails(t *testing.T) {
root := repoRoot(t)
referenceName := "reference-flat.jsonl"
referenceName := "reference-ns.jsonl"
if profile := eval.GoldenProfile(); profile != "amd64-fma" {
referenceName = "reference-flat-" + profile + ".jsonl"
referenceName = "reference-ns-" + profile + ".jsonl"
}
reference := filepath.Join(root, "eval", "golden", referenceName)
if err := runEval(goldenArgs(t, "--golden-check", reference)); err != nil {
Expand All @@ -35,12 +35,12 @@ func TestGoldenLastBitWeightMutationFails(t *testing.T) {
if err != nil {
t.Fatal(err)
}
old := w.Weight["resource_velocity_1h"]
old := w.Weight["core.resource_velocity_1h"]
raw, err := os.ReadFile(path)
if err != nil {
t.Fatal(err)
}
changed := strings.Replace(string(raw), "resource_velocity_1h: "+strconv.FormatFloat(old, 'g', -1, 64), "resource_velocity_1h: "+strconv.FormatFloat(math.Float64frombits(math.Float64bits(old)^1), 'g', -1, 64), 1)
changed := strings.Replace(string(raw), "core.resource_velocity_1h: "+strconv.FormatFloat(old, 'g', -1, 64), "core.resource_velocity_1h: "+strconv.FormatFloat(math.Float64frombits(math.Float64bits(old)^1), 'g', -1, 64), 1)
if changed == string(raw) {
t.Fatal("test did not mutate weight")
}
Expand Down
11 changes: 11 additions & 0 deletions cmd/abusekit/main.go
Original file line number Diff line number Diff line change
Expand Up @@ -210,7 +210,18 @@ func runServe(args []string) error {
// returns, so a caller can rely on the worker having actually stopped
// touching the store by the time it does. The `/v1/*` HTTP surface itself
// is still S3's job — nothing here listens on a port.
func requireProductionKeySpace(environment, keySpace string) error {
if environment == "production" && keySpace != "ns-v1" {
return fmt.Errorf("production requires feature key space ns-v1")
}
return nil
}

func runServeWithContext(ctx context.Context, c serveConfig) error {
if err := requireProductionKeySpace(os.Getenv("ABUSEKIT_ENV"), feature.KeySpace); err != nil {
return err
}

s, cfg, deps, err := boot(ctx, c)
if err != nil {
return err
Expand Down
2 changes: 1 addition & 1 deletion cmd/abusekit/main_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -151,7 +151,7 @@ func TestBoot_RejectsMissingRulesFile(t *testing.T) {
func TestBoot_RejectsInvalidRules(t *testing.T) {
c := shippedConfig(t)
bad := filepath.Join(t.TempDir(), "bad-rules.yaml")
if err := os.WriteFile(bad, []byte("tiers: {medium: 0.4, high: 0.8}\nrules:\n - name: r\n mode: advise\n scorer: does_not_exist\n inputs: [subject_age_h]\n labels: [benign, abusive]\n benign_label: benign\n threshold: 0.5\n"), 0o644); err != nil {
if err := os.WriteFile(bad, []byte("tiers: {medium: 0.4, high: 0.8}\nrules:\n - name: r\n mode: advise\n scorer: does_not_exist\n inputs: [core.subject_age_h]\n labels: [benign, abusive]\n benign_label: benign\n threshold: 0.5\n"), 0o644); err != nil {
t.Fatalf("write bad rules file: %v", err)
}
c.rulesPath = bad
Expand Down
54 changes: 54 additions & 0 deletions cmd/abusekit/namespace_test.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,54 @@
package main

import (
"github.com/tokencanopy/abusekit/internal/feature"
"os"
"path/filepath"
"strings"
"testing"
)

func TestScoreRejectsFlatFeatures(t *testing.T) {
p := filepath.Join(t.TempDir(), "input.jsonl")
if err := os.WriteFile(p, []byte(`{"id":"example","input":{"features":{"key_total":3}}}`), 0600); err != nil {
t.Fatal(err)
}
f, err := os.Open(p)
if err != nil {
t.Fatal(err)
}
defer f.Close()
_, err = readScoreRows(f)
if err == nil || !strings.Contains(err.Error(), "feature_renamed: key_total is now core.credential_total") {
t.Fatalf("wrong error: %v", err)
}
}

func TestNoProductionBeforeRename(t *testing.T) {
if err := requireProductionKeySpace("production", "flat-v0"); err == nil {
t.Fatal("pre-rename production accepted")
}
if err := requireProductionKeySpace("production", feature.KeySpace); err != nil {
t.Fatal(err)
}
for _, adapter := range []string{"gemini", "jev", "laya"} {
if _, err := os.Stat(filepath.Join(repoRoot(t), "internal", "model", adapter)); err == nil && feature.KeySpace != "ns-v1" {
t.Fatalf("adapter %s landed before rename", adapter)
}
}
}

func TestScoreRejectsTrailingJSON(t *testing.T) {
p := filepath.Join(t.TempDir(), "input.jsonl")
if err := os.WriteFile(p, []byte(`{"id":"example","input":{"features":{"core.subject_age_h":3}}} {"input":{"features":{"subject_age_h":4}}}`), 0600); err != nil {
t.Fatal(err)
}
f, err := os.Open(p)
if err != nil {
t.Fatal(err)
}
defer f.Close()
if _, err = readScoreRows(f); err == nil {
t.Fatal("accepted second JSON value on a score line")
}
}
17 changes: 15 additions & 2 deletions cmd/abusekit/score_cmd.go
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,8 @@ import (
"encoding/json"
"flag"
"fmt"
"github.com/tokencanopy/abusekit/internal/feature/registry"
"io"
"os"
"strings"

Expand All @@ -21,8 +23,9 @@ import (
// (design §4.6) minus the label/split/source/meta bookkeeping fields —
// score is a raw scoring pipe, not an evaluation.
type scoreRow struct {
ID string `json:"id"`
Input struct {
FeatureKeySpace string `json:"feature_key_space,omitempty"`
ID string `json:"id"`
Input struct {
Features map[string]float64 `json:"features"`
Text map[string][]string `json:"text"`
Context string `json:"context"`
Expand Down Expand Up @@ -157,9 +160,19 @@ func readScoreRows(r *os.File) ([]scoreRow, error) {
if err := dec.Decode(&row); err != nil {
return nil, fmt.Errorf("stdin:%d: invalid JSON or unknown field: %w", line, err)
}
if err := dec.Decode(new(any)); err != io.EOF {
return nil, fmt.Errorf("stdin:%d: expected exactly one JSON value per line", line)
}

if row.ID == "" {
return nil, fmt.Errorf("stdin:%d: id is required", line)
}
if row.FeatureKeySpace != "" && row.FeatureKeySpace != registry.KeySpace {
return nil, fmt.Errorf("stdin:%d: feature_key_space must be %s", line, registry.KeySpace)
}
if err := registry.ValidateKeys(row.Input.Features); err != nil {
return nil, fmt.Errorf("stdin:%d: %w", line, err)
}
rows = append(rows, row)
}
if err := scanner.Err(); err != nil {
Expand Down
Loading
Loading