diff --git a/CHANGELOG.md b/CHANGELOG.md index 8da14e3..69e4484 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -261,6 +261,24 @@ proposed; any of them that is taken lands as an editor change recorded here. ### Clarified +- **§4 — a `top_ngrams` `sequence` is `ngram_size` consecutive records of one + observation stream.** §4 sized the block by `ngram_size` (*"size of n-grams"*) and + described a *"sequence/transition fingerprint"*, but never said what a sequence IS, so + a producer could read it as any template-to-template relation it observes. A new + sentence under the example says it: the template ids of `ngram_size` consecutive + records of one observation stream, in observation order, so every entry holds exactly + `ngram_size` ids. What a producer treats as one observation stream (the whole window, + or a narrower scope such as one trace) stays its own; the sentence names no scope. + The probability paragraph loses the clause about a `top_ngrams` holding sequences of + more than one length, which no conforming document can now hold, and reads n as + `ngram_size`. The clause was added in this unreleased 0.10.0 line, after the + reference implementation was found putting a second relation (a declared link between + two records, not their adjacency) into the block at `ngram_size` 3; that producer now + carries such relations in `extensions` (§7), which is where a relation another + producer would not compute from the same records belongs. **Against every released + version this changes no document's validity**; against the unreleased 0.10.0 draft it + withdraws a clause that admitted two lengths. No schema changed. + - **§8 — clause 2 is mechanically decidable for `window` (§2.2), and the closing paragraph now says so.** It read *"clauses 2 and 3 are not mechanically decidable today: clause 2 only as far as the schema expresses it"*. That understated what is diff --git a/SPEC.md b/SPEC.md index ac6ba35..7c96d71 100644 --- a/SPEC.md +++ b/SPEC.md @@ -807,7 +807,7 @@ Captures *how* templates follow each other, beyond raw frequency. "ngram_size": 2, // integer, required, size of n-grams (2 = bigrams) "top_ngrams": [ // array, required, ordered by count desc { - "sequence": ["h:8a3f...", "h:b104..."], // array of template_ids + "sequence": ["h:8a3f...", "h:b104..."], // array of ngram_size template_ids — see below "count": 8421, "probability": 0.677 // p(last | first n−1), among sequences of this length — see below } @@ -829,15 +829,17 @@ Captures *how* templates follow each other, beyond raw frequency. } ``` +**`sequence` — consecutive records of one observation stream.** An entry's `sequence` +is the template ids of `ngram_size` consecutive records of one observation stream, in +observation order, so every entry of `top_ngrams` holds exactly `ngram_size` ids. + **`probability` — a conditional, never a joint.** An entry's `probability` is -p(last | prefix): its `count` divided by the summed `count` of every sequence **of the -same length** the producer counted in the window whose first n − 1 template ids equal -the entry's own, n being the length of the entry's `sequence`. At `ngram_size` 2 it is -p(next | prev); at 3 it is p(third | first two). It is computed over every counted -sequence, **before** the `top_ngrams_size` cut, so the retained entries sharing a -prefix need not sum to 1. A sequence of another length never enters the sum: in a -`top_ngrams` holding sequences of more than one length, each length is conditioned -among its own. It is a conditional and not a joint probability because frequency is +p(last | prefix): its `count` divided by the summed `count` of every sequence the +producer counted in the window whose first n − 1 template ids equal the entry's own, +n being `ngram_size`. At `ngram_size` 2 it is p(next | prev); at 3 it is +p(third | first two). It is computed over every counted sequence, **before** the +`top_ngrams_size` cut, so the retained entries sharing a prefix need not sum to 1. +It is a conditional and not a joint probability because frequency is already carried by `count`, and because §13's `ngram_delta.rate_changed` compares this value across two windows as a transition RATE: a joint value would move with any change in traffic mix and would mean a different quantity at each `ngram_size`.