Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 18 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -261,6 +261,24 @@ proposed; any of them that is taken lands as an editor change recorded here.

### Clarified

- **§4 — a `top_ngrams` `sequence` is `ngram_size` consecutive records of one
observation stream.** §4 sized the block by `ngram_size` (*"size of n-grams"*) and
described a *"sequence/transition fingerprint"*, but never said what a sequence IS, so
a producer could read it as any template-to-template relation it observes. A new
sentence under the example says it: the template ids of `ngram_size` consecutive
records of one observation stream, in observation order, so every entry holds exactly
`ngram_size` ids. What a producer treats as one observation stream (the whole window,
or a narrower scope such as one trace) stays its own; the sentence names no scope.
The probability paragraph loses the clause about a `top_ngrams` holding sequences of
more than one length, which no conforming document can now hold, and reads n as
`ngram_size`. The clause was added in this unreleased 0.10.0 line, after the
reference implementation was found putting a second relation (a declared link between
two records, not their adjacency) into the block at `ngram_size` 3; that producer now
carries such relations in `extensions` (§7), which is where a relation another
producer would not compute from the same records belongs. **Against every released
version this changes no document's validity**; against the unreleased 0.10.0 draft it
withdraws a clause that admitted two lengths. No schema changed.

- **§8 — clause 2 is mechanically decidable for `window` (§2.2), and the closing
paragraph now says so.** It read *"clauses 2 and 3 are not mechanically decidable
today: clause 2 only as far as the schema expresses it"*. That understated what is
Expand Down
20 changes: 11 additions & 9 deletions SPEC.md
Original file line number Diff line number Diff line change
Expand Up @@ -807,7 +807,7 @@ Captures *how* templates follow each other, beyond raw frequency.
"ngram_size": 2, // integer, required, size of n-grams (2 = bigrams)
"top_ngrams": [ // array, required, ordered by count desc
{
"sequence": ["h:8a3f...", "h:b104..."], // array of template_ids
"sequence": ["h:8a3f...", "h:b104..."], // array of ngram_size template_ids — see below
"count": 8421,
"probability": 0.677 // p(last | first n−1), among sequences of this length — see below
}
Expand All @@ -829,15 +829,17 @@ Captures *how* templates follow each other, beyond raw frequency.
}
```

**`sequence` — consecutive records of one observation stream.** An entry's `sequence`
is the template ids of `ngram_size` consecutive records of one observation stream, in
observation order, so every entry of `top_ngrams` holds exactly `ngram_size` ids.

**`probability` — a conditional, never a joint.** An entry's `probability` is
p(last | prefix): its `count` divided by the summed `count` of every sequence **of the
same length** the producer counted in the window whose first n − 1 template ids equal
the entry's own, n being the length of the entry's `sequence`. At `ngram_size` 2 it is
p(next | prev); at 3 it is p(third | first two). It is computed over every counted
sequence, **before** the `top_ngrams_size` cut, so the retained entries sharing a
prefix need not sum to 1. A sequence of another length never enters the sum: in a
`top_ngrams` holding sequences of more than one length, each length is conditioned
among its own. It is a conditional and not a joint probability because frequency is
p(last | prefix): its `count` divided by the summed `count` of every sequence the
producer counted in the window whose first n − 1 template ids equal the entry's own,
n being `ngram_size`. At `ngram_size` 2 it is p(next | prev); at 3 it is
p(third | first two). It is computed over every counted sequence, **before** the
`top_ngrams_size` cut, so the retained entries sharing a prefix need not sum to 1.
It is a conditional and not a joint probability because frequency is
already carried by `count`, and because §13's `ngram_delta.rate_changed` compares this
value across two windows as a transition RATE: a joint value would move with any change
in traffic mix and would mean a different quantity at each `ngram_size`.
Expand Down
Loading