Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion documents/reference/voice_for_models.md
Original file line number Diff line number Diff line change
Expand Up @@ -98,7 +98,7 @@ Vocabulary entries are reference entries with their own discipline. The schema s

**semantic_domains**: required classification for content words. Each rationale says why this word belongs in that domain rather than repeating the domain's title.

**Migration-only fields**: `concept` and `grammatical_notes` remain readable while old entries move to the target contract. Do not add them to a new entry or retain them after a complete prose revision.
**Required prose fields**: every entry carries `description`, `articulatory_notes`, and structured `examples`. The schema rejects the retired `concept` and `grammatical_notes` properties. Discovery handles belong in `search_terms`; syntax and contrasts belong in `usage_notes` or the examples themselves.

**Never in an entry**: "beautifully," "perfectly embodies," "reminds us," rhetorical questions, or philosophy that repeats what the project's protocol documents already say. An entry's philosophy lives in its specifics.

Expand Down
1 change: 1 addition & 0 deletions project/development_log.md
Original file line number Diff line number Diff line change
Expand Up @@ -130,6 +130,7 @@ Implementation status, dependencies, optional evaluation work, and the next deve
| D112 | Accepted | Certify Gibran's *On Love* under the isolated translation process and give its grain work direct optional vocabulary. | Thirty-three aligned units reconstruct all 2,403 normalized Gibran characters. The source-only pass preserves love as actor through the teaching, settles stable onyms and participant roles, gives possessive plurality an audible boundary, unfolds a subject-gap relative, separates necessary desire from love through a simultaneity frame, and adds `thapori` thresh, `shireku` sift, `napewa` husk, and `pumeli` knead. The four roots add nine memberships across Ecological Systems and Material Life, Household and Daily Life, and Work, Craft, and Repair. The final Phi freezes at SHA-256 `7494350b9f12a5a616c046057cf8dffcfc5e9ad8ea54f7e7c35e0f48a85d6125`. A fresh source-blind context derives all 33 units; independent comparison and unit-scoped retries repair the remaining English disagreements and Phi attachment ambiguities. No registered compound, function word, or grammar is added. |
| D113 | Accepted | Certify Gibran's *On Children* under the isolated translation process. | Eighteen aligned units reconstruct all 980 normalized Gibran characters. The source-only pass gives the framing exchange explicit participants, establishes `ne musatapha`, reserves reflexive `miso` for a true object, states thought possession and plurality directly, and unfolds the archery image through visible materials and roles. Anonymous derivation and independent source-blind review expose three invalid possessive structures, an incorrectly attached request topic, and ambiguous strength and gladness relations. Each affected English layer is discarded before the Phi is repaired and freshly derived. The final stream freezes at SHA-256 `028a864fb7bb333245159dcf9e2048b9021c3dab04a34f0e299506da9d059455`. No root, module membership, registered compound, function word, or grammar is added. |
| D114 | Accepted | Give certified translations a public mark and explanation generated from the process register. | The website links every certified translation page to a public account of the two isolated directions, shows compact marks in the short-work and collection catalogues, and lists the current certified works from `project/translation_process_status.json`. The explanation states prominently that certified texts can still contain errors. The build validates the register before rendering and stops when a translation page, catalogue, or public list disagrees with it. Pending translations receive no mark. |
| D115 | Accepted | Close the completed lexicon prose migration at the schema boundary. | Every lexicon entry must carry `description`, `articulatory_notes`, and structured `examples`. The schema rejects the retired `concept` and `grammatical_notes` properties instead of accepting them as fallbacks. Regression tests prove both requirements and both rejections, while the completed migration ledger remains available as evidence. No vocabulary entry or generated lexicon reference changes. |

## Completed discretionary lexical migrations

Expand Down
2 changes: 1 addition & 1 deletion project/development_protocol.md
Original file line number Diff line number Diff line change
Expand Up @@ -230,7 +230,7 @@ Every addition to Phi should make the following practices available without clai
### When Using JSON Schema
[`vocabulary/schema.json`](../vocabulary/schema.json) is the executable entry contract. Stable fields identify the `word`, `gloss`, `ipa`, `syllables`, one scalar `pos`, and `description`. The target prose contract adds required `articulatory_notes` and structured `examples`; optional `search_terms`, `usage_notes`, `sound_symbolism`, and `pillars` appear only when useful. Content entries also have a `semantic_domains` object and may list one or more `modules`.

The lexicon prose migration is complete: every entry is target-shaped, and no legacy `concept` or `grammatical_notes` fields remain. The schema retains its migration-era tolerance for the legacy shapes, but that tolerance is historical; new and revised entries use the target contract only, and the committed coverage report records the completed state.
The lexicon prose migration is complete. The schema requires `description`, `articulatory_notes`, and structured `examples`, and it rejects the retired `concept` and `grammatical_notes` properties. The committed coverage report records the completed target state and exposes any attempted return to a legacy, partial, or dual shape.

Particles alone have `slot`. A Slot 1 particle also has `slot1_rank`, whose value places it in the canonical sequence of tense, aspect, voice, evidentiality, modality, and negation. The validator reads that ordering from the entries themselves.

Expand Down
2 changes: 1 addition & 1 deletion project/handoff/vocabulary_migration.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ The target prose contract is not a cosmetic rewrite. A completed content entry h
- any useful `search_terms`, `usage_notes`, `sound_symbolism`, direct `pillars`, and established `modules` memberships;
- no legacy `concept` or `grammatical_notes` field.

The schema retains its migration-era tolerance for legacy forms. The migration is complete, and that tolerance is not permission to reintroduce old fields in any entry.
The schema rejects legacy forms. The migration is complete, and every entry must carry the full target prose shape.

Entry state is computed by `scripts/vocabulary_prose_coverage.py`:

Expand Down
10 changes: 6 additions & 4 deletions project/near_term_development_plan.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,8 +24,8 @@ Community engagement is outside this plan. Phi can strengthen its language, corp

| ID | Status | Work package | Completion gate |
|---|---|---|---|
| NTP-01 | **NEXT** | Close the migrated vocabulary schema | Current entries pass a schema that requires the target prose fields and no longer accepts the retired fields. |
| NTP-02 | **READY** | Add independent source-reconstruction checking | At least the locally stored Morris source can be compared mechanically with each chapter's ordered citation stream. |
| NTP-01 | **DONE** | Close the migrated vocabulary schema | Current entries pass a schema that requires the target prose fields and no longer accepts the retired fields. |
| NTP-02 | **NEXT** | Add independent source-reconstruction checking | At least the locally stored Morris source can be compared mechanically with each chapter's ordered citation stream. |
| NTP-03 | **READY** | Establish the *News from Nowhere* continuity record | Recurring source-to-Phi choices have one maintained home that is excluded from Phi-to-English derivation. |
| NTP-04 | **READY** | Reconcile stale project records | The roadmap, development protocol, and deferred records describe the actual completed corpus and closed grammar boundary. |
| NTP-05 | **READY** | Finish the short-work certification queue | The present Gibran, Taoist, and Buddhist short works have completed the isolated process and the generated register passes. |
Expand All @@ -37,9 +37,11 @@ Community engagement is outside this plan. Phi can strengthen its language, corp

## NTP-01: Close the vocabulary schema

The lexicon migration is complete, but [`vocabulary/schema.json`](../vocabulary/schema.json) still permits the deprecated `concept` and `grammatical_notes` fields and still allows old alternatives to `articulatory_notes` and structured `examples`. That tolerance has changed from a migration aid into a regression path.
The lexicon migration was complete while [`vocabulary/schema.json`](../vocabulary/schema.json) still permitted the deprecated `concept` and `grammatical_notes` fields and allowed old alternatives to `articulatory_notes` and structured `examples`. That tolerance had changed from a migration aid into a regression path.

This package should require `articulatory_notes` and `examples` directly, remove the two retired properties and their fallback branches, replace tests that prove legacy acceptance with tests that prove legacy rejection, and revise the development and voice references that still describe migration tolerance. It should not rewrite vocabulary prose or change any word. Completion means the complete inventory and all schema, example, sentence, generation, and site checks pass under the stricter contract.
This package requires `articulatory_notes` and `examples` directly, removes the two retired properties and their fallback branches, replaces tests that prove legacy acceptance with tests that prove legacy rejection, and revises the development and voice references that described migration tolerance. It does not rewrite vocabulary prose or change any word. Completion means the complete inventory and all schema, example, sentence, generation, and site checks pass under the stricter contract.

Completed under D115. The target fields are direct requirements, the retired properties are invalid, and the migration report remains as regression evidence rather than a compatibility path.

## NTP-02: Verify source reconstruction independently

Expand Down
9 changes: 7 additions & 2 deletions scripts/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -47,7 +47,12 @@ python3 scripts/validate_sentences.py --docs

The default command checks every structured vocabulary example and the maintained teaching corpus: canon, grammar references, manual, pamphlets, primer, book, Kia, and the Short Road. `--paths` checks selected active Markdown beside the lexicon, `--lexicon-only` narrows the run to dictionary examples, and `--docs` extends the parser to every recognized complete example in active Markdown; `archive/` is excluded. Literary texts and older evaluation material remain a migration audit until each area has been reviewed. The parser's contract and limits are recorded in [`documents/validation/sentence_validator.md`](../documents/validation/sentence_validator.md).

The focused regression suites cover the executable vocabulary contract, its transitional and target prose shapes, Slot 1 metadata, Slot 2 quantity alternatives, the productive-name open class, unconditional four-syllable name acceptance, current non-content exclusion, short retired lexical forms, the completed migration ledger, and the absence of long lexicon entries:
The focused regression suites cover:

- the executable vocabulary contract, its required prose shape, and rejection of retired prose fields;
- Slot 1 metadata and Slot 2 quantity alternatives;
- the productive-name open class, four-syllable names, and non-content exclusions;
- retired lexical forms, the completed migration ledger, and the three-syllable vocabulary limit.

```bash
python3 scripts/test_vocabulary_schema.py
Expand All @@ -67,7 +72,7 @@ python3 scripts/test_content_vocabulary_decisions.py

## vocabulary_prose_coverage.py

Writes the committed migration report at `documents/validation/vocabulary_prose_coverage.json`. Each entry is classified as legacy, partial, dual, or target according to its prose fields. The main validator compares the report with the live lexicon and fails when a vocabulary edit leaves it stale.
Writes the committed contract report at `documents/validation/vocabulary_prose_coverage.json`. Each entry is classified as legacy, partial, dual, or target according to its prose fields. The first three shapes are schema-invalid, but keeping them visible in the report makes a regression plain. The main validator compares the report with the live lexicon and fails when a vocabulary edit leaves it stale.

```bash
python3 scripts/vocabulary_prose_coverage.py
Expand Down
46 changes: 19 additions & 27 deletions scripts/test_vocabulary_schema.py
Original file line number Diff line number Diff line change
Expand Up @@ -50,19 +50,6 @@ def target_entry(self, word="sileta"):
}
entry["articulatory_notes"] = notes[word]
entry["examples"] = [examples[word]]
entry.pop("concept", None)
entry.pop("grammatical_notes", None)
return entry

def legacy_entry(self, word="sileta"):
entry = copy.deepcopy(self.by_word[word])
entry.pop("articulatory_notes", None)
entry.pop("examples", None)
entry.pop("search_terms", None)
entry.pop("usage_notes", None)
entry["concept"] = "Legacy discovery label"
entry["sound_symbolism"] = "A legacy account of the word's sounds."
entry["grammatical_notes"] = "A legacy account of grammar and examples."
return entry

def test_schema_uses_and_satisfies_draft_2020_12(self):
Expand Down Expand Up @@ -343,21 +330,23 @@ def test_target_prose_contract_is_valid(self):
entry.pop("pillars", None)
self.assertEqual(validate_examples.entry_schema_errors(entry), [])

def test_legacy_prose_contract_remains_valid_during_migration(self):
entry = self.legacy_entry()
self.assertNotIn("articulatory_notes", entry)
self.assertNotIn("examples", entry)
self.assertEqual(validate_examples.entry_schema_errors(entry), [])
def test_articulatory_notes_are_required_even_with_sound_symbolism(self):
entry = self.target_entry()
entry["sound_symbolism"] = "The open ending gives the word an audible lift."
del entry["articulatory_notes"]
self.assert_invalid(entry, "articulatory_notes")

def test_articulatory_or_legacy_sound_field_is_required(self):
entry = self.legacy_entry()
del entry["sound_symbolism"]
self.assert_invalid(entry, "not valid under any of the given schemas")
def test_structured_examples_are_required(self):
entry = self.target_entry()
del entry["examples"]
self.assert_invalid(entry, "examples")

def test_structured_examples_or_legacy_grammar_field_is_required(self):
entry = self.legacy_entry()
del entry["grammatical_notes"]
self.assert_invalid(entry, "not valid under any of the given schemas")
def test_retired_prose_fields_are_rejected(self):
for field in ("concept", "grammatical_notes"):
with self.subTest(field=field):
entry = self.target_entry()
entry[field] = "Retired prose must not return."
self.assert_invalid(entry, field)

def test_structured_example_shape_is_closed(self):
entry = self.target_entry()
Expand All @@ -380,7 +369,10 @@ def test_search_terms_are_nonempty_and_unique(self):
self.assert_invalid(entry, "should be non-empty")

def test_prose_coverage_states(self):
legacy = self.legacy_entry()
legacy = {
"concept": "Retired discovery label",
"grammatical_notes": "Retired grammar prose.",
}
partial = copy.deepcopy(legacy)
partial["articulatory_notes"] = "A physical pronunciation note."
dual = self.target_entry()
Expand Down
2 changes: 0 additions & 2 deletions scripts/validate_examples.py
Original file line number Diff line number Diff line change
Expand Up @@ -646,11 +646,9 @@ def check_lexicon(entries):
if len(s) >= 2:
own_units.add(s[:-1])
prose_fields = {
"concept": data.get("concept", ""),
"description": data.get("description", ""),
"articulatory_notes": data.get("articulatory_notes", ""),
"sound_symbolism": data.get("sound_symbolism", ""),
"grammatical_notes": data.get("grammatical_notes", ""),
"usage_notes": data.get("usage_notes", ""),
}
for k, v in data.get("pillars", {}).items():
Expand Down
10 changes: 5 additions & 5 deletions scripts/vocabulary_prose_coverage.py
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
#!/usr/bin/env python3
"""Measure progress from the legacy lexicon prose shape to the target contract."""
"""Record lexicon prose-contract state and expose any migration regression."""

import json
from collections import Counter
Expand All @@ -11,7 +11,7 @@
)

TARGET_FIELDS = ("articulatory_notes", "examples")
LEGACY_FIELDS = ("concept", "grammatical_notes")
RETIRED_FIELDS = ("concept", "grammatical_notes")
COUNTED_FIELDS = (
"articulatory_notes",
"examples",
Expand All @@ -26,14 +26,14 @@


def entry_state(entry):
"""Classify one entry by which old and target prose fields it carries."""
"""Classify one entry by which retired and target prose fields it carries."""
target_count = sum(field in entry for field in TARGET_FIELDS)
has_legacy = any(field in entry for field in LEGACY_FIELDS)
has_retired = any(field in entry for field in RETIRED_FIELDS)
if target_count == 0:
return "legacy"
if target_count < len(TARGET_FIELDS):
return "partial"
return "dual" if has_legacy else "target"
return "dual" if has_retired else "target"


def load_entries(root=PROJECT_ROOT):
Expand Down
4 changes: 2 additions & 2 deletions vocabulary/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,9 +20,9 @@ An entry's filename comes from its English gloss, and the JSON keeps the field o

The target prose contract requires a definition in `description`, a physical account of pronunciation in `articulatory_notes`, and worked `examples` stored as Phi and English pairs. `search_terms` and `usage_notes` are optional aids. `sound_symbolism` is also optional: it records an honest Phi-specific phonesthetic association when one exists, not a universal claim about sounds or a hidden analysis of the word. `pillars` records only direct, specific relationships, so an entry may have none.

The lexicon prose migration is complete: every entry is target-shaped, and no legacy `concept` or `grammatical_notes` fields remain. The schema retains its migration-era tolerance for the legacy shapes, but that tolerance is historical; new entries and complete prose revisions use the target contract only.
Every entry requires `description`, `articulatory_notes`, and structured `examples`. The schema rejects the retired `concept` and `grammatical_notes` properties, so an older entry shape cannot return unnoticed.

The validator first runs the complete entry through Draft 2020-12 JSON Schema. It then checks the rules the schema cannot settle: the spoken form, cross-entry collisions, canonical file layout, Phi examples, and the word's place in the lexicon. Slot 1 ordering comes from each particle's `slot1_rank` metadata rather than a second list hidden in the validator. [`documents/validation/vocabulary_prose_coverage.json`](../documents/validation/vocabulary_prose_coverage.json) records how many entries are legacy, partial, dual, or fully target-shaped; validation fails when that report is stale.
The validator first runs the complete entry through Draft 2020-12 JSON Schema. It then checks the rules the schema cannot settle: the spoken form, cross-entry collisions, canonical file layout, Phi examples, and the word's place in the lexicon. Slot 1 ordering comes from each particle's `slot1_rank` metadata rather than a second list hidden in the validator. [`documents/validation/vocabulary_prose_coverage.json`](../documents/validation/vocabulary_prose_coverage.json) records whether an entry has fallen into a legacy, partial, or dual shape, even though those shapes are schema-invalid; validation also fails when that report is stale.

## Working on an entry

Expand Down
Loading
Loading