Skip to content

Translate the Japanese navigation command prefixes - #752

Merged
moritz-gross merged 1 commit into
daisy:jafrom
yasumorishima:ja-navigate-prefix
Sep 5, 2026
Merged

moritz-gross merged 1 commit into
daisy:jafrom
yasumorishima:ja-navigate-prefix

Conversation

@yasumorishima

Copy link
Copy Markdown
Contributor

say-command still set the four prefixes to the English words, so a Japanese
synthesiser said them in English: "move 右", "zoom イン", "read 右",
"describe 右". They are now 移動 / ズーム / 読み上げ / 説明, the way ru, nb,
fr, hu and sv already translate theirs.

This is safe because of $CommandOffset (ec36e05): the offset is the length
of the English $NavCommand stem, not of the spoken prefix, so the suffix test
still matches. ja already carried the right offsets (5/5/5/9) — see #740 for
the two languages that do not.

Word order

Japanese puts the target before the verb, and the particle belongs to the
verb, so a $Particle is now set beside the prefix (に for move, を for read
and describe) and emitted inside each direction branch:

command before after
MoveNext move 右 右 に 移動
ReadNext read 右 右 を 読み上げ
ZoomIn zoom イン ズーム イン
ZoomInAll zoomインを最大にしました ズームインを最大にしました

Keeping the particle inside the branch matters: a command whose suffix matches
nothing (MoveCellUp, MoveStart, …) still falls back to the bare verb, the
way it does in en.

The Zoom commands keep the English order on purpose. ズーム + イン is the
ordinary loanword, and the concatenation that produces
ズームインを最大にしました needs the prefix first. The two suffix sets are
disjoint — Zoom only produces In/InAll/Out/OutAll and the others only
Next/Previous/Current/LineStart/LineEnd — so splitting the branch loses
nothing.

Two rules that would have contradicted it

Reviewing the change surfaced two other rules in the same file speaking the
same words in the other order, so they move too:

  • current (ReadCurrent / DescribeCurrent) never goes through
    say-command; it said 読み上げ 現在 and now says 現在 を読み上げ.
  • move-next-no-auto-zoom-at-edge-math used 右に for all three verbs, which
    gives read and describe the wrong particle. The direction and its particle
    are now part of each verb branch: 右に移動 / 右を読み上げ / 右を説明,
    followed by できません.

Checks

  • audit-translations ja: untranslated text 3577 → 3574; missing rules 0,
    extra rules 0. Rule differences 28 → 30, both of them in navigate.yaml and
    both intended — diffing the audit output against ja shows exactly one
    "variable difference" (the added $Particle) and one "structure difference"
    (the reordered branch).
  • Adds tests/Languages/ja/navigate.rs, the first navigation tests for ja.
    They assert the command prefix with starts_with, because what follows it
    comes from NavigationParts and is not what this touches.

Wiring the new test module in needs one line in tests/languages.rs, which is
the only change outside Rules/Languages/ja and tests/Languages/ja.

🤖 Generated with Claude Code

@yasumorishima

Copy link
Copy Markdown
Contributor Author

The one red check here is not from this PR.

  • Failing test: Languages::en::alphabets::cap_cyrillic, in Test (no-unsafe, Rules.zip) — job. The panic is 'RefCell already borrowed' at src/speech.rs:2799, i.e. re-entry while the full unicode table is being loaded, not a speech mismatch.
  • The same tree passes that same job on my fork: success.
  • This PR only touches Rules/Languages/ja/navigate.yaml, tests/Languages/ja/navigate.rs and one line of tests/languages.rs; it does not touch anything the Cyrillic alphabet test reads.
  • I have seen the same job fail non-deterministically before, on 2026-08-29, where the two runs of one commit disagreed (one red, one green) with this identical panic.

I could not find an existing issue for it (searched RefCell already borrowed and cap_cyrillic). Happy to open one with these two runs as evidence if that would be useful — I did not want to file it uninvited.

`say-command` still set the four prefixes to the English words, so a Japanese
synthesiser read them out as English: "move 右", "zoom イン", "read 右",
"describe 右". They are now 移動 / ズーム / 読み上げ / 説明.

`$CommandOffset` (from ec36e05) already makes this safe: the offset is the
length of the English `$NavCommand` stem, not of the spoken prefix, so the
suffix test still matches. ja already carried the correct offsets (5/5/5/9).
ru, nb, fr, hu and sv translate their prefixes the same way.

Word order is the other half. Japanese puts the target before the verb, and
the particle belongs to the verb, so a new `$Particle` is set beside the
prefix (に for move, を for read and describe) and emitted inside each
direction branch: 右 に 移動, 右 を 読み上げ. Keeping the particle inside the
branch means a command whose suffix matches nothing still falls back to the
bare verb, the way it does in en.

The Zoom commands keep the English order: ズーム + イン is the ordinary
loanword, and the concatenation that produces ズームインを最大にしました
requires the prefix to come first. The two suffix sets are disjoint (Zoom only
produces In/InAll/Out/OutAll, the others only Next/Previous/Current/LineStart/
LineEnd), so splitting the branch loses nothing.

Two other rules in the same file spoke the same words in the other order and
would have contradicted this, so they move too:

- `current` (ReadCurrent / DescribeCurrent) said 読み上げ 現在; it now says
  現在 を読み上げ.
- `move-next-no-auto-zoom-at-edge-math` said 右に for all three verbs, so
  read and describe got the wrong particle; the direction and particle are now
  part of each verb branch (右に移動 / 右を読み上げ / 右を説明 できません).

audit-translations ja: untranslated text 3577 -> 3574; missing rules 0, extra
rules 0. Rule differences 28 -> 30, both in navigate.yaml and both intended:
one "variable difference" for the added `$Particle`, one "structure
difference" for the reordered branch.

Adds tests/Languages/ja/navigate.rs, the first navigation tests for ja. They
assert the whole command prefix -- everything up to the first pause -- because
the description that follows comes from NavigationParts and is not what this
change touches.
Wiring it up needs one line in tests/languages.rs, which is outside
Rules/Languages/ja and tests/Languages/ja.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016JCoREgn1pJzcbdnUx4iuh
@moritz-gross
moritz-gross merged commit f8a0178 into daisy:ja Sep 5, 2026
10 checks passed
@github-project-automation github-project-automation Bot moved this from Triage to Done in MathCAT Project Board Sep 5, 2026
yasumorishima added a commit to yasumorishima/MathCAT that referenced this pull request Sep 18, 2026
…d order

Three of the navigation announcements were still in the shape en gives them.

- The 2D move announcements spoke $Move2D as it stands, and that variable holds
  an English literal: 'in', 'out of', 'end of', 'start of'. Zooming into x
  squared said "ズーム イン; in 底; x", with the "in" read out in English. Those
  four are the only values this file ever sets, so the two rules that emit it
  now map them (に入る / から出る / の終わり / の始め) and say them after the part
  they apply to, which is where Japanese wants them: "ズーム イン; 底 に入る; x".
- Undo named the verb first: 元に戻す ズームイン is "undo zoom in" word for word.
  Japanese names what is being undone first, ズームイン を 元に戻す.
- The placemarker announcements had the same order -- 読み上げ プレースホルダー 3
  for "read placeholder 3". They now name the placeholder and its number first
  and then take the particle the matching move command already uses (を 読み上げ,
  を 説明, に 移動), the pattern daisy#752 set up. セット becomes 設定, the ordinary
  Japanese for setting a value.

The four rules whose children are reordered carry `# audit-ignore`. The
divergence from en is the point of the change, and without the marker the rule
difference count goes 30 -> 34 and stops being usable as evidence that a PR did
not break the rules. The marker sits at the end of the `- name:` line: that is
the position that works today, and the one 170 of the 298 markers under
Rules/Languages already use. (daisy#742 is deciding which position becomes canonical;
if it moves, these move with the rest.)

This also promotes 底 / 上付き文字 / 下付き文字 to T:, checked against en's base /
superscript / subscript. In en the `phrase(...)` hints on those lines and the
values they annotate do not agree with each other; ja follows the values, which
are what is actually spoken.

audit-translations ja: untranslated text 3447 -> 3436, which is the 8 lines this
changes plus those 3 promotions; rule differences 30 -> 30; missing, extra and
definition counts unchanged at 0.

Tests: four in navigate.rs, with a helper that compares the whole announcement
rather than the command prefix, since none of these three is a prefix.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XzNYYoiVRzKGnDEWTJnuHh
yasumorishima added a commit to yasumorishima/MathCAT that referenced this pull request Sep 19, 2026
Each of these was translated token by token, in English order, so the Japanese
carries the words and not the meaning. Where a phrase is split across several
`t:` entries the fix moves material between them, which is what the merged
navigation work (daisy#752, daisy#773) already did; the en `phrase(...)` comments are left
alone.

  {x | x > 2}
    集合 すべて x そのようなこと x は 大なり 2
    -> 集合 すべての x ただし x は 大なり 2
    そのようなこと is "such a thing": it translates the words of "such that" and
    none of its work, which is to attach a condition to what was just named.
    ただし is what Japanese mathematical prose uses. 集合 stays in front of the
    variable rather than moving to the head-noun position Japanese would
    normally give it, because in a speech stream it tells the listener what
    they are about to hear, and the first argument can be long ({x ∈ ℤ : x > 5}).

  P(A | B)
    A 与えられた B -> A 条件は B
    与えられた is the past participle "given", which in Japanese modifies the
    noun after it, so B was being described rather than named as the condition.

  a labelled row
    行 1 ラベルを使って 第1式 -> 行 1 ラベルは 第1式
    ラベルを使って is "using a label", which says the row does something with
    one. The label is simply what the row is called.

  f|_a^b
    f 評価される b 同じ式が評価されるマイナス a
    -> f 次の点で評価 b マイナス 同じ式を次の点で評価 a
    評価される is the passive "is evaluated" with nothing to say where, and the
    minus came last, so the subtraction arrived after the thing subtracted.

  round()        ラウンド値 ("round" in katakana + 値) -> 丸めた値
  fenced-group   フェンスグループ, an English term in katakana -> 括弧のまとまり

Each phrase is swept across every file that carries it, including the sibling
branches: unicode.yaml has two more `given`s, general.yaml a second
`evaluated at`, and definitions.yaml the intent mappings for such-that, given
and conditional-probability. The copy of `such that` in unicode-full.yaml is
left alone, since that file is waiting on a MathPlayer seed.

audit-translations ja: untranslated 3415 -> 3400, which is exactly the number of
rule lines promoted; rule differences unchanged at 30; missing, extra and
definition counts unchanged at 0. definitions.yaml entries carry no case marker,
so they do not move that count.

Tests: five new in ja.rs and one updated, since set_builder_member_symbol had
pinned the old reading. Reverting the readings and nothing else turns exactly
those six red.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XzNYYoiVRzKGnDEWTJnuHh
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

2 participants