Skip to content

Create a Japanese (ja) translation #715

Description

@yasumorishima

What is missing

Rules/Languages/ currently contains de el en es fi fr hu id nb pl ru sv vi zh plus the zz template — there is no ja. Rules/Braille/ contains ASCIIMath, CMU, French, LaTeX, Nemeth, Russian, Swedish, UEB and Vietnam — there is no Japanese math braille code either.

The practical effect is that Japanese screen reader users get no math speech in their own language from MathCAT.

I searched the tracker for existing Japanese requests (open and closed) and found none, so I am opening one following the pattern of #466 (Catalan), #257 (Italian) and #548 (Portuguese).

Background

I am a native Japanese speaker and I would like to take this on. I contributed #665 to this repo earlier this year, so I am somewhat familiar with the codebase, though I have not worked on the rule files before.

I have read the translator's guide (https://daisy.github.io/MathCAT/helpers.html) and AGENTS.md, including the t: / T: convention and uv run --project PythonScripts audit-translations. I understand the guide's estimate that a thorough TTS translation is on the order of 300–450 hours and that it needs review by someone with a mathematics background followed by user testing, so I am not treating this as a quick patch.

Questions before I start

  1. Seeding. For Catalan you created a ca branch with an initial seeding and had the translator correct the strings from there (Create a Catalan translation #466). Would you like to do the same for ja, or would you prefer that I open a PR adding Rules/Languages/ja myself?

  2. Which language to start from. The guide recommends starting from a grammatically similar language when one exists. I do not think any of the current languages is a good base for Japanese: word order is subject-object-verb, and quantities are expressed with counter words that vary with the kind of object being counted, so zh is less close than it may appear. My plan is to start from en with zz as the structural reference, unless you advise otherwise.

  3. Braille. Other Braille codes to implement #232 invites an issue from anyone willing to help with an additional braille code. Japanese mathematics braille is a separate and substantial piece of work. Would you prefer that tracked as its own issue when the speech work is far enough along, or kept here?

Happy to adjust the plan to whatever fits your process best.

Activity

  1. NSoiffer commented on Aug 26, 2026

    @NSoiffer
    Collaborator
    1. I have some Japanese from MathPlayer that can be used to help seed the unicode files. So that would I think provide a better start.
    2. 'en' is the typical place to start as it is the most up-to-date. "zz" is a dummy test language. Maybe you meant 'zh' (zh-tw)?
    3. Braille is different and as you said, much harder. That should be a separate issue. I typically work on it, asking others to provide a pointer to the spec, ~200 examples (MathML and corresponding braille), along with having someone available to answer questions. However, if you want to do it yourself, that would likely be faster. I don't think I can start on it until January at the earliest. Someone implemented Russian braille themselves, so it is doable. You need to be able to write YAML rules and do a little Rust coding (in src/braille.rs).

    When were you thinking on starting on the speech? I'm on holiday right and don't have access to the MathPlayer translations. We could do a seeding and you could work mainly on the Rules files and I could get you better unicode files later.

  2. yasumorishima commented on Aug 26, 2026

    @yasumorishima
    ContributorAuthor

    Thanks — and you're right, I misspoke about zz; I'll start from en.

    I can start on the speech right away. The seeding-first plan works well for me: if you seed a ja branch whenever it's convenient, I'll work through the Rules/Languages/ja files from en and drop in the better unicode files when you're back from holiday and have the MathPlayer translations to hand. No rush on your side — I'd rather you enjoy the holiday than dig those out now.

    I'll send the work in reviewable chunks rather than one large PR, and I'll flag anything where Japanese math reading conventions look genuinely ambiguous rather than guessing.

    On braille: understood that it's a separate issue. I'd rather not open it until the speech work is actually moving — I'll come back to it once there's something to show here.

  3. moritz-gross commented on Aug 27, 2026

    @moritz-gross
    Collaborator

    @yasumorishima great to have somebody working on the Japanese translation!

    If you're using the audit-translations Python tool, feel free to open an issue for any suggestions/bugs/etc and I'll look into it.

    Regarding the seeding, that was just me using gpt-5.6 in OpenAI Codex and told it to start off with some basic language rules. From our experience, using AI agents helps to get the translation off the ground faster, but human quality control is still very important.

  4. moritz-gross commented on Aug 28, 2026

    @moritz-gross
    Collaborator

    @yasumorishima I let gpt-5.6 create a few rules and tests #718. A good start is to review if the unit tests are correct, and increase their coverage, and then fix / adjust rules where needed. Once you are happy with what's on the "ja" branch, we can merge it into main.

  5. yasumorishima commented on Aug 28, 2026

    @yasumorishima
    ContributorAuthor

    Thanks @moritz-gross, and thanks for doing the seeding — that is a big head start.

    I went through tests/Languages/ja/ja.rs and the seeded rules. Summary up front: the tests pass, but most of them encode Japanese that a Japanese reader would not accept, so tests and rules have to move together. One of them is not a wording problem but a semantic one.

    The one that has to be fixed first: fractions are inverted

    simple_fraction expects 21/22 → 「21 分の 22」. In Japanese 「A 分の B」 means B/A — the denominator is spoken first. So this reads 21/22 as 22/21. The rule has the same order (SimpleSpeak_Rules.yaml, fraction/simple: *[1] → 「分の」 → *[2]), which is exactly why the test is green.

    The reference I am using

    Rather than going by feel, I am anchoring on the standard proposal for reading mathematics aloud in Japanese:

    山口雄仁・川根深・澤崎陽彦「日本語による数式読み上げ法の基本構成について」日本数学教育学会誌 78(9), 239–247 (1996)
    K. Yamaguchi et al., On the basic structure of a method for reading mathematical formulas aloud in Japanese, J. Japan Society of Mathematical Education — https://www.jstage.jst.go.jp/article/jjsme/78/9/78_7/_pdf

    Katsuhito Yamaguchi is behind ChattyInfty / InftyReader, the math TTS actually used by blind students in Japan, so this is the closest thing to a de facto standard. (It is a scanned PDF with no text layer, so I read the page images.)

    Two things in it are worth knowing before reviewing any wording:

    1. It deliberately follows English word order. The stated goal is to avoid 逆読み — making the listener jump backwards — and it explicitly imports English prepositions as katakana: オブ (of), オーバー (over), オア (or). So the seeding's instinct to use katakana is right. What goes wrong is katakana content words (スクエア for "squared", サブ for "sub"), which is not how anyone reads these.
    2. It defines two levels: 厳密読み上げ法 (strict — every group gets an explicit end marker) and 簡略読み上げ法 (simplified — end markers dropped). That maps naturally onto machinery MathCAT already has, which is my one question below.

    Concrete rules from it that the seeding gets wrong:

    • Fractions — if numerator and denominator are both plain numbers, use ordinary Japanese 「denominator 分の numerator」. Otherwise read the numerator first, as 「分数 A オーバー B 分数終了」. This lines up strikingly well with MathCAT's existing simple vs default fraction rules.
    • Roots — a single token inside → 「平方根 a」 or 「ルート a」; otherwise 「平方根オブ a プラス b 根号終了」.
    • Sub/superscripts — x₁ → 「x 下付き 1」, x² → 「x の 2 乗」. Multi-token ones get an end marker: 「x の上付き (1) 上付き終了」.
    • Brackets — ( ) → 「(丸) カッコ」 / 「(丸) カッコ閉じ」, [ → 角カッコ, { → 中カッコ. Exceptions: f(x) → 「f オブ x カッコ閉じ」, |x| → 「絶対値 x 絶対値閉じ」.
    • Big operators — ∑ → 「サム i イコール 1 から n まで (オブ) …」, ∫ → 「インテグラル a から b まで (オブ) …」, ∏ → 「プロダクト …」, lim → 「リミット x 右向き矢印 a オブ …」.
    • Symbol table (appendix) — ∈ 要素オブ, ⊂ 部分集合オブ / 含まれる (イン), ∪ ユニオン, ∩ 交わり, ∅ 空集合, ∴ ゆえに, ∵ なぜならば, ≡ 同値 / 合同, ± プラス・オア・マイナス, ≠ ノット・イコール, ≤ 小なり・オア・イコール, ≒ 近似的イコール, ∝ 比例, × 掛ける / クロス, ÷ 割る, ∇ ナブラ / デル, ∫ インテグラル, ∮ 周回インテグラル, ∑ サム / 大文字シグマ, ∞ 無限大, → 右向き矢印.

    Good news: the gradient test's デル and 大文字 f are both consistent with this.

    The 16 tests

    test seeded expectation should be
    simple_fraction 21 分の 22 22 分の 21 (inverted — semantic)
    squared 3 スクエア 3 の 2 乗
    subscript x サブ 1 x 下付き 1
    set_membership x は 属する 実数 x 要素オブ 実数
    square_root 平方根 の 9 平方根 9 / ルート 9 (no の)
    cube_root 立方根 の 8 立方根 8
    sine_function サイン の x サイン オブ x / サイン x
    absolute_value 絶対値 の x 絶対値 x 絶対値閉じ (簡略: 絶対値 x)
    parenthesized_expression 開き丸括弧 … 閉じ丸括弧 丸カッコ … 丸カッコ閉じ
    less_than x は 小なり 5 x 小なり 5 (は + 小なり mixes two registers)
    summation 総和 から i イコール 1, に n の i サム i イコール 1 から n まで オブ i
    definite_integral 積分 から 0, に 1 の; x 微分 d x インテグラル 0 から 1 まで オブ x d x

    Three are fine as they stand: arithmetic_operators, multiplication_and_division, greek_letters. gradient is close enough to keep (the の in the Verbose form is the only thing I would revisit).

    Beyond the tests

    definitions.yaml has a number of machine-translation accidents. A few, so you can gauge the scale: secant → 種目 (means "event / category" — unrelated), imaginary-part → 想像上の部分 ("imaginary" in the fictional sense; should be 虚部), real-part → 実際の部分 (should be 実部), complex-conjugate → 複雑なコンジュゲート ("complicated conjugate"), identity-matrix → アイデンティティ マトリックス (should be 単位行列), matrix → マトリクス (should be 行列), set → セット (should be 集合), mode → モード (should be 最頻値). In SimpleSpeak_Rules.yaml, "per" is rendered パーカー, which is a hooded sweatshirt.

    Plenty of other entries are correct (床関数, 天井関数, 線形包, 行列式, 標準偏差, 濃度), so this needs to be walked entry by entry rather than swept.

    One question before I start

    The strict/simplified split above is real in Japanese practice, and MathCAT already has three axes that could carry it: SimpleSpeak vs ClearSpeak, Verbosity, and Impairment = Blindness. My inclination is:

    • keep 簡略 (no explicit end markers) as the default, and
    • emit the 厳密 end markers (分数終了, 根号終了, 絶対値閉じ, 上付き終了 …) under Verbose / Blindness — which is close to what the seeded fraction/default rule already does with 分数 … 分数終わり.

    Would you rather I do that, or keep ja structurally parallel to en and not introduce a language-specific policy? I would rather ask than bake it in.

    Plan

    Small PRs against ja, in this order, each moving the rule and its test together and adding coverage for the cases the reference names explicitly:

    1. fractions (the inversion)
    2. powers and sub/superscripts
    3. brackets, absolute value, roots
    4. sums / integrals / limits
    5. definitions.yaml terminology

    I will also run audit-translations ja and open an issue if I run into anything, as you offered.

  6. NSoiffer commented on Sep 25, 2026

    @NSoiffer
    Collaborator

    I'm back from holiday and created the unicode files from speech that was in MathPlayer. You also created this files. I don't want to overwrite what you wrote because I know there are some problems with the new versions including several strings that didn't translate at all. So instead, here are diffs for those files. I am not able to determine which is better when both show Japanese characters:
    ja_unicode_diff_no_comments.txt
    ja_unicode_full_diff_no_comments.txt

  7. yasumorishima commented on Sep 25, 2026

    @yasumorishima
    ContributorAuthor

    Welcome back, and thanks for generating these. For unicode-full.yaml the MathPlayer side is clearly better wherever it is in Japanese — the current file is mostly literal machine translation (¥ comes out as 入会金, "entrance fee", and "dagger with right guard" as 正しいガード, "correct guard"). About 900 of its entries came through still in English, though, so I'd take the MathPlayer file as the base and fill those in, fixing the bad ones from the current file as I go. For unicode.yaml I'd keep the current file, since it was checked against the Yamaguchi reading rules and the MathPlayer version has a few entries that look shifted onto the wrong character (" → backslash, \ → 大カッコ, fraktur → 実部オブ), but I'd pull over the terms that are better there, such as カッコ / カッコ閉じ. If that works for you, I can do both as a PR against the ja branch.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestrulesPertains to RulestranslationLanguage translation of math/code

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions