Repository navigation
Create a Japanese (ja) translation #715
Description
Activity
- I have some Japanese from MathPlayer that can be used to help seed the unicode files. So that would I think provide a better start.
- 'en' is the typical place to start as it is the most up-to-date. "zz" is a dummy test language. Maybe you meant 'zh' (zh-tw)?
- Braille is different and as you said, much harder. That should be a separate issue. I typically work on it, asking others to provide a pointer to the spec, ~200 examples (MathML and corresponding braille), along with having someone available to answer questions. However, if you want to do it yourself, that would likely be faster. I don't think I can start on it until January at the earliest. Someone implemented Russian braille themselves, so it is doable. You need to be able to write YAML rules and do a little Rust coding (in src/braille.rs).
When were you thinking on starting on the speech? I'm on holiday right and don't have access to the MathPlayer translations. We could do a seeding and you could work mainly on the Rules files and I could get you better unicode files later.
Thanks — and you're right, I misspoke about
zz; I'll start fromen.I can start on the speech right away. The seeding-first plan works well for me: if you seed a
jabranch whenever it's convenient, I'll work through theRules/Languages/jafiles fromenand drop in the better unicode files when you're back from holiday and have the MathPlayer translations to hand. No rush on your side — I'd rather you enjoy the holiday than dig those out now.I'll send the work in reviewable chunks rather than one large PR, and I'll flag anything where Japanese math reading conventions look genuinely ambiguous rather than guessing.
On braille: understood that it's a separate issue. I'd rather not open it until the speech work is actually moving — I'll come back to it once there's something to show here.
@yasumorishima great to have somebody working on the Japanese translation!
If you're using the
audit-translationsPython tool, feel free to open an issue for any suggestions/bugs/etc and I'll look into it.Regarding the seeding, that was just me using
gpt-5.6in OpenAI Codex and told it to start off with some basic language rules. From our experience, using AI agents helps to get the translation off the ground faster, but human quality control is still very important.- addedenhancementNew feature or requestNew feature or requestrulesPertains to RulesPertains to RulestranslationLanguage translation of math/codeLanguage translation of math/code
on Aug 27, 2026 @yasumorishima I let gpt-5.6 create a few rules and tests #718. A good start is to review if the unit tests are correct, and increase their coverage, and then fix / adjust rules where needed. Once you are happy with what's on the "ja" branch, we can merge it into main.
Thanks @moritz-gross, and thanks for doing the seeding — that is a big head start.
I went through
tests/Languages/ja/ja.rsand the seeded rules. Summary up front: the tests pass, but most of them encode Japanese that a Japanese reader would not accept, so tests and rules have to move together. One of them is not a wording problem but a semantic one.The one that has to be fixed first: fractions are inverted
simple_fractionexpects21/22→ 「21 分の 22」. In Japanese 「A 分の B」 means B/A — the denominator is spoken first. So this reads 21/22 as 22/21. The rule has the same order (SimpleSpeak_Rules.yaml,fraction/simple:*[1]→ 「分の」 →*[2]), which is exactly why the test is green.The reference I am using
Rather than going by feel, I am anchoring on the standard proposal for reading mathematics aloud in Japanese:
山口雄仁・川根深・澤崎陽彦「日本語による数式読み上げ法の基本構成について」日本数学教育学会誌 78(9), 239–247 (1996)
K. Yamaguchi et al., On the basic structure of a method for reading mathematical formulas aloud in Japanese, J. Japan Society of Mathematical Education — https://www.jstage.jst.go.jp/article/jjsme/78/9/78_7/_pdfKatsuhito Yamaguchi is behind ChattyInfty / InftyReader, the math TTS actually used by blind students in Japan, so this is the closest thing to a de facto standard. (It is a scanned PDF with no text layer, so I read the page images.)
Two things in it are worth knowing before reviewing any wording:
- It deliberately follows English word order. The stated goal is to avoid 逆読み — making the listener jump backwards — and it explicitly imports English prepositions as katakana: オブ (of), オーバー (over), オア (or). So the seeding's instinct to use katakana is right. What goes wrong is katakana content words (スクエア for "squared", サブ for "sub"), which is not how anyone reads these.
- It defines two levels: 厳密読み上げ法 (strict — every group gets an explicit end marker) and 簡略読み上げ法 (simplified — end markers dropped). That maps naturally onto machinery MathCAT already has, which is my one question below.
Concrete rules from it that the seeding gets wrong:
- Fractions — if numerator and denominator are both plain numbers, use ordinary Japanese 「denominator 分の numerator」. Otherwise read the numerator first, as 「分数 A オーバー B 分数終了」. This lines up strikingly well with MathCAT's existing
simplevsdefaultfraction rules. - Roots — a single token inside → 「平方根 a」 or 「ルート a」; otherwise 「平方根オブ a プラス b 根号終了」.
- Sub/superscripts — x₁ → 「x 下付き 1」, x² → 「x の 2 乗」. Multi-token ones get an end marker: 「x の上付き (1) 上付き終了」.
- Brackets —
()→ 「(丸) カッコ」 / 「(丸) カッコ閉じ」,[→ 角カッコ,{→ 中カッコ. Exceptions:f(x)→ 「f オブ x カッコ閉じ」,|x|→ 「絶対値 x 絶対値閉じ」. - Big operators — ∑ → 「サム i イコール 1 から n まで (オブ) …」, ∫ → 「インテグラル a から b まで (オブ) …」, ∏ → 「プロダクト …」, lim → 「リミット x 右向き矢印 a オブ …」.
- Symbol table (appendix) — ∈ 要素オブ, ⊂ 部分集合オブ / 含まれる (イン), ∪ ユニオン, ∩ 交わり, ∅ 空集合, ∴ ゆえに, ∵ なぜならば, ≡ 同値 / 合同, ± プラス・オア・マイナス, ≠ ノット・イコール, ≤ 小なり・オア・イコール, ≒ 近似的イコール, ∝ 比例, × 掛ける / クロス, ÷ 割る, ∇ ナブラ / デル, ∫ インテグラル, ∮ 周回インテグラル, ∑ サム / 大文字シグマ, ∞ 無限大, → 右向き矢印.
Good news: the
gradienttest's デル and 大文字 f are both consistent with this.The 16 tests
test seeded expectation should be simple_fraction21 分の 22 22 分の 21 (inverted — semantic) squared3 スクエア 3 の 2 乗 subscriptx サブ 1 x 下付き 1 set_membershipx は 属する 実数 x 要素オブ 実数 square_root平方根 の 9 平方根 9 / ルート 9 (no の) cube_root立方根 の 8 立方根 8 sine_functionサイン の x サイン オブ x / サイン x absolute_value絶対値 の x 絶対値 x 絶対値閉じ (簡略: 絶対値 x) parenthesized_expression開き丸括弧 … 閉じ丸括弧 丸カッコ … 丸カッコ閉じ less_thanx は 小なり 5 x 小なり 5 (は + 小なり mixes two registers) summation総和 から i イコール 1, に n の i サム i イコール 1 から n まで オブ i definite_integral積分 から 0, に 1 の; x 微分 d x インテグラル 0 から 1 まで オブ x d x Three are fine as they stand:
arithmetic_operators,multiplication_and_division,greek_letters.gradientis close enough to keep (the の in the Verbose form is the only thing I would revisit).Beyond the tests
definitions.yamlhas a number of machine-translation accidents. A few, so you can gauge the scale:secant→ 種目 (means "event / category" — unrelated),imaginary-part→ 想像上の部分 ("imaginary" in the fictional sense; should be 虚部),real-part→ 実際の部分 (should be 実部),complex-conjugate→ 複雑なコンジュゲート ("complicated conjugate"),identity-matrix→ アイデンティティ マトリックス (should be 単位行列),matrix→ マトリクス (should be 行列),set→ セット (should be 集合),mode→ モード (should be 最頻値). InSimpleSpeak_Rules.yaml, "per" is rendered パーカー, which is a hooded sweatshirt.Plenty of other entries are correct (床関数, 天井関数, 線形包, 行列式, 標準偏差, 濃度), so this needs to be walked entry by entry rather than swept.
One question before I start
The strict/simplified split above is real in Japanese practice, and MathCAT already has three axes that could carry it: SimpleSpeak vs ClearSpeak,
Verbosity, andImpairment = Blindness. My inclination is:- keep 簡略 (no explicit end markers) as the default, and
- emit the 厳密 end markers (分数終了, 根号終了, 絶対値閉じ, 上付き終了 …) under Verbose / Blindness — which is close to what the seeded
fraction/defaultrule already does with 分数 … 分数終わり.
Would you rather I do that, or keep
jastructurally parallel toenand not introduce a language-specific policy? I would rather ask than bake it in.Plan
Small PRs against
ja, in this order, each moving the rule and its test together and adding coverage for the cases the reference names explicitly:- fractions (the inversion)
- powers and sub/superscripts
- brackets, absolute value, roots
- sums / integrals / limits
definitions.yamlterminology
I will also run
audit-translations jaand open an issue if I run into anything, as you offered.I'm back from holiday and created the unicode files from speech that was in MathPlayer. You also created this files. I don't want to overwrite what you wrote because I know there are some problems with the new versions including several strings that didn't translate at all. So instead, here are diffs for those files. I am not able to determine which is better when both show Japanese characters:
ja_unicode_diff_no_comments.txt
ja_unicode_full_diff_no_comments.txtWelcome back, and thanks for generating these. For unicode-full.yaml the MathPlayer side is clearly better wherever it is in Japanese — the current file is mostly literal machine translation (¥ comes out as 入会金, "entrance fee", and "dagger with right guard" as 正しいガード, "correct guard"). About 900 of its entries came through still in English, though, so I'd take the MathPlayer file as the base and fill those in, fixing the bad ones from the current file as I go. For unicode.yaml I'd keep the current file, since it was checked against the Yamaguchi reading rules and the MathPlayer version has a few entries that look shifted onto the wrong character (
"→ backslash,\→ 大カッコ, fraktur → 実部オブ), but I'd pull over the terms that are better there, such as カッコ / カッコ閉じ. If that works for you, I can do both as a PR against the ja branch.
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsTriage
What is missing
Rules/Languages/currently containsde el en es fi fr hu id nb pl ru sv vi zhplus thezztemplate — there is noja.Rules/Braille/contains ASCIIMath, CMU, French, LaTeX, Nemeth, Russian, Swedish, UEB and Vietnam — there is no Japanese math braille code either.The practical effect is that Japanese screen reader users get no math speech in their own language from MathCAT.
I searched the tracker for existing Japanese requests (open and closed) and found none, so I am opening one following the pattern of #466 (Catalan), #257 (Italian) and #548 (Portuguese).
Background
I am a native Japanese speaker and I would like to take this on. I contributed #665 to this repo earlier this year, so I am somewhat familiar with the codebase, though I have not worked on the rule files before.
I have read the translator's guide (https://daisy.github.io/MathCAT/helpers.html) and
AGENTS.md, including thet:/T:convention anduv run --project PythonScripts audit-translations. I understand the guide's estimate that a thorough TTS translation is on the order of 300–450 hours and that it needs review by someone with a mathematics background followed by user testing, so I am not treating this as a quick patch.Questions before I start
Seeding. For Catalan you created a
cabranch with an initial seeding and had the translator correct the strings from there (Create a Catalan translation #466). Would you like to do the same forja, or would you prefer that I open a PR addingRules/Languages/jamyself?Which language to start from. The guide recommends starting from a grammatically similar language when one exists. I do not think any of the current languages is a good base for Japanese: word order is subject-object-verb, and quantities are expressed with counter words that vary with the kind of object being counted, so
zhis less close than it may appear. My plan is to start fromenwithzzas the structural reference, unless you advise otherwise.Braille. Other Braille codes to implement #232 invites an issue from anyone willing to help with an additional braille code. Japanese mathematics braille is a separate and substantial piece of work. Would you prefer that tracked as its own issue when the speech work is far enough along, or kept here?
Happy to adjust the plan to whatever fits your process best.