Skip to content

Commit d7243f3

Browse files
committed
fix: boundary treats combining marks as word characters, like Ruby
Ruby's \b counts combining marks (Arabic diacritics) as word characters; Python's \w does not — Mn is not alphanumeric. At a hamza-carrier + kasra junction Ruby saw no boundary while Python did, so every word-final rule (X + boundary -> "'a") fired wrongly and doubled vowels: دَائِم became dā'aim instead of dā'im. The boundary now compiles to an explicit word/non-word transition over a word class that includes combining marks. Through the Ruby bridge, alalc-ara: 10 failures -> 4.
1 parent b0f3fea commit d7243f3

2 files changed

Lines changed: 28 additions & 1 deletion

File tree

‎src/interscript/expr.py‎

Lines changed: 11 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -29,6 +29,16 @@
2929

3030
SPACE = re.escape(" ")
3131

32+
# Ruby's \b counts combining marks as word characters; Python's \w
33+
# does not (Mn is not alphanumeric). At a hamza-carrier + kasra
34+
# junction Ruby sees no boundary while Python does — word-final rules
35+
# fired wrongly and doubled vowels. Express the boundary as an
36+
# explicit word/non-word transition over a word class that includes
37+
# combining marks (Mnemonic ranges: combining diacritics 0300-036F,
38+
# Arabic diacritics 064B-065F, 0670, and Quranic annotation 06D6-06ED).
39+
_WORD = r"[\w\u0300-\u036F\u064B-\u065F\u0670\u06D6-\u06ED]"
40+
_BOUNDARY = "(?:(?<=" + _WORD + ")(?!" + _WORD + ")|(?<!" + _WORD + ")(?=" + _WORD + "))"
41+
3242

3343
def _unesc(s: str) -> str:
3444
return _UNESC.sub(lambda m: chr(int(m.group(1), 16)), s)
@@ -95,7 +105,7 @@ def expr_to_regex(expr: str) -> str:
95105
elif kind == "space":
96106
parts.append(SPACE)
97107
elif kind == "boundary":
98-
parts.append(r"\b")
108+
parts.append(_BOUNDARY)
99109
elif kind == "nwb":
100110
parts.append(r"\B")
101111
elif kind == "grp":

‎tests/test_engine.py‎

Lines changed: 17 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -135,3 +135,20 @@ def test_parallel_selection_matches_ruby_max_length():
135135
' sub "abc", "Y"\n }\n}\n'
136136
)
137137
assert Engine(tree2).transliterate("abc") == "Y"
138+
139+
140+
def test_boundary_treats_combining_marks_as_word_chars():
141+
"""Ruby's \\b counts combining marks (Arabic diacritics) as word
142+
characters — a word-final rule must not fire when a kasra follows
143+
the hamza carrier (dā'im, not dā'aim)."""
144+
tree = parse_imp(
145+
'stage {\n parallel {\n'
146+
' sub "ئ" + boundary, "\'a"\n'
147+
' sub "ئ", "\'"\n'
148+
' sub "ِ", "i"\n'
149+
' sub "d", "d"\n }\n}\n'
150+
)
151+
# ئ + kasra: no boundary — the bare-ئ rule fires, not the final one.
152+
assert Engine(tree).transliterate("dئِ") == "d'i"
153+
# ئ at a true word end: the boundary rule fires.
154+
assert Engine(tree).transliterate("dئ") == "d'a"

0 commit comments

Comments
 (0)