Skip to content

fix: boundary treats combining marks as word characters, like Ruby - #11

Merged
ronaldtse merged 1 commit into
mainfrom
fix/parallel-selection-max-length
Oct 1, 2026
Merged

ronaldtse merged 1 commit into
mainfrom
fix/parallel-selection-max-length

Conversation

@ronaldtse

Copy link
Copy Markdown
Contributor

Ruby's \b counts combining marks (Arabic diacritics) as word characters; Python's \w does not — Mn is not alphanumeric. Proven with a two-line test: ئ\b on دَائِم does not match in Ruby, does in Python. Every word-final rule (X + boundary → "’a") therefore fired wrongly, doubling vowels — دَائِم → "dā’aim" instead of "dā’im".

The boundary now compiles to an explicit word/non-word transition whose word class includes combining marks (0300–036F, Arabic diacritics 064B–065F, 0670, Quranic annotation 06D6–06ED), pinned by a spec on the hamza-carrier/kasra junction both ways.

Through the Ruby bridge against the ISC corpus: alalc-ara goes 10 failures → 4 (dā’im, mala’a, al-miṣrī, al-ghazzī, al-aṣl all heal; zero regressions). Builds on #10's Ruby-exact selection ordering.

Ruby's \b counts combining marks (Arabic diacritics) as word
characters; Python's \w does not — Mn is not alphanumeric. At a
hamza-carrier + kasra junction Ruby saw no boundary while Python
did, so every word-final rule (X + boundary -> "'a") fired wrongly
and doubled vowels: دَائِم became dā'aim instead of dā'im. The
boundary now compiles to an explicit word/non-word transition over a
word class that includes combining marks.

Through the Ruby bridge, alalc-ara: 10 failures -> 4.
@ronaldtse
ronaldtse merged commit 988f240 into main Oct 1, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant