How the Unicode Bidirectional Algorithm works
Learn how the Unicode bidi algorithm orders mixed Arabic and Latin text, why numbers and punctuation move, and when HTML isolation fixes it.
By Omran Khleifat, Founder of Ritla
Guide13 min read
A seed-bank dashboard stores an accession code as 618-742-93. On one screen the code looks exactly like that. After an Arabic label, its groups appear in the opposite visual order: 93-742-618. The database has not changed the value. The browser has not translated it. It has applied the Unicode Bidirectional Algorithm to a different text context.
That is the problem this algorithm solves, and occasionally the problem it reveals. A line containing Arabic, Latin letters, digits, and punctuation cannot be drawn by simply walking through the stored characters from one edge of the screen to the other. The Unicode Bidirectional Algorithm, or UBA, gives those characters a display order while preserving their logical, stored order. You do not need to implement its rules in your UI. You do need enough of a mental model to stop fixing a text-order problem with flex-direction: row-reverse (which is rather like repairing a watch with a ladder).
Here is the short version. The browser establishes a paragraph's base direction, classifies characters by their directional behavior, resolves the characters whose direction depends on context, assigns levels to the resulting runs, then reorders each displayed line. The complicated part is not the existence of RTL text. It is the punctuation and numbers between runs, where the algorithm must infer what belongs together.
Stored order is not display order#
In the DOM, the accession code remains 618-742-93. Its digits and hyphens are stored in that sequence, and JavaScript reads the same sequence from textContent. The UBA determines how it is presented. It does not rewrite the DOM, reverse the database field, or change the characters sent to an API. W3C's explanation of logical and visual ordering is worth reading if a teammate is tempted to save the characters in screen order. Please stop them before they make a migration plan.
This distinction changes how you test. Selecting or copying the code gives the logical sequence, even when the screen shows the groups in a different order. The accessibility tree also represents the logical content. A copy-and-paste assertion can prove the stored value is intact, but it cannot prove a reader sees the right sequence. Check the rendered line with your eyes or a visual test as well.
Some visual reordering is correct. An Arabic reader reads the surrounding sentence from right to left. A numeric range written with a hyphen may therefore be read in its intended sequence even if its groups appear reversed to someone scanning from left to right. An accession code is different: its groups have a fixed identifier order. Before calling a display a bug, decide what the value means and how an Arabic reader should encounter it. The algorithm cannot know that 618-742-93 is a seed-bank key rather than a range. That bit is your job.
The paragraph sets the starting direction#
The algorithm works within a paragraph. Its base direction supplies a starting level: even for LTR, odd for RTL. When a higher-level format such as HTML provides the direction, the UBA uses that context instead of guessing solely from the characters. In an Arabic page, declare the page's language and direction explicitly:
<html lang="ar" dir="rtl">The HTML dir attribute sets base direction. lang="ar" identifies the language but does not set that direction. In the Chromium 151 behavior checked for this guide, an Arabic page with only lang="ar" still computes to LTR. concerns document direction that is missing, contradicted, or declared only in CSS. The HTML direction guide covers page-level markup in more detail; here we care about what that direction gives the algorithm.
The base direction matters most where text does not announce its own direction. Latin letters have a strong LTR type. Arabic letters have a strong RTL Arabic type. A space, bracket, or hyphen cannot always decide where it belongs without the runs and paragraph around it. The paragraph level supplies the fallback when context remains ambiguous. A phrase may therefore look different when moved into an RTL paragraph even though the phrase's characters are unchanged.
With dir="auto", HTML chooses direction from the first strongly directional character in the element. The W3C guidance on inline bidi markup explains the mechanism and its limits. It is not language detection. A user comment beginning with a Latin product name and continuing in Arabic can choose LTR because the first strong letter is Latin. A value containing only Western digits and punctuation has no strong letter to choose from; in the checked Chromium behavior, dir="auto" resolves that isolated value to LTR. Use an explicit direction when your data contract already tells you what the value is.
Characters have types, not little arrows#
The UBA assigns each character a bidirectional class. It does not classify a whole string as “Arabic” or “English” and finish the job. The classes relevant to this accession-code example are:
| Character in the interface | UBA class | What that means here |
|---|---|---|
| Arabic letter | AL, strong Arabic RTL | Supplies strong context for following weak characters. |
| Latin letter | L, strong LTR | Keeps its letter run LTR. |
| Western digit | EN, European Number | A weak type whose resolved behavior depends on context. |
| Arabic-Indic digit | AN, Arabic Number | A numeric type with different separator behavior from EN. |
| Hyphen-minus | ES, European Separator | Can join EN groups under a specific rule, but not every pair of numbers. |
| Space | WS, whitespace | Does not glue numeric groups into one identifier. |
These names come from the Unicode character-type table. “European Number” and “Arabic Number” are bidi classes, not a request to convert the glyphs. When rule W2 changes a Western digit's resolved type from EN to AN after Arabic text, the browser does not turn 6 into another numeral. It changes how the algorithm treats the digit while ordering the line. This is an easy detail to miss, mostly because Unicode's terminology is trying to be precise and is succeeding at being a little unfriendly.
There are more classes: R for Hebrew and other RTL letters, CS for common separators such as the colon and full stop, ON for many other neutral characters, and explicit formatting and isolate controls. You do not need to memorize the table. You need to recognize that a hyphen, a space, and a slash are not interchangeable just because they all sit between groups of digits.
The algorithm also does not do translation, choose a font, or design your grid. Arabic glyph shaping and the ordering of text runs are related during rendering but are not the same operation. If the text box is positioned on the wrong side of a card, inspect CSS layout. If the box is in the right place but its punctuation or numeric groups look wrong, inspect bidi context. Moving the box will not repair the string inside it.
Weak types explain the numeric surprise#
Let's put the code after an Arabic label, in logical source order:
<p lang="ar" dir="rtl">رقم العينة 618-742-93</p>In the checked Chromium behavior, the numeric part appears as 93-742-618 when described left to right. The groups themselves retain their digit order; their positions reverse. The reason is a short chain of rules in UAX #9's weak-type phase:
- The Arabic letters are strong
ALcharacters. The Western digits begin asEN. - Rule W2 looks backward from each
ENgroup to the nearest strong type in the relevant isolating run sequence. Here it finds ArabicAL, so those digits are resolved asANfor bidi processing. - Rule W4 can absorb an
EShyphen between twoENgroups. It does not absorb anEShyphen betweenANgroups. The remaining hyphens become neutral under W6. - Neutral resolution takes the surrounding context into account. The hyphens do not hold the three numeric groups together as one LTR run, so the groups take their positions in the RTL line.
You need not run W2 through W6 by hand for every field. The useful conclusion is narrower: a hyphenated Western-digit identifier can appear intact alone and change when Arabic text precedes it in the same paragraph. That makes a screenshot of a bare fixture a poor test for the production label.
Compare the same stored value on its own:
<p dir="rtl">618-742-93</p>With no preceding Arabic strong character in that paragraph, the checked Chromium behavior displays the numeric groups as written. The context changed; the value did not. This is why a developer can inspect a detail pane, see a correct identifier, and still ship a broken inline notification using the same data.
Spaces give another useful contrast. In a standalone RTL paragraph, this fixture:
<p dir="rtl">618 742 93</p>displays its groups as 93 742 618 in the checked behavior. Spaces do not become numeric separators under W4. They resolve with the RTL context between the groups. So a quick “fix” that replaces hyphens with spaces is not a fix. It may make the visual order less predictable while also changing an identifier that was supposed to remain exact. A fine result, if your goal was to annoy both QA and the database.
Digit shape is an edge case, not decoration#
The EN to AN change above concerns Western digits following Arabic letters. Arabic-Indic digits begin with the AN class. That difference is not a font preference. It changes which weak-type rules apply before the line is reordered. The Unicode character-type table makes the distinction explicit.
Imagine a second seed-bank system that writes grouped accession values with Arabic-Indic digits and spaces:
<p lang="ar" dir="rtl">٦١٨ ٧٤٢ ٩٣</p>In the checked Chromium behavior, those spaced groups exchange positions even when the value is alone. More surprisingly, putting that exact value inside <bdi dir="ltr"> does not preserve the group sequence. The spaces between AN groups still resolve in a way that lets the groups reorder. This is a limit to the earlier bdi dir="ltr" recommendation, not a reason to abandon isolation for Western-digit identifiers.
If the product requires those Arabic-Indic groups in a fixed left-to-right sequence, first ask whether spaces are genuinely part of the identifier format. If they are, a narrowly scoped <bdo dir="ltr"> can force the display order for that known notation:
<p lang="ar" dir="rtl"><bdo dir="ltr">٦١٨ ٧٤٢ ٩٣</bdo></p>Test it with the real value, including adjacent punctuation, because an override changes more than an isolate does. Do not put bdo around the whole Arabic sentence. It is a special tool for a special representation, not the default response to every bidi screenshot.
W2 also has a script-specific boundary. It looks backward for AL, the strong Arabic type. Hebrew letters are strong R, not AL, so Western digits after Hebrew do not take the same W2 path. “RTL language” is not a single numeric behavior. This matters when a shared component is tested with Arabic and Hebrew fixtures: one passing screenshot does not settle the other script's punctuation and number cases.
Neutrals and brackets use their surroundings#
Numbers are only half the story. Much of the apparent punctuation trouble comes from characters with neutral behavior: spaces, quotes, and many symbols. After weak numeric types have been resolved, the UBA resolves neutral types from surrounding strong text. Where the sides disagree, the embedding direction helps decide. This is why a punctuation mark at the end of an English phrase inside Arabic text may appear at the wrong visual end of that phrase if the phrase has no isolating boundary.
Paired brackets have extra treatment. Rule N0 in UAX #9 identifies bracket pairs and considers the text inside them and the established context. The opening and closing bracket are not resolved as two unrelated guesses. That does not mean every parenthesized Latin phrase will be safe in every surrounding sentence. The context still matters, particularly if the phrase contains mixed scripts or is moved between labels. Inspect the rendered result rather than deciding that parentheses are a universal bidi shield.
For example, a seed-bank record might name a storage location in Latin inside an Arabic sentence:
<p lang="ar" dir="rtl">مكان الحفظ (North Vault) في السجل</p>The Latin words form a strong LTR run. The parentheses and neighboring spaces need the surrounding bidi rules. If North Vault is a self-contained location name whose punctuation must travel with it, wrap the whole name and its punctuation in an isolated element, then inspect the output in the actual sentence. Do not insert a directional mark at a random edge until the screenshot looks agreeable; invisible characters are difficult to review and easy to carry into copied data.
Levels produce the final visual order#
After weak and neutral types have been resolved, the algorithm assigns implicit embedding levels. An RTL paragraph starts at level 1. LTR letter runs and numeric runs inside it receive even levels above that base. Those levels let the renderer keep each run's internal reading order while placing runs in the paragraph's overall direction.
The final reversal is performed per displayed line. UAX #9's reordering phase determines levels, accounts for line breaks, then applies rule L2 to reorder runs on each line. This is why “reverse the string” is not an implementation strategy. A plain reversal would flip digits within a group, scramble Latin words, and ignore where the browser wrapped the line. It would also change the stored value instead of just its presentation.
Line wrapping deserves a real test. A notification may fit on one line on desktop and wrap before the identifier on a narrow phone. The UBA still processes the paragraph, but its final display reordering is line-specific. If a punctuation problem appears only at one width, inspect the same text before and after the wrap rather than assuming a responsive CSS rule moved it. CSS can still be involved if it changes direction, unicode-bidi, or the inline structure, but a line break alone can change where neutral characters land on the displayed line.
The algorithm can also mirror the glyph used for certain characters, such as paired brackets, in an RTL context. That is separate from reversing the code points in your data and separate again from mirroring a CSS icon. The Unicode reordering and mirroring rules define the text behavior. If an SVG arrow points the wrong way, that is a UI-direction decision, not something the UBA will fix for you.
Isolation gives the identifier its own boundary#
The accession code is a fixed-order identifier. Give it an explicitly LTR isolated span within the RTL sentence:
Problem
<p lang="ar" dir="rtl">رقم العينة 618-742-93</p>Better
<p lang="ar" dir="rtl">رقم العينة <bdi dir="ltr">618-742-93</bdi></p>The text and component are the same. The difference is that the identifier is now a separate directional unit. The bdi element isolates its contents from surrounding text; its explicit dir="ltr" tells the browser that this identifier's order is known. In the checked Chromium behavior, the numeric groups display as 618-742-93 inside that unit. The surrounding Arabic sentence remains RTL. W3C's inline-markup guidance explains why isolation keeps an inserted value from affecting, or being affected by, its neighbors.
A plain <span> without dir does not create that boundary. The same value can reorder inside it exactly as it did without a wrapper. Styling a span with text-align: left is no substitute either; alignment positions the line in a box, not the bidi runs within the line. When the direction is known, an element with dir="ltr" can also be suitable. When direction is unknown, <bdi> or dir="auto" lets the browser use the first strong character while isolating the value. The choice depends on whether you know the content's direction, not on which tag looks more technical.
This is also where not to overcorrect. Do not force an entire Arabic card to LTR because one code needs LTR isolation. That would change the base direction for labels, punctuation, and layout descendants. concerns a forced-LTR container or bidi override over Arabic content. Isolate the small known-direction value, and leave the Arabic container alone.
For plain-text output where HTML markup is unavailable, Unicode supplies isolate controls (LRI, RLI, FSI, and PDI). They are real tools, but they are invisible in normal editing and must be paired and handled carefully. In HTML, semantic direction markup is easier to inspect and maintain. Avoid bdo and bidi overrides for an ordinary identifier: they force character direction rather than simply giving a self-contained run its base direction. W3C distinguishes isolation from override, and the distinction is not decorative.
Test the rendered string, not just its value#
Make a small fixture with three contexts for the same accession code: after the Arabic label, alone in an RTL paragraph, and after the label with isolation. Keep the exact source string constant. In Chromium 151, the checked results for this shape are 93-742-618, 618-742-93, and 618-742-93 respectively. Check those visually in the browser you support; do not infer the result from textContent or a copy operation.
Then test one value that begins with Latin letters. After Arabic text, a value beginning with Latin letters and followed by hyphenated digits can remain intact without isolation in the checked behavior. That is not evidence that the numeric-only identifier is safe. Bidi bugs are often sensitive to the shape of the value. A test suite that uses only one convenient code will happily certify a rule it never exercised.
Inspect the element's dir attribute and computed direction when a result surprises you. Check whether the value is inside the same paragraph as the Arabic label or has its own element, whether punctuation is inside or outside the isolated span, and whether a formatter inserted spaces or different digit shapes. The format can change across locales and browser data, so use the actual product string. For a multiline message, inspect both a wide and a narrow viewport. The line break is part of the display case.
If the visible problem is an unisolated phone number, email address, or Latin token inside Arabic text, covers that hazard. concerns punctuation rendering at the wrong visual end of Arabic text. Neither check should be claimed as an automatic verdict on every numeric accession code; verify that example directly. The broader bidi implementation guide covers more value types and markup choices once this algorithm's moving parts make sense.
A free Ritla scan can point you to the unisolated mixed-text and punctuation cases those checks describe. For the accession code here, the decisive test is simpler: leave the stored bytes alone, put the value in its intended Arabic sentence, and confirm that an Arabic reader sees the identifier in its fixed order. The algorithm handles the rest. It has enough work already.