Arabic and English mixed text in HTML
Learn when mixed Arabic-English UI text needs direction, isolation, or language markup, with examples for names, status phrases, and contact values.
By Omran Khleifat, Founder of Ritla
Guide12 min read
An Arabic status message says a water sensor was checked. The sensor's model is AquaSense+, but the plus sign appears before the name. A developer wraps the name in a <span>, refreshes, and sees the same result. The browser is not refusing to cooperate. An unmarked span has not told it anything about where the English value ends.
Mixed Arabic and English text needs two things that are easy to confuse: a base direction for the surrounding sentence, and a boundary around any inserted run whose order must stay intact. A third attribute, lang, identifies the language of a phrase for uses such as pronunciation; it does not set text direction. This guide shows when each matters in HTML, using the water-lab message rather than trying to solve every bidi problem with one increasingly worried CSS rule.
Most mixed text works without intervention. Latin letters keep their left-to-right order inside an Arabic paragraph. The difficult cases are at the edges: a model name ending in punctuation, a copied English sentence with its own final period, a phone or email address, or a number whose groups have a fixed order. Start by deciding what the embedded value is, then mark up that value at its own boundary.
Give the Arabic sentence an RTL base#
For a page whose main content is Arabic, declare language and direction at the document level:
<html lang="ar" dir="rtl">The HTML dir attribute establishes base direction. The lang attribute identifies the language. One is not a shorthand for the other. In the Chromium 151 behavior checked for this guide, lang="ar" alone leaves the page LTR. concerns document direction that is missing, contradicted, or declared only in CSS. If the page is known to be Arabic, set dir="rtl" on its root instead of asking every paragraph to repair the missing context. The HTML direction guide covers document-level choices in depth.
The base direction tells the browser how to order a paragraph when it contains characters that cannot establish direction on their own. Arabic letters are strongly RTL; Latin letters are strongly LTR. Digits and punctuation have weaker or neutral behavior and can depend on the surrounding text. The Unicode Bidirectional Algorithm resolves those runs for display. It does not translate the sentence, change the stored string, or know that AquaSense+ is one product name.
Do not set the entire message to LTR to protect a model name. That changes the context for the Arabic words and every other value in the message. A forced-LTR container or bidi override over Arabic content is the specific concern of . Keep the sentence RTL and address the inserted value, which is the smaller problem.
Do not isolate every English word by reflex#
Consider a model name made only of Latin letters, such as AquaSense. In an Arabic sentence, its letters form an LTR run. In the checked Chromium behavior, a value beginning with Latin letters stays intact even with hyphens or spaces after it. Adding markup around every English word because it is English creates noise without necessarily changing the display.
Now add the actual model's trailing plus sign:
<p lang="ar" dir="rtl">اكتمل فحص جهاز AquaSense+ في المختبر</p>This is different. In the checked behavior, a Latin name shaped X+ between Arabic words displays with its plus sign on the other visual side of the name, described left to right as +AquaSense. The letters are fine; the neutral punctuation has been resolved using the surrounding Arabic context. The product name has a fixed spelling, so the intended display keeps the plus sign at its end. This is the moment to provide a boundary.
Problem
<p lang="ar" dir="rtl">اكتمل فحص جهاز <span>AquaSense+</span> في المختبر</p>The <span> gives you a styling hook but no directional isolation. The value can display the same way it did as bare text.
Better
<p lang="ar" dir="rtl">اكتمل فحص جهاز <bdi dir="ltr">AquaSense+</bdi> في المختبر</p><bdi> tells the bidi algorithm to treat the name as an isolated unit. Because this model name is known to be LTR, dir="ltr" states that explicitly. In the checked Chromium behavior, the plus sign stays with the name at its end. The rest of the Arabic paragraph remains RTL. An existing semantic element with dir="ltr" could provide the same boundary; <bdi> is useful here because the name needs no other element.
The fix is not “English needs bdi.” The fix is “this complete value has a known reading direction and punctuation that the surrounding sentence can pull away.” That distinction saves you from wrapping harmless text while missing the one token that actually breaks.
An English name can move a neighboring number#
There is another case where the Latin name itself looks fine. The lab might report how many times a model was checked:
<p lang="ar" dir="rtl">فحصنا AquaSense 3 مرات</p>The AquaSense letters remain in their expected order. In the checked Chromium behavior, though, a Western digit immediately after a Latin name can join that name's LTR run. The quantity then appears between the Arabic verb and the model name, away from the Arabic noun it counts. Someone testing only the model's spelling would call this correct while the sentence says something much less clear to a reader.
Isolate the model name, not the whole sentence and not the quantity:
<p lang="ar" dir="rtl">فحصنا <bdi dir="ltr">AquaSense</bdi> 3 مرات</p>The bdi boundary stops the Latin name from absorbing the following number as part of its own run. Check the resulting visual relationship between the number and مرات in the rendered line; copying the sentence only confirms its stored order. This example is why the rule cannot be “wrap English only when its punctuation looks wrong.” An inserted run may change the placement of its neighbors while appearing perfectly respectable itself. concerns unisolated Latin tokens inside Arabic text, which is the relevant hazard here.
Keep punctuation inside an English phrase#
Suppose the lab device emits a fixed English status phrase that the interface quotes for troubleshooting. It is intentionally English content, not a button that someone forgot to translate. If the period is part of the phrase, it belongs inside the phrase's directional boundary:
Problem
<p lang="ar" dir="rtl">عند ظهور <span lang="en">Calibration ready.</span>، احفظ التقرير.</p>lang="en" identifies the language, but it does not isolate the phrase or set its base direction. In the checked Chromium behavior, an English phrase shaped like this in an RTL paragraph can show its final period at the left visual end, described as .Calibration ready. The punctuation has followed the paragraph context rather than the intended English phrase.
Better
<p lang="ar" dir="rtl">عند ظهور <span lang="en" dir="ltr">Calibration ready.</span>، احفظ التقرير.</p>The phrase, including its final period, is one LTR unit. The comma and the rest of the Arabic instruction remain outside it. W3C's inline bidi guidance recommends wrapping an opposite-direction phrase tightly rather than placing direction on a much larger parent. The boundary matters as much as the value of dir: if the final period were outside the span, it could still take its direction from the Arabic sentence.
There is a product decision before the markup decision. If this is ordinary interface copy, translate it. Directional markup is not a license to leave buttons, labels, and help text in English. If it is a quoted device message or an English title users must match exactly, preserve the original and mark it up. concerns intentional foreign-language runs missing a lang attribute; it does not make an untranslated Arabic UI acceptable simply because the English words are tagged.
lang and dir solve different problems#
lang="en" can help software choose the right language treatment for an English phrase, including pronunciation in assistive technology when supported. W3C explains why language tagging matters, and WCAG's language-of-parts criterion addresses changes of language within a page, with exceptions for proper names and certain technical terms. It is not a layout command. dir="ltr" supplies the phrase's base direction and, on an HTML element, an isolated boundary. It does not declare what human language the phrase is in.
That distinction is useful even when both attributes have the same apparent value in an example. An Arabic person's name on an English page is lang="ar" and normally dir="rtl"; a Latin-script Arabic transliteration could be lang="ar" while its written direction remains LTR. The language and script direction can diverge. Do not write a rule that infers dir mechanically from lang for every piece of content.
For the model name AquaSense+, an explicit English language tag may be unnecessary: proper names and technical terms are exceptions in the language-of-parts criterion. For the complete English phrase Calibration ready., lang="en" has a clearer purpose. This is not permission to omit language markup from a paragraph of English troubleshooting instructions. It is a reminder to tag passages according to their actual language, not to chase every brand name with an attribute.
Also distinguish language from character shape. Western digits do not make a value English. An email address is not an English sentence. A model number may have no natural-language pronunciation at all. Those values may need LTR direction for visual order while having no meaningful lang="en" claim. Mark the property you know, and leave the other one alone when it does not apply.
Give each dynamic value its own node#
Static HTML makes the boundary obvious. Real messages are assembled from localized copy and values arriving from a database. Do not flatten a product name into one large text string and then wonder how to isolate it. Keep the localized words and the inserted value as separate nodes.
This browser-console example builds the same water-lab notice without treating a database value as HTML. It uses a fixed model value here so the snippet runs as shown; in the product, modelName would come from the record:
const modelName = "AquaSense+";
const notice = document.createElement("p");
notice.lang = "ar";
notice.dir = "rtl";
notice.append("اكتمل فحص جهاز ");
const model = document.createElement("bdi");
model.dir = "ltr";
model.textContent = modelName;
notice.append(model, " في المختبر");
document.body.append(notice);textContent inserts the model as text, not parsed markup. The bdi node keeps the value separate for bidi processing. The surrounding Arabic copy stays in its natural source order. This pattern is more important than the exact API: in a component framework, render a child element for the value instead of generating one undifferentiated sentence and trying to repair its visual order later.
The snippet builds one known Arabic sentence. It is not a localization architecture for every language. If your product uses message templates, let the translation determine where the model value belongs in the sentence, while the renderer keeps that value as a distinct node with its own direction. English and Arabic grammar do not promise the same word order. A hard-coded English prefix, then a value, then a suffix may preserve the bidi boundary and still produce unnatural Arabic. Keep the translator's sentence structure and the developer's value boundary as separate responsibilities.
Keep punctuation with the value when it belongs to the value. The + in AquaSense+ is part of the model name, so it is inside the isolated node. The period at the end of Calibration ready. is part of the quoted English phrase, so it is inside its LTR span. Arabic sentence punctuation belongs outside those inserted values. A boundary drawn one character too early is still the wrong boundary. Standards can only work with the markup you give them (a recurring theme, unfortunately).
Avoid adding invisible directional marks to translated strings as a first repair. They can be useful in plain-text channels that have no markup, but in HTML they are harder to inspect, copy, and maintain than an element with dir. W3C's inline-markup guidance describes both approaches. Prefer a visible DOM boundary where you control the HTML.
Known direction, unknown direction, and numeric-only values#
Use an explicit direction when you know it. A product model written in Latin script, an email address, a URL, or a fixed-order code should not ask the browser to infer its direction from whichever character happens to come first. An element with dir="ltr" isolates that value in HTML. It can be <bdi>, an existing link, or another appropriate inline element.
When the direction really is unknown, such as a user-entered note that could begin in Arabic or English, dir="auto" or <bdi> without an explicit direction can choose from the first strong character in the value. That is a first-strong heuristic, not language detection. A note beginning with an English model name and continuing in Arabic can choose LTR. A value containing only Western digits, spaces, and punctuation has no strong letter; in the checked Chromium behavior, it resolves to LTR when isolated with dir="auto" or <bdi>. If you already know that a field is a fixed-order identifier, explicit LTR is easier to reason about.
Scope dir="auto" to the unknown value, not a whole paragraph that also includes a known Arabic label. Otherwise the label's Arabic letter supplies the first strong character and the value never receives an independent base direction. This is a subtle failure because the markup contains the word auto, so it looks considerate. It is simply considering the wrong unit.
There is an edge to the explicit-LTR advice as well. Spaced groups of Arabic-Indic digits have different bidi behavior from Western digits. In the checked Chromium behavior, their groups can reverse even inside <bdi dir="ltr">. If your product has a fixed-order spaced Arabic-Indic notation, test that exact value rather than assuming the Western-digit fix applies. The broader bidi guide goes into those cases; this guide keeps its main example on Arabic-English inline markup.
Links, addresses, and values that must survive copying#
The water-lab page might include a support address in an Arabic help sentence. The visible address is one LTR value, while the sentence stays RTL:
<p lang="ar" dir="rtl">
للمساعدة، راسل <a href="mailto:service@riverlab.example" dir="ltr">service@riverlab.example</a>.
</p>The link element is already a suitable boundary; there is no need to add a nested <bdi> just to make the DOM look thorough. Keep the visible address and its destination consistent, and test the whole line in the browser. concerns unisolated email addresses, phone numbers, and Latin tokens inside Arabic text. For phone fields and tel: links, the phone-number guide covers input direction and dialing behavior in more detail.
Do not repair an address by reversing characters in the stored string. The W3C distinction between logical and visual order matters here: copying text gives its logical order, not a trustworthy report of what was displayed. An address can copy correctly and still look broken to the person trying to read it. Conversely, manually reversing the source to make one screenshot look right can corrupt the copied address or link target. Keep the data in its real order and fix presentation at the value boundary.
An English filename or URL may include punctuation, slashes, dots, and digits. Some combinations render acceptably without markup, others do not. Avoid claiming that every Latin token breaks or that one dir value fixes every possible format. Put the exact production value in an Arabic sentence, check the display, and isolate the value when its characters or neighboring punctuation move. If the string has meaning as a URL or filename, preserve its real syntax rather than inserting spaces to make the screenshot easier on the eye.
Test the sentence, not a detached sample#
A mixed-text fixture should contain the Arabic words around the value. The bare AquaSense+ token on its own does not reproduce the punctuation problem described above. Nor does a value that starts with Latin letters and has no trailing punctuation necessarily show any visible failure. Test the troublesome shape in its actual sentence and at the widths where that sentence wraps.
For this page's examples, a useful visual pass includes the unisolated model name, the isolated model name, and the intentional English status phrase with its final period inside an LTR span. At each step, check the actual rendered order and where punctuation lands. The measured behavior described here comes from Chromium 151; verify Firefox and Safari in your supported browser set rather than treating a Chromium fixture as universal proof.
Inspect the DOM source or textContent to confirm the data remains in logical order, but do not stop there. Selection and copying preserve the order the text was entered in and therefore cannot catch a display-only bidi problem. An accessibility-tree check also reads logical order. A visual assertion, screenshot review, or a person reading the rendered line must cover the presentation. If the phrase wraps onto two lines, inspect both lines; the Unicode algorithm's final reordering happens per displayed line.
Check the nearby language treatment too. A quoted English passage that intentionally remains English should carry lang="en"; concerns intentional foreign-language runs missing a language attribute. A plain English UI label on an Arabic route is a translation issue, not automatically an intentional foreign passage. concerns punctuation at the wrong visual end of Arabic text. Keep those findings separate so the remedy matches the defect.
A free Ritla scan can point to unisolated mixed-text hazards and misplaced punctuation on the page being examined. Then inspect the exact sentence: is the inserted value really one unit, is its punctuation inside that unit, and is an English passage tagged for the language it uses? The browser is quite willing to do the ordering. It does ask you to draw the boundaries.
Checks in this guide
- document direction missing, contradicted, or declared only in CSS
- Forced-LTR container, or a bidi override, over Arabic content
- Unisolated bidi hazards: phones, emails, Latin tokens inside Arabic text
- Intentional foreign-language runs missing a lang attribute
- Punctuation rendering at the wrong visual end of Arabic text