How to detect text direction in HTML

Learn how dir="auto", bdi and JavaScript direction checks handle unknown text, where first-strong detection fails and what metadata to preserve.

Guide14 min read

Text direction is easy when the application owns the content. Arabic interface copy is RTL; English interface copy is LTR. Detection becomes necessary when the application renders a seller name, support message, CMS entry or search query whose direction was not known when the component was built.

The first mistake is treating that as language detection. dir="auto" does not decide whether a string is Arabic or English. It scans for a strong directional character and uses the first one it finds. JavaScript implementations often imitate that heuristic with a narrower script check.

Use explicit direction whenever the content contract already tells you the answer. Use detection only at boundaries where direction is genuinely unknown, and preserve an explicit answer when the producer can supply one.

Choose declaration before detection#

Start by classifying the content source.

ContentDirection strategy
Arabic application interfaceInherit dir="rtl" from the document
Known English quotation inside ArabicSet lang="en" dir="ltr" on the quotation
Fixed-format email, URL or machine identifierUse the value's documented LTR boundary
User comment with unknown languageUse dir="auto" on the comment container
Unknown inline name inside surrounding textIsolate it with <bdi> or a tight dir="auto" boundary
Text submitted with direction metadataRender the stored ltr or rtl value explicitly
Plain-text system with no HTML direction supportDetect or carry direction before rendering

Detection is a fallback for missing metadata. It should not replace known product rules. An Arabic route should not put dir="auto" on <html> and hope the first heading chooses correctly. Set the document to dir="rtl"; detect only the unknown content inside it.

The same applies to machine values. If every marketplace listing ID is defined as LTR, declare that boundary. Running a heuristic on each ID adds uncertainty without learning anything.

How dir="auto" decides#

The HTML dir attribute accepts ltr, rtl and auto. With auto, the browser examines the content and derives a base direction from the first strongly directional character. MDN's dir reference documents the algorithm and the elements skipped during the scan. W3C guidance recommends the value for inserted text whose direction is not known in advance.

html
<article class="seller-message" dir="auto">
  <!-- User-provided message is inserted here as text. -->
</article>

If the first strong character is Arabic, Hebrew or another RTL type, the element resolves RTL. A Latin letter resolves it LTR. Leading spaces, punctuation, emoji and digits do not establish the base direction, so the scan continues until it finds a strong character.

In Chromium 151, measured on 2026-09-24, a value containing only Western digits, spaces and punctuation resolves LTR. That makes a numeric-only value look plausible, but it does not tell you whether the value should be treated as an LTR identifier or an RTL-readable range. Meaning still owns the final decision.

You can inspect the resolved direction with computed style:

js
// element is the node carrying dir="auto".
const resolvedDirection = getComputedStyle(element).direction;

The element's dir property still reports auto; computed style reports the resolved ltr or rtl. Test both when the distinction matters.

First strong is not majority script#

The heuristic stops at the first qualifying character. It does not count Arabic and Latin letters, translate the text, identify the dominant language or reconsider its answer later in the string.

This user message resolves LTR because the brand name begins with a Latin letter:

html
<p dir="auto">Nova أرسلت المنتج الخطأ وأحتاج إلى استبداله</p>

Most of the sentence is Arabic, but that fact is irrelevant to first-strong detection. If the product knows that messages in this queue are Arabic, set dir="rtl". If the language is genuinely mixed and unknown, accept that auto is a heuristic and give users or moderators a way to correct the rare wrong result.

The reverse can happen when an English sentence begins with an Arabic quotation:

html
<p dir="auto">خصم خاص was the heading shown in the promotion.</p>

The first Arabic letter makes the paragraph RTL even though the surrounding sentence is English. Detection cannot infer which phrase is the quotation and which is the container language.

W3C's guidance on string direction metadata makes the practical hierarchy clear: explicit direction metadata is preferable; first-strong is a fallback when metadata is unavailable. Store the producer's answer when your content model can carry it.

Neutral prefixes are skipped#

Developers sometimes strip punctuation or emoji before detecting direction. HTML's algorithm already skips characters that do not establish strong direction, so preprocessing can change content without improving the decision.

These prefixes do not decide direction by themselves:

  • whitespace and line breaks;
  • common punctuation;
  • emoji;
  • Western digits;
  • many symbols and separators.

The letter after the prefix becomes decisive. A hashtag beginning with an Arabic letter resolves RTL; one beginning with a Latin letter resolves LTR. A numeric-only coupon token has no strong letter, so the fallback behavior applies.

Do not remove a leading mention, URL or product code to make the detector find the “real” sentence unless that rule is part of the product's content model. Once you start skipping domain-specific prefixes, you have designed a custom heuristic. Name it, test it and keep a manual override.

HTML's scan also ignores text inside certain nested elements when determining an ancestor's auto direction, including <bdi> and elements with their own valid dir attribute. That lets an isolated username avoid deciding the direction of the surrounding message. The details are documented by MDN and covered by the W3C direction tests.

Use <bdi> for unknown inline values#

An unknown inline value needs two things: a direction for itself and isolation from the surrounding sentence. <bdi> provides an isolated boundary and behaves with automatic direction when no explicit dir is supplied.

html
<p dir="rtl">
  كتب <bdi class="seller-name">North Gate</bdi> تعليقًا جديدًا
</p>

The seller name determines its own base direction without allowing its characters and punctuation to merge into the surrounding Arabic bidi run.

A plain span does not isolate anything:

html
<p dir="rtl">
  كتب <span class="seller-name">North Gate</span> تعليقًا جديدًا
</p>

Changing the tag name is not cosmetic here. R021 reports unisolated bidi hazards such as phones, emails and Latin tokens inside Arabic text.

Use the smallest complete semantic unit. If a seller name ends with punctuation that belongs to the name, keep that punctuation inside the <bdi>. If punctuation belongs to the Arabic sentence, leave it outside. The bidirectional text guide explains the boundary choices in more depth.

When the direction is known, you may use an existing semantic element with explicit dir instead of <bdi>:

html
<cite lang="en" dir="ltr">Design Systems Weekly!</cite>

Detection is unnecessary because the content contract already supplies language and direction.

Use dir="auto" on unknown blocks#

For a standalone comment, review or CMS block, put dir="auto" on the element that owns the whole string.

html
<article class="review" dir="auto">
  <!-- One review, stored as plain text. -->
</article>

Do not put dir="auto" on a wrapper containing labels, metadata and the unknown text. The first strong character in a fixed label can choose direction before the browser reaches the review.

Problem

html
<article dir="auto">
  <h3>Review</h3>
  <p class="review-text">...</p>
</article>

The Latin heading decides the whole article's direction.

Better

html
<article>
  <h3>Review</h3>
  <p class="review-text" dir="auto">...</p>
</article>

The unknown content gets its own boundary while the fixed interface keeps the document's direction. The component performs the same job in both examples; only the detection boundary changes.

For rich text, direction can vary by paragraph. A single dir="auto" on the rich-text root chooses from the first qualifying content for that element, not an independent result for every arbitrary descendant. Store paragraph direction or render paragraph-level boundaries when the editor supports multilingual documents.

Inputs need an empty-state decision#

dir="auto" can improve an input for text that may be Arabic or English:

html
<label for="marketplace-search">Search</label>
<input
  id="marketplace-search"
  name="query"
  type="search"
  dir="auto"
>

As the user types, the first strong character determines direction. This works for a search box that genuinely accepts both Arabic and English.

The empty state is the edge case. In the measured Chromium version, an empty auto-direction input resolves LTR. An Arabic placeholder can therefore appear at the left edge until an Arabic character is entered. Do not use auto for a field known to expect Arabic solely to obtain dynamic alignment. Let it inherit RTL.

Separate content types:

  • known Arabic names, addresses and notes inherit RTL;
  • unknown multilingual search or message input can use auto;
  • email, URL and phone fields stay LTR;
  • fixed-format account and listing codes use their documented direction.

R023 reports text inputs rendered LTR where Arabic input is expected. R070 reports tel, email or url inputs rendered RTL. Automatic detection does not remove the need for field contracts.

Browsers may also let users change an input's direction manually. Do not overwrite that choice on every keystroke with JavaScript. If direction matters after submission, capture it as metadata.

Capture submitted direction with dirname#

HTML can submit a control's direction alongside its value through the dirname attribute. MDN's dirname reference documents the submitted ltr or rtl field.

html
<label for="seller-reply">Reply</label>
<textarea
  id="seller-reply"
  name="reply"
  dir="auto"
  dirname="reply.dir"
></textarea>

A form submission includes the reply plus a reply.dir value. Validate the metadata on the server, store only supported values and render it back as an explicit dir attribute.

This is useful when:

  • the user can manually change input direction;
  • the first-strong result should survive editing and later rendering;
  • the same text appears in email, web and exported documents;
  • downstream systems need direction without rerunning an HTML heuristic.

Direction metadata is not language metadata. Store language separately when pronunciation, translation, spellchecking or font choice depends on it.

Do not accept arbitrary attribute text from the database and concatenate it into HTML. Map the stored value to a closed ltr or rtl enum through the rendering framework.

textarea and pre handle paragraphs specially#

Multiline plain-text controls cannot wrap each paragraph in its own HTML element. HTML defines special handling for dir="auto" on <textarea> and <pre> so paragraph direction can be determined separately.

That behavior helps multilingual notes, but it also means the visible control may contain lines aligned in different directions. Test the editing experience with actual paragraph breaks, not just one line.

Check:

  • an Arabic paragraph followed by an English paragraph;
  • an English paragraph beginning with an Arabic quotation;
  • a blank first line;
  • pasted text with multiple newline conventions;
  • punctuation at each paragraph boundary;
  • selection and caret movement across paragraphs;
  • submitted direction metadata if the form stores only one direction value.

One dirname result cannot describe every paragraph of a multilingual textarea. If downstream rendering must preserve independent paragraph direction, the data model needs paragraph-level structure or a safe rich-text format. A single field-level flag is not enough.

Use JavaScript only when HTML cannot answer in time#

Prefer dir="auto" when the browser is rendering unknown HTML text. It implements the platform algorithm and keeps the direction declaration next to the content.

JavaScript detection is justified when:

  • a non-DOM renderer needs direction before drawing;
  • a component API requires ltr or rtl before mounting;
  • the server stores direction metadata for another channel;
  • analytics need a coarse script-direction category;
  • layout selection happens before the text enters HTML.

Avoid client detection followed by a second independent server detector. If their character tables differ, the saved and rendered directions can disagree.

JavaScript has Unicode property escapes for scripts and letters, but it does not expose the full Unicode bidi class algorithm as a simple built-in API. A regex based on Arabic Unicode blocks is not equivalent to HTML first-strong. It can classify digits, combining marks or punctuation incorrectly, and it can omit RTL scripts outside the ranges the developer remembered.

If your product contract permits only Arabic and Latin letters, implement and name that narrower detector honestly.

A product-scoped JavaScript detector#

This function detects the first Arabic-script letter or Latin-script letter. It deliberately returns null for digits, punctuation, emoji and unsupported scripts.

js
const ARABIC_LETTER = /^(?=\p{Letter}$)\p{Script=Arabic}$/u;
const LATIN_LETTER = /^(?=\p{Letter}$)\p{Script=Latin}$/u;

export function detectArabicLatinDirection(text) {
  for (const character of text) {
    if (ARABIC_LETTER.test(character)) return "rtl";
    if (LATIN_LETTER.test(character)) return "ltr";
  }

  return null;
}

The lookahead ensures that an Arabic-script character is also a Unicode letter. That prevents Arabic-script digits or non-letter marks from becoming the first “strong” result in this limited detector.

The function is not a complete Unicode direction algorithm. It does not classify Hebrew, Syriac, Thaana, N'Ko, Adlam or other scripts. It also does not reproduce HTML's tree traversal, because it receives plain text rather than elements. Use it only where the accepted languages are Arabic and Latin-script languages, and keep null as a real outcome.

Do not default null to RTL because the surrounding product is Arabic. A numeric identifier can require LTR, while a value with no textual direction may be better rendered from explicit schema metadata. Let the caller apply the field contract.

Test the function with:

js
// expect is supplied by the test runner.
expect(detectArabicLatinDirection("   مرحبًا")).toBe("rtl");
expect(detectArabicLatinDirection("✨ Welcome")).toBe("ltr");
expect(detectArabicLatinDirection("738-05-4162")).toBeNull();
expect(detectArabicLatinDirection("Nova مرحبًا")).toBe("ltr");

The last result demonstrates the heuristic's limitation rather than a bug in the function.

Avoid “any RTL character” detection#

A common JavaScript shortcut searches the entire string for one character in an RTL Unicode range. If it finds one, the function returns RTL. That is not first-strong detection.

Consider an English support reply that quotes one Arabic word near the end. An any-RTL detector marks the whole reply RTL even though its first strong character and container language are LTR. The opposite shortcut, “RTL only when Arabic letters are the majority,” fails a short Arabic sentence containing a long URL or product code.

Each heuristic answers a different question:

HeuristicQuestion it actually answersFailure to expect
HTML first strongWhat direction does the first strong character suggest?Leading opposite-script names or quotations
Any RTL characterDoes the string contain at least one matched RTL character?LTR text containing an RTL quotation
Majority scriptWhich counted script has more matched characters?Codes, URLs and short mixed sentences distort the count
Language detectorWhich language resembles this text?Language confidence is not direction metadata
Explicit metadataWhat direction did the producer declare?Incorrect or stale producer data

Do not label these functions detectDirection without documenting which rule they implement. A precise name such as detectArabicLatinDirection makes the limited script contract visible at each call site.

If product requirements favor a custom heuristic, keep examples of its known failures in tests and expose an override. Do not silently replace HTML first-strong with “any RTL” because it appeared simpler in a regular expression.

Make direction part of the component API#

Unknown text should enter the component tree through a boundary that accepts direction metadata. Prefer an explicit value, then fall back to auto.

jsx
// direction is "ltr", "rtl", or undefined when the producer has no answer.
export function InlineDirectedValue({ value, direction }) {
  return <bdi dir={direction ?? "auto"}>{value}</bdi>;
}

For block content, use a block owner rather than reusing the inline component:

jsx
// className is supplied by the caller for presentation only.
export function DirectedParagraph({ value, direction, className }) {
  return (
    <p className={className} dir={direction ?? "auto"}>
      {value}
    </p>
  );
}

Keep direction out of styling props such as align="right". Alignment and direction are different. A centered paragraph still needs a base direction for punctuation and run ordering.

Also keep language separate:

jsx
// language and direction are independent metadata supplied by the content model.
export function DirectedQuote({ value, language, direction }) {
  return (
    <blockquote lang={language} dir={direction ?? "auto"}>
      {value}
    </blockquote>
  );
}

Central components reduce improvised regex checks across cards, lists and dialogs. They also give tests one place to verify that explicit metadata wins and the fallback is exactly auto, not a second client-side guess.

Preserve direction in the content model#

If authors, translators or users can supply direction, store it with the string. A useful record separates value, language and direction:

json
{
  "value": "...",
  "language": "ar",
  "direction": "rtl"
}

Each field answers a different question. Language supports pronunciation, spellchecking, translation and locale behavior. Direction controls bidi layout. The text is the content itself.

Validate direction as ltr, rtl or an explicitly supported auto state. Prefer storing the resolved value when a trusted producer knows it. Use auto when consumers are expected to run their own first-strong detection and the heuristic is acceptable.

Preserving metadata matters when content leaves the web page. Email templates, native apps, PDFs and support exports may not share the browser's surrounding direction. Re-detecting in every channel invites inconsistent answers.

Do not assume a language tag always determines direction. Some languages can be written in more than one script, and product content can mix languages. Treat the two fields independently even when the current Arabic use case maps neatly to RTL.

Sanitize content without deleting direction#

User-generated HTML needs sanitization, but sanitization and direction detection solve different problems.

Prefer storing user text as text rather than HTML where rich formatting is not required. When rich text is allowed, use an allowlist that preserves supported dir and lang metadata on permitted elements while removing scripts, event handlers and unsafe URLs.

Do not run direction detection against raw markup. Tag names, attributes and hidden content are not the visible string. Detect after parsing and sanitization, or let the browser apply dir="auto" to the rendered content boundary.

Be explicit about nested boundaries. A sanitized paragraph with dir="ltr" should not be overwritten by a parent-level script counter. A <bdi> username should not decide the direction of the surrounding auto block. HTML's algorithm already defines those traversal rules; a plain-text regex does not.

Test detection with adversarial fixtures#

A passing Arabic sentence and English sentence prove only the obvious cases. Build fixtures around the boundaries of the algorithm.

FixtureExpected decisionWhy it matters
Arabic letter firstRTLOrdinary Arabic content
Latin letter firstLTROrdinary Latin content
Emoji then ArabicRTLNeutral prefix is skipped
Digits and punctuation onlyProduct fallbackNo strong letter
Latin brand then Arabic sentenceLTR under first strongKnown false classification for majority content
Arabic quotation then English sentenceRTL under first strongOpposite false classification
Isolated name before Arabic messageMessage uses its own first strongNested boundary must not take over
Empty multilingual inputBrowser fallbackPlaceholder and alignment edge case
Multiline textareaPer-paragraph behaviorOne field can show both directions

For browser tests, assert more than alignment:

js
// page and expect are supplied by the browser test runner.
const message = page.locator("[data-test=seller-message]");

await expect(message).toHaveAttribute("dir", "auto");

const direction = await message.evaluate(
  element => getComputedStyle(element).direction
);

expect(direction).toBe("rtl");

Run a separate fixture for every row whose result matters. Avoid a screenshot-only assertion: a centered card can hide the direction while punctuation and mixed runs remain wrong.

Test in every supported browser when the product depends on a specific edge. The browser facts called out in this guide were measured in that version only; the cited HTML behavior defines the target, but the supported matrix still needs product verification.

Debug the boundary before the algorithm#

When auto direction appears wrong, inspect what the browser actually scanned.

  1. Confirm the dir="auto" attribute sits on the text owner, not a wrapper with fixed labels.
  2. Inspect the first strong character, including invisible or nested content.
  3. Check for a child with its own dir or a <bdi> boundary that the parent scan skips.
  4. Read the element's computed direction.
  5. Compare the resolved result with the product's intended content direction.
  6. Decide whether the content needs explicit metadata instead of a heuristic.

Do not “fix” the result with text-align. Alignment does not change base direction, punctuation behavior or run ordering. Do not prepend invisible direction marks to every stored string either; that mutates content and spreads a rendering workaround into data.

If the text is known Arabic, replace auto with rtl. If it is truly unknown but first-strong is wrong for a valid case, expose an author override or store direction with the content. The detector is doing what it was designed to do.

Use detection as a fallback, not a guess everywhere#

Text direction works best as metadata. Declare it for application copy and typed values, capture it when users can supply it, and use dir="auto" only where the content arrives without an answer. JavaScript should fill a specific non-HTML or pre-render need, not compete with the browser's direction model.

A Ritla scan can surface mixed-direction values and field-direction problems after unknown content reaches the rendered page. Use the bidirectional text guide to fix inline boundaries, and the HTML dir="rtl" guide to keep detected content inside a correctly directed document.

Checks in this guide

See what your Arabic pages are hiding

Paste a URL. Ritla renders the page on desktop and mobile, runs every check, and shows the top issues with screenshot evidence.