The bug report was one sentence: "Arabic goes to the left." It was reproducible in about
fifteen seconds. Open a note, type Hello, press Enter, paste a paragraph of Arabic.
The Arabic renders with its glyphs shaped and joined correctly — the browser handles that part
perfectly — but the paragraph sits hard against the left margin, with its ragged edge on the
right. In every word processor a reader of Arabic has ever used, it would be the other way round.
The editor had dir="auto" on the contenteditable element. That attribute exists
precisely to solve this. It was doing exactly what it is specified to do.
dir="auto" is not broken here — it is answering a different question than the one a notepad needs answered.What the attribute actually promises
The rule underneath dir="auto" is not a heuristic somebody at a browser vendor made
up. It is
Unicode Annex #9, the
Bidirectional Algorithm, and the relevant part is three rules long. Rule P1
splits text into paragraphs. Rule P2 says:
In each paragraph, find the first character of type L, AL, or R while skipping over any characters between an isolate initiator and its matching PDI.
And rule P3: if that character is AL or R, the
paragraph embedding level is 1 — right-to-left. Otherwise it is 0.
That is the whole "first strong character" rule, and it is genuinely good. It is what Word does.
It is what every messaging app does when it flips a chat bubble. The problem is the word
paragraph. In UAX #9 a paragraph is a real unit of text. In HTML,
dir="auto" resolves once for the element it sits on, and a
contenteditable div is one element no matter how many paragraphs the user types into
it.
The browser resolves one direction for the whole contenteditable from its very
first strong character. A note that opens with the word "Hello" will keep every Arabic paragraph
in it left-aligned for the life of the document — and because the glyph shaping still looks
perfect, nothing about it reads as a rendering bug. It reads as the app being careless.
Moving the decision down to the block
Once you accept that direction is a property of a paragraph and not of a document, the shape of
the fix is obvious: run the first-strong rule yourself, per block, and write the answer onto the
block as a real dir attribute. The browser then does everything else — glyph
shaping, mirrored punctuation, list marker placement, caret movement — from that one attribute.
The scanner is fifteen lines and it is the core of the module:
/** Any letter. Everything that is a letter but not RTL is strong LTR. */
const LETTER = /\p{L}/u;
export function detectDirection(text: string): 'ltr' | 'rtl' | null {
for (const ch of text) {
const cp = ch.codePointAt(0)!;
if (isRtlCodePoint(cp)) return 'rtl';
if (cp === 0x200e) return 'ltr'; // LEFT-TO-RIGHT MARK
if (LETTER.test(ch)) return 'ltr';
}
return null;
}
Three details in there are load-bearing and none of them are obvious.
for (const ch of text) iterates code points, not UTF-16 code units.
A for (let i = 0; i < text.length; i++) loop would hand you a lone surrogate for
anything above U+FFFF, and several of the scripts that matter here — Adlam, Mende Kikakui, Old
Hebrew, Avestan — live in the astral planes. With an index loop they silently never match.
\p{L} with the u flag means "any Unicode letter", which is what makes
the LTR branch work for Cyrillic, Greek, Devanagari, Thai, Han and everything else without
enumerating a single one of them. The RTL set has to be enumerated; the LTR set is the complement,
and complements are free.
And the return type has three states, not two. That third one is the subject of a section further down, because getting it wrong produces a caret that jumps around while you type.
Digits are not strong characters, and the range table has to say so
The naive RTL test is "is this code point in the Arabic block". The Arabic block is U+0600–U+06FF, so you write that range, and everything works until somebody starts a paragraph with an Arabic-Indic numeral.
U+0660–U+0669 are the digits ٠١٢٣٤٥٦٧٨٩. U+06F0–U+06F9 are the Extended Arabic-Indic
digits ۰۱۲۳۴۵۶۷۸۹ used in Persian and Urdu. Both sit inside the Arabic block, and neither is a
strong character — UAX #9 classes them as AN and EN. The same is true in
the other direction: a paragraph beginning "2024" must not resolve as left-to-right just because
European digits happen to come first.
So the table has holes punched in it, deliberately, and the holes are the interesting part:
const RTL_RANGES: ReadonlyArray<readonly [number, number]> = [
[0x0590, 0x05ff], // Hebrew
[0x0600, 0x065f], // Arabic — stops before the Arabic-Indic digits
[0x066a, 0x066a], // Arabic percent sign
[0x066d, 0x06ef], // Arabic — resumes after the digit separators
[0x06fa, 0x08ff], // Arabic ext., Syriac, Thaana, N'Ko, Samaritan, Mandaic
[0x200f, 0x200f], // RIGHT-TO-LEFT MARK
[0x202b, 0x202b], // RIGHT-TO-LEFT EMBEDDING
[0x202e, 0x202e], // RIGHT-TO-LEFT OVERRIDE
[0x2067, 0x2067], // RIGHT-TO-LEFT ISOLATE
[0xfb1d, 0xfdff], // Hebrew + Arabic presentation forms A
[0xfe70, 0xfefc], // Arabic presentation forms B
[0x10800, 0x10fff], // Cypriot, Phoenician, Old Hebrew, Avestan, …
[0x1e800, 0x1efff], // Mende Kikakui, Adlam, Arabic Mathematical
];
[0x0600, 0x06ff].The lookup exploits the fact that the table is sorted, so a Latin character exits on the first comparison rather than testing thirteen ranges:
function isRtlCodePoint(cp: number): boolean {
for (const [lo, hi] of RTL_RANGES) {
if (cp < lo) return false; // ranges are sorted — nothing later can match
if (cp <= hi) return true;
}
return false;
}
For an English note that early exit is the difference between one integer comparison per character and thirteen. This runs on every keystroke, so it matters more than it looks.
"No strong character" is a real answer, not a missing one
Press Enter at the end of an Arabic paragraph. The new paragraph is empty. It has no first strong character, because it has no characters at all.
If detectDirection returned 'ltr' for that — the obvious default — the
caret would snap from the right edge of the paper to the left the instant you pressed Enter, then
snap back to the right the instant you typed an Arabic letter. Every new line, twice. It is the
kind of thing that is hard to describe in a bug report and impossible to ignore once you have felt
it.
So the third state exists, and the block-level sync treats it as "keep whatever you have":
function syncBlock(el: HTMLElement, root: HTMLElement): void {
if (el.hasAttribute(MANUAL_ATTR)) return;
const dir = detectDirection(el.textContent ?? '');
if (dir === null) return; // undecidable — leave it alone
if (dir === 'rtl') setDir(el, 'rtl');
// Only spell out "ltr" when something above would otherwise make it RTL.
else setDir(el, inheritsRtl(el, root) ? 'ltr' : null);
}
A new empty paragraph after an Arabic one inherits rtl from the browser's normal
attribute inheritance, keeps it because the scanner declined to overrule, and the caret stays
where the user's eye already is. The moment they type a Latin letter it flips, once, deliberately.
Write "rtl" always, write "ltr" only when you have to
The last line of syncBlock is asymmetric on purpose. rtl is always
written out. ltr is written only when inheritsRtl() finds an RTL
ancestor that would otherwise cascade down onto this block — otherwise the attribute is removed
entirely.
The reason is that dir is in the saved HTML. Every tab's content is serialised to
storage as markup, and a dir="ltr" on every paragraph of every English note is pure
weight in every save, every export and every paste. Left-to-right is already the default;
restating it is only useful when something is actively contradicting it — an Arabic list item
inside an Arabic list, with one English entry in the middle.
Which elements count as a block
The list is longer than it first looks, and two of the entries are there for reasons that are not about text at all:
const BLOCK_SELECTOR =
'p,div,h1,h2,h3,h4,h5,h6,li,ul,ol,blockquote,pre,table,thead,tbody,tr,td,th';
ul and ol are in there alongside li because the list
indent is padding-inline-start on the list, not on the item. Stamping only the
li puts the marker on the correct side but leaves the whole list indented from the
wrong edge.
table is in there alongside td and th because an RTL table
lays its columns out right-to-left. Stamp only the cells and you get right-aligned text in
left-to-right columns, which is worse than either consistent option.
The CSS has to be logical, or the attribute achieves nothing
Stamping dir="rtl" tells the browser which way the text runs. It does not override a
stylesheet that has hardcoded which side things sit on. A single padding-left inside
the editor paper is enough to indent every RTL list from the wrong edge while its markers sit
correctly on the right — a layout that looks less like a bug and more like a design decision
nobody thought through.
padding-inline-start resolves against the element's own direction, so one declaration is correct in both. There is no RTL override sheet anywhere in this codebase.| Do not use inside the editor | Use instead |
|---|---|
padding-left / padding-right | padding-inline-start / -end |
margin-left / margin-right | margin-inline |
border-left / border-right | border-inline-start / -end |
text-align: left / right | text-align: start / end |
left / right offsets | inset-inline-start / -end |
The blockquote rule is the clearest example. It is
border-inline-start: 3px solid …, so on an English quote the rule sits on the left
and on an Arabic quote it sits on the right, from one declaration. The physical version needs two
rules and a selector that knows about direction, and it will be forgotten the first time somebody
adds a new block type.
Running this on every keystroke without walking the document
Direction has to update as you type — the first Arabic letter in an empty paragraph should flip it immediately, not on blur. But re-deriving every block in the note on every character is quadratic in a way that shows up on a long note.
So there are two entry points with very different scopes. A full pass runs on load and on paste, where the content is arbitrary and every block is suspect. A scoped pass runs on input, walking only from the caret's element up to the root:
export function updateDirectionAt(root: HTMLElement, node: Node | null): void {
if (isUnwrapped(root)) {
applyAutoDirection(root);
return;
}
let el = nearestElement(node);
if (!el || !root.contains(el)) return;
while (el && el !== root) {
if (el.matches(BLOCK_SELECTOR)) syncBlock(el, root);
el = el.parentElement;
}
}
input, deliberately outside the 100 ms change debounce — so the caret flips on the first Arabic keystroke rather than a tenth of a second later, and so the HTML autosave writes already carries the attribute.The bare-text case
A plain-text tab, or a fresh paste of unwrapped text, has no block elements at all — just text nodes directly under the paper. There is nothing to stamp. That case is checked explicitly and the direction goes onto the paper element itself. This is the one situation where the whole-region behaviour is correct, because the region genuinely is one paragraph.
Letting the user overrule the detection
First-strong is right almost always, and wrong in a specific, predictable case: a paragraph that
opens with a Latin brand name, a URL or a code identifier but is otherwise Arabic. The detector
sees the Latin letter, answers ltr, and is technically correct by the rule while
being obviously wrong to the person reading it.
The toolbar's LTR and RTL buttons and the Format menu pin a block by adding a valueless sentinel
attribute. syncBlock short-circuits on it in its first line, so a pinned block is
never reconsidered:
/** Marks a block whose direction the user set by hand — auto-detection skips it. */
const MANUAL_ATTR = 'data-dir-lock';
export function setSelectionDirection(root: HTMLElement, dir: 'ltr' | 'rtl' | 'auto'): void {
const blocks = selectedBlocks(root);
for (const el of blocks) {
if (dir === 'auto') {
el.removeAttribute(MANUAL_ATTR);
el.removeAttribute('dir');
} else {
el.setAttribute(MANUAL_ATTR, '');
el.setAttribute('dir', dir);
}
}
if (dir === 'auto') applyAutoDirection(root);
}
Three states rather than two — left, right, and auto. Without an explicit way back, a user who pins a paragraph once has no way to hand it back to the detector, and the third option costs one branch.
selectedBlocks uses range.intersectsNode() rather than walking from the
anchor and focus nodes, because a selection dragged across three paragraphs has boundary points in
only the first and last of them. The middle paragraph is fully inside the range and appears in
neither endpoint.
The state that leaks between tabs
This one was not on anybody's list. The editor paper is a single DOM element shared by every open tab — switching tabs replaces its contents, it does not create a new element. So an attribute written onto the paper itself belongs to the element, not to the note.
Pin direction on a plain-text tab and the lock goes onto the paper (the bare-text case above).
Switch to another tab and the lock is still there, silently suppressing detection on a note that
never asked for it. The fix is a single call in setContent, which runs on every tab
switch:
export function clearRootDirection(root: HTMLElement): void {
root.removeAttribute('dir');
root.removeAttribute(MANUAL_ATTR);
}
Any state written onto a reused element has to be cleared when the content behind it changes. Obvious in hindsight, invisible in the moment.
Paste, and the two attributes worth keeping
Pasted HTML goes through an allowlist before it enters the document — thirty tags, eight style properties, and exactly two attributes:
const ALLOWED_ATTRS = new Set(['dir', 'lang']);
Everything else is stripped, including id, class, and every event
handler. dir survives because the source document may have got the direction right
and there is no reason to throw that away and re-derive it. lang survives because it
drives font selection and hyphenation, and because a language tag is the only thing that
distinguishes Persian from Arabic once the text is in the DOM — they share a script and the
detector cannot tell them apart, but a font stack can.
After the insert, a full applyAutoDirection pass runs anyway. Pasted markup brings
its own blocks, and those blocks have never been through the detector.
Export is a separate problem entirely
Everything above produces a correctly-directioned document on screen. Word does not read any of it.
A .docx has no concept of "infer direction from the text". Both halves have to be
stated explicitly: <w:bidi/> in the paragraph properties to flip the
paragraph's layout, and <w:rtl/> in each run's properties to set that run's
reading order. Emit the first without the second and the text lands in the right place with its
characters running the wrong way.
The useful part is that the exporter does not re-implement the rule. It imports the same
detectDirection the editor uses and applies it per run, so a mostly-Latin sentence
with one Arabic word in it is not flipped wholesale, and the file agrees with the screen by
construction rather than by coincidence. That is written up in
writing a .docx in the browser with no
dependencies.
Direct PDF export sidesteps the problem by declining it. The writer draws with PDF's Standard-14
fonts, which cover Latin-1 only, and complex scripts additionally need shaping — contextual forms
in Arabic and Urdu, reordering and conjuncts in Bengali and Devanagari — that no font file alone
provides. Anything outside that repertoire is handed to the browser's own print pipeline, which
embeds real Unicode fonts, shapes correctly, and honours the dir attributes we spent
this whole article stamping. The reasoning is in
generating a PDF in the browser.
What we would tell someone starting this
Bidirectional text has a reputation for being a specialist topic, and the shaping and reordering parts genuinely are — but the browser already does all of that for free. What is left is a narrower and much more tractable problem: deciding, per paragraph, which way it runs, and then writing CSS that does not contradict the answer.
- Direction is a property of a block. Never of a document, and never of an editable region.
-
Never put
dir="auto"on acontenteditable. It resolves once for the whole region and the failure is invisible until the second paragraph. -
Digits are weak types. Punch
U+0660–0669andU+06F0–06F9out of any Arabic range you write by hand. - Iterate code points, not code units, or every astral-plane RTL script silently fails to match.
- "No strong character" is a third answer. Inherit on it — do not fall back to LTR, or the caret jumps on every Enter.
-
Write
rtlalways; writeltronly when an ancestor would otherwise override it. The attribute ends up in every save and every export. -
Ban physical
left/rightCSS inside the editor. Logical properties are one declaration that is correct in both directions. - Give the user a manual override and a way back to automatic.
- Clear per-element state on tab switch. A shared DOM node outlives the content inside it.
- Export formats infer nothing. State the direction explicitly, and share one detector between the screen and the file so they cannot drift apart.
The whole module is 197 lines. Most of that is the range table and comments explaining why the table has holes in it.