Notepad Neo

Why dir="auto" breaks a mixed-language editor

One attribute is supposed to solve right-to-left text. It solves it exactly once per editable region, which in a notepad is the wrong granularity — and the failure is invisible until somebody writes their second paragraph.

DH

— builds and maintains Notepad Neo

· updated · 12 min read

The bug report was one sentence: "Arabic goes to the left." It was reproducible in about fifteen seconds. Open a note, type Hello, press Enter, paste a paragraph of Arabic. The Arabic renders with its glyphs shaped and joined correctly — the browser handles that part perfectly — but the paragraph sits hard against the left margin, with its ragged edge on the right. In every word processor a reader of Arabic has ever used, it would be the other way round.

The editor had dir="auto" on the contenteditable element. That attribute exists precisely to solve this. It was doing exactly what it is specified to do.

One resolved direction for a whole editor, versus one per paragraph On the left, dir="auto" on the editable root resolves a single direction from the note's very first strong character — the H of Hello — so the Arabic paragraphs below inherit left-to-right and align to the left edge. On the right, each block is stamped from its own first strong character, so the English paragraph stays left-aligned and the Arabic paragraphs align right. ONE DECISION PER EDITOR ONE DECISION PER BLOCK <div contenteditable dir="auto"> <div contenteditable> Hello — a note about travel first strong char: H → level 0 Hello — a note about travel dir="ltr" inherited مرحبا بالعالم ✕ left مرحبا بالعالم dir="rtl" الرحلة الأولى الرحلة الثانية ✕ bullets left الرحلة الأولى الرحلة الثانية bullets right Back to English. Back to English. The whole region resolved once, from the H in “Hello”. Every paragraph after it inherits that answer. Each block resolves from its own first strong character. Mixing scripts in one note costs nothing.
The same note, same content, same browser. The only difference is where the direction decision is made. dir="auto" is not broken here — it is answering a different question than the one a notepad needs answered.

What the attribute actually promises

The rule underneath dir="auto" is not a heuristic somebody at a browser vendor made up. It is Unicode Annex #9, the Bidirectional Algorithm, and the relevant part is three rules long. Rule P1 splits text into paragraphs. Rule P2 says:

In each paragraph, find the first character of type L, AL, or R while skipping over any characters between an isolate initiator and its matching PDI.

And rule P3: if that character is AL or R, the paragraph embedding level is 1 — right-to-left. Otherwise it is 0.

That is the whole "first strong character" rule, and it is genuinely good. It is what Word does. It is what every messaging app does when it flips a chat bubble. The problem is the word paragraph. In UAX #9 a paragraph is a real unit of text. In HTML, dir="auto" resolves once for the element it sits on, and a contenteditable div is one element no matter how many paragraphs the user types into it.

Rules P2 and P3 of the Unicode Bidirectional Algorithm Scanning a string left to right, weak types such as digits, punctuation and whitespace are skipped. The scan halts on the first strong character. If that character is class R or AL the paragraph level is set to 1 (right to left); if it is class L the level is 0. If no strong character exists the rule yields no answer. SCANNING “2024 — مرحبا” 2 0 2 4 EN ␠ WS — ON ␠ WS م AL ر ح skip every weak type… …halt on the first strong one P2 — first L, AL or R found: AL P3 — AL or R ? yes embedding level 1 right-to-left Digits are class EN or AN — weak, never strong. “2024” cannot decide anything, which is why the scan has to keep going until it reaches a letter.
Rules P2 and P3 in full. The interesting half is what gets skipped: whitespace, punctuation, and — the one that catches people — digits.
Never put dir="auto" on the editable root

The browser resolves one direction for the whole contenteditable from its very first strong character. A note that opens with the word "Hello" will keep every Arabic paragraph in it left-aligned for the life of the document — and because the glyph shaping still looks perfect, nothing about it reads as a rendering bug. It reads as the app being careless.

Moving the decision down to the block

Once you accept that direction is a property of a paragraph and not of a document, the shape of the fix is obvious: run the first-strong rule yourself, per block, and write the answer onto the block as a real dir attribute. The browser then does everything else — glyph shaping, mirrored punctuation, list marker placement, caret movement — from that one attribute.

The scanner is fifteen lines and it is the core of the module:

/** Any letter. Everything that is a letter but not RTL is strong LTR. */
const LETTER = /\p{L}/u;

export function detectDirection(text: string): 'ltr' | 'rtl' | null {
  for (const ch of text) {
    const cp = ch.codePointAt(0)!;
    if (isRtlCodePoint(cp)) return 'rtl';
    if (cp === 0x200e) return 'ltr';   // LEFT-TO-RIGHT MARK
    if (LETTER.test(ch)) return 'ltr';
  }
  return null;
}

Three details in there are load-bearing and none of them are obvious.

for (const ch of text) iterates code points, not UTF-16 code units. A for (let i = 0; i < text.length; i++) loop would hand you a lone surrogate for anything above U+FFFF, and several of the scripts that matter here — Adlam, Mende Kikakui, Old Hebrew, Avestan — live in the astral planes. With an index loop they silently never match.

\p{L} with the u flag means "any Unicode letter", which is what makes the LTR branch work for Cyrillic, Greek, Devanagari, Thai, Han and everything else without enumerating a single one of them. The RTL set has to be enumerated; the LTR set is the complement, and complements are free.

And the return type has three states, not two. That third one is the subject of a section further down, because getting it wrong produces a caret that jumps around while you type.

Digits are not strong characters, and the range table has to say so

The naive RTL test is "is this code point in the Arabic block". The Arabic block is U+0600–U+06FF, so you write that range, and everything works until somebody starts a paragraph with an Arabic-Indic numeral.

U+0660–U+0669 are the digits ٠١٢٣٤٥٦٧٨٩. U+06F0–U+06F9 are the Extended Arabic-Indic digits ۰۱۲۳۴۵۶۷۸۹ used in Persian and Urdu. Both sit inside the Arabic block, and neither is a strong character — UAX #9 classes them as AN and EN. The same is true in the other direction: a paragraph beginning "2024" must not resolve as left-to-right just because European digits happen to come first.

So the table has holes punched in it, deliberately, and the holes are the interesting part:

const RTL_RANGES: ReadonlyArray<readonly [number, number]> = [
  [0x0590, 0x05ff],   // Hebrew
  [0x0600, 0x065f],   // Arabic — stops before the Arabic-Indic digits
  [0x066a, 0x066a],   // Arabic percent sign
  [0x066d, 0x06ef],   // Arabic — resumes after the digit separators
  [0x06fa, 0x08ff],   // Arabic ext., Syriac, Thaana, N'Ko, Samaritan, Mandaic
  [0x200f, 0x200f],   // RIGHT-TO-LEFT MARK
  [0x202b, 0x202b],   // RIGHT-TO-LEFT EMBEDDING
  [0x202e, 0x202e],   // RIGHT-TO-LEFT OVERRIDE
  [0x2067, 0x2067],   // RIGHT-TO-LEFT ISOLATE
  [0xfb1d, 0xfdff],   // Hebrew + Arabic presentation forms A
  [0xfe70, 0xfefc],   // Arabic presentation forms B
  [0x10800, 0x10fff], // Cypriot, Phoenician, Old Hebrew, Avestan, …
  [0x1e800, 0x1efff], // Mende Kikakui, Adlam, Arabic Mathematical
];
The Arabic block with its digit ranges excluded A number line from U+0590 to U+08FF. Hebrew and the Arabic letter ranges are marked as strong right-to-left. Two gaps are cut out of the Arabic block: U+0660 to U+0669, the Arabic-Indic digits, and U+06F0 to U+06F9, the extended Arabic-Indic digits, both of which are weak types and must not decide a paragraph's direction. U+0590 → U+08FF Hebrew 0590–05FF Arabic letters 0600–065F digits 0660–69 Arabic 066D–06EF digits 06F0–F9 Syriac · Thaana · N'Ko · Samaritan 06FA–08FF class AN / EN — weak, not strong ٠١٢٣٤٥٦٧٨٩ ۰۱۲۳۴۵۶۷۸۹ “2024 مرحبا” must resolve right-to-left, from the م. If the digit ranges were inside the table, “٢٠٢٤ Hello” would resolve RTL — and that is equally wrong.
Three separate ranges where one would look tidier. The two gaps are the whole reason the table is not simply [0x0600, 0x06ff].

The lookup exploits the fact that the table is sorted, so a Latin character exits on the first comparison rather than testing thirteen ranges:

function isRtlCodePoint(cp: number): boolean {
  for (const [lo, hi] of RTL_RANGES) {
    if (cp < lo) return false;   // ranges are sorted — nothing later can match
    if (cp <= hi) return true;
  }
  return false;
}

For an English note that early exit is the difference between one integer comparison per character and thirteen. This runs on every keystroke, so it matters more than it looks.

"No strong character" is a real answer, not a missing one

Press Enter at the end of an Arabic paragraph. The new paragraph is empty. It has no first strong character, because it has no characters at all.

If detectDirection returned 'ltr' for that — the obvious default — the caret would snap from the right edge of the paper to the left the instant you pressed Enter, then snap back to the right the instant you typed an Arabic letter. Every new line, twice. It is the kind of thing that is hard to describe in a bug report and impossible to ignore once you have felt it.

So the third state exists, and the block-level sync treats it as "keep whatever you have":

function syncBlock(el: HTMLElement, root: HTMLElement): void {
  if (el.hasAttribute(MANUAL_ATTR)) return;
  const dir = detectDirection(el.textContent ?? '');
  if (dir === null) return;                 // undecidable — leave it alone
  if (dir === 'rtl') setDir(el, 'rtl');
  // Only spell out "ltr" when something above would otherwise make it RTL.
  else setDir(el, inheritsRtl(el, root) ? 'ltr' : null);
}

A new empty paragraph after an Arabic one inherits rtl from the browser's normal attribute inheritance, keeps it because the scanner declined to overrule, and the caret stays where the user's eye already is. The moment they type a Latin letter it flips, once, deliberately.

Write "rtl" always, write "ltr" only when you have to

The last line of syncBlock is asymmetric on purpose. rtl is always written out. ltr is written only when inheritsRtl() finds an RTL ancestor that would otherwise cascade down onto this block — otherwise the attribute is removed entirely.

The reason is that dir is in the saved HTML. Every tab's content is serialised to storage as markup, and a dir="ltr" on every paragraph of every English note is pure weight in every save, every export and every paste. Left-to-right is already the default; restating it is only useful when something is actively contradicting it — an Arabic list item inside an Arabic list, with one English entry in the middle.

Which elements count as a block

The list is longer than it first looks, and two of the entries are there for reasons that are not about text at all:

const BLOCK_SELECTOR =
  'p,div,h1,h2,h3,h4,h5,h6,li,ul,ol,blockquote,pre,table,thead,tbody,tr,td,th';

ul and ol are in there alongside li because the list indent is padding-inline-start on the list, not on the item. Stamping only the li puts the marker on the correct side but leaves the whole list indented from the wrong edge.

table is in there alongside td and th because an RTL table lays its columns out right-to-left. Stamp only the cells and you get right-aligned text in left-to-right columns, which is worse than either consistent option.

Which blocks in a mixed note receive a dir attribute A DOM tree under the editor paper. An English paragraph gets no attribute because left-to-right is already the default. An Arabic heading and paragraph get dir="rtl". An Arabic list gets dir="rtl" on both the ul and its li children so the indent and the markers both move. One paragraph carries data-dir-lock, set by the user from the Format menu, and auto-detection skips it entirely. #editor-paper <p> Trip notes no attribute — LTR is the default <h2 dir="rtl"> الرحلة first strong char is AL <ul dir="rtl"> flips padding-inline-start <li dir="rtl"> flips the ::marker side <li dir="rtl"> <p dir="ltr" data-dir-lock> pinned by hand — sync skips it
The stamping is sparse by design. Only blocks that need to contradict the default carry an attribute, and a locked block is never touched again by the detector.

The CSS has to be logical, or the attribute achieves nothing

Stamping dir="rtl" tells the browser which way the text runs. It does not override a stylesheet that has hardcoded which side things sit on. A single padding-left inside the editor paper is enough to indent every RTL list from the wrong edge while its markers sit correctly on the right — a layout that looks less like a bug and more like a design decision nobody thought through.

Physical versus logical CSS on a right-to-left list The same Arabic list styled two ways. With padding-left the list indents from the left edge while its bullets sit on the right, leaving a gap on the wrong side. With padding-inline-start the indent follows the reading direction and the list sits correctly against the right margin. padding-left: 2em padding-inline-start: 2em indent العنصر الأول العنصر الثاني العنصر الثالث ✕ dead space on the wrong side indent العنصر الأول العنصر الثاني العنصر الثالث ✓ indent follows the reading edge The markers are on the right in both. Only the box the list sits in is wrong on the left — which is exactly why this survives a casual look at the screenshot.
padding-inline-start resolves against the element's own direction, so one declaration is correct in both. There is no RTL override sheet anywhere in this codebase.
Do not use inside the editorUse instead
padding-left / padding-rightpadding-inline-start / -end
margin-left / margin-rightmargin-inline
border-left / border-rightborder-inline-start / -end
text-align: left / righttext-align: start / end
left / right offsetsinset-inline-start / -end

The blockquote rule is the clearest example. It is border-inline-start: 3px solid …, so on an English quote the rule sits on the left and on an Arabic quote it sits on the right, from one declaration. The physical version needs two rules and a selector that knows about direction, and it will be forgotten the first time somebody adds a new block type.

Running this on every keystroke without walking the document

Direction has to update as you type — the first Arabic letter in an empty paragraph should flip it immediately, not on blur. But re-deriving every block in the note on every character is quadratic in a way that shows up on a long note.

So there are two entry points with very different scopes. A full pass runs on load and on paste, where the content is arbitrary and every block is suspect. A scoped pass runs on input, walking only from the caret's element up to the root:

export function updateDirectionAt(root: HTMLElement, node: Node | null): void {
  if (isUnwrapped(root)) {
    applyAutoDirection(root);
    return;
  }
  let el = nearestElement(node);
  if (!el || !root.contains(el)) return;
  while (el && el !== root) {
    if (el.matches(BLOCK_SELECTOR)) syncBlock(el, root);
    el = el.parentElement;
  }
}
Full-document sync versus caret-scoped sync On load and paste every block in the note is re-derived. On each keystroke only the chain of block ancestors above the caret is re-derived, typically two or three elements, regardless of how long the note is. ON LOAD · ON PASTE ON EVERY KEYSTROKE re-derived re-derived re-derived re-derived re-derived applyAutoDirection — every block caret updateDirectionAt — ancestors only Two or three matches() calls per character, and the cost does not grow with the note.
The scoped pass runs synchronously on input, deliberately outside the 100 ms change debounce — so the caret flips on the first Arabic keystroke rather than a tenth of a second later, and so the HTML autosave writes already carries the attribute.

The bare-text case

A plain-text tab, or a fresh paste of unwrapped text, has no block elements at all — just text nodes directly under the paper. There is nothing to stamp. That case is checked explicitly and the direction goes onto the paper element itself. This is the one situation where the whole-region behaviour is correct, because the region genuinely is one paragraph.

Letting the user overrule the detection

First-strong is right almost always, and wrong in a specific, predictable case: a paragraph that opens with a Latin brand name, a URL or a code identifier but is otherwise Arabic. The detector sees the Latin letter, answers ltr, and is technically correct by the rule while being obviously wrong to the person reading it.

The toolbar's LTR and RTL buttons and the Format menu pin a block by adding a valueless sentinel attribute. syncBlock short-circuits on it in its first line, so a pinned block is never reconsidered:

/** Marks a block whose direction the user set by hand — auto-detection skips it. */
const MANUAL_ATTR = 'data-dir-lock';

export function setSelectionDirection(root: HTMLElement, dir: 'ltr' | 'rtl' | 'auto'): void {
  const blocks = selectedBlocks(root);
  for (const el of blocks) {
    if (dir === 'auto') {
      el.removeAttribute(MANUAL_ATTR);
      el.removeAttribute('dir');
    } else {
      el.setAttribute(MANUAL_ATTR, '');
      el.setAttribute('dir', dir);
    }
  }
  if (dir === 'auto') applyAutoDirection(root);
}

Three states rather than two — left, right, and auto. Without an explicit way back, a user who pins a paragraph once has no way to hand it back to the detector, and the third option costs one branch.

selectedBlocks uses range.intersectsNode() rather than walking from the anchor and focus nodes, because a selection dragged across three paragraphs has boundary points in only the first and last of them. The middle paragraph is fully inside the range and appears in neither endpoint.

The state that leaks between tabs

This one was not on anybody's list. The editor paper is a single DOM element shared by every open tab — switching tabs replaces its contents, it does not create a new element. So an attribute written onto the paper itself belongs to the element, not to the note.

Pin direction on a plain-text tab and the lock goes onto the paper (the bare-text case above). Switch to another tab and the lock is still there, silently suppressing detection on a note that never asked for it. The fix is a single call in setContent, which runs on every tab switch:

export function clearRootDirection(root: HTMLElement): void {
  root.removeAttribute('dir');
  root.removeAttribute(MANUAL_ATTR);
}

Any state written onto a reused element has to be cleared when the content behind it changes. Obvious in hindsight, invisible in the moment.

Paste, and the two attributes worth keeping

Pasted HTML goes through an allowlist before it enters the document — thirty tags, eight style properties, and exactly two attributes:

const ALLOWED_ATTRS = new Set(['dir', 'lang']);

Everything else is stripped, including id, class, and every event handler. dir survives because the source document may have got the direction right and there is no reason to throw that away and re-derive it. lang survives because it drives font selection and hyphenation, and because a language tag is the only thing that distinguishes Persian from Arabic once the text is in the DOM — they share a script and the detector cannot tell them apart, but a font stack can.

After the insert, a full applyAutoDirection pass runs anyway. Pasted markup brings its own blocks, and those blocks have never been through the detector.

Export is a separate problem entirely

Everything above produces a correctly-directioned document on screen. Word does not read any of it.

A .docx has no concept of "infer direction from the text". Both halves have to be stated explicitly: <w:bidi/> in the paragraph properties to flip the paragraph's layout, and <w:rtl/> in each run's properties to set that run's reading order. Emit the first without the second and the text lands in the right place with its characters running the wrong way.

The useful part is that the exporter does not re-implement the rule. It imports the same detectDirection the editor uses and applies it per run, so a mostly-Latin sentence with one Arabic word in it is not flipped wholesale, and the file agrees with the screen by construction rather than by coincidence. That is written up in writing a .docx in the browser with no dependencies.

Direct PDF export sidesteps the problem by declining it. The writer draws with PDF's Standard-14 fonts, which cover Latin-1 only, and complex scripts additionally need shaping — contextual forms in Arabic and Urdu, reordering and conjuncts in Bengali and Devanagari — that no font file alone provides. Anything outside that repertoire is handed to the browser's own print pipeline, which embeds real Unicode fonts, shapes correctly, and honours the dir attributes we spent this whole article stamping. The reasoning is in generating a PDF in the browser.

What we would tell someone starting this

Bidirectional text has a reputation for being a specialist topic, and the shaping and reordering parts genuinely are — but the browser already does all of that for free. What is left is a narrower and much more tractable problem: deciding, per paragraph, which way it runs, and then writing CSS that does not contradict the answer.

The whole module is 197 lines. Most of that is the range table and comments explaining why the table has holes in it.

← All engineering write-ups Try the editor