Notepad Neo

Generating a PDF in the browser by letting the browser do the layout

A PDF page is a list of glyphs at coordinates. Nothing in the format wraps text, indents a list, or breaks a page — so either you write a layout engine, or you borrow the one already running.

DH

— builds and maintains Notepad Neo

· updated · 13 min read

The DOCX exporter has an easy job it does not deserve credit for: it hands Word a stream of paragraphs and runs, and Word decides where the lines break, where the pages end, and how far a nested list indents. The file describes structure. Something else does layout.

PDF does not work like that. A PDF page is a content stream of drawing operators, and text is positioned with an explicit matrix. There is no wrapping, no indentation, no pagination, no concept of a paragraph. Every glyph on the page is somewhere because a number put it there.

Which means a from-scratch PDF exporter needs a layout engine. Writing one that handles line breaking, bidirectional text, list markers, table cells and inline images correctly is not a feature — it is a browser.

The browser is already a layout engine

The way out is that the layout has already happened. The note is on screen, laid out, wrapped and measured, at the exact width the paper will be. Every box the exporter needs to know about already exists in the render tree; the DOM will hand them over through getClientRects() for the asking.

So the pipeline never computes a line break. It reproduces one:

The five stages of the PDF export pipeline The live editor paper is cloned into an offscreen sandbox at exact content width. The clone is measured, producing word boxes, baselines and decoration fills in paper-space coordinates. Those are paginated into pages, then written out as PDF objects and byte offsets. Nothing in the pipeline computes a line break. live paper #editor-paper sandbox clone left: -10000px measure boxes + baselines paginate split by band write bytes objects + xref The layout engine is stage two, and it is Blink or WebKit or Gecko. Line breaking, justification, list indentation, table cell sizing and bidirectional reordering are all read back rather than reimplemented — which is why the output matches what Print produces.
Stage one is a clone rather than the live element because measuring the live paper would mean mutating what the user is looking at, and because the paper carries zoom, theme colours and padding the PDF must not inherit.

The clone cannot be hidden with display:none

The instinct for an offscreen measuring element is display: none. It is exactly wrong. An element with display: none generates no boxes at all — it is not in the render tree — so every getClientRects() call comes back empty and every measurement is zero.

The clone has to be laid out for real, just somewhere nobody can see:

Why an offscreen clone is positioned rather than hidden An element with display none generates no boxes, so getClientRects returns an empty list and all measurements are zero. An element positioned fixed at minus ten thousand pixels is fully laid out and returns real rectangles, while remaining invisible to the user. display: none position: fixed; left: -10000px render tree: no boxes generated el.getClientRects() → DOMRectList (length 0) render tree: el.getClientRects() → real rects, real baselines The clone is given zero padding and exactly the paper's content width, so its border box is its content box and its top-left corner is the origin of paper space. No padding arithmetic later.
Setting the clone's width to the content width rather than the paper width removes an entire class of off-by-a-margin bugs from every downstream coordinate.

Three things the clone has to be protected from

const host = document.createElement('div');
host.className = HOST_CLASS;
host.setAttribute('aria-hidden', 'true');
// Offscreen rather than display:none — a display:none subtree generates no
// boxes at all and every getClientRects() call would come back empty.
host.style.cssText =
  'position:fixed;left:-10000px;top:0;opacity:0;pointer-events:none;' +
  'z-index:-1;margin:0;padding:0;border:0;background:#fff;' +
  `width:${contentWidthPx}px;`;

const clone = plainText === undefined ? prepareClone(paper) : plainPaper(plainText);
for (const [prop, value] of FORCED) clone.style.setProperty(prop, value, 'important');
clone.style.setProperty('width', `${contentWidthPx}px`, 'important');

The theme. Dark mode is applied via [data-theme="dark"] on the <html> element, which is an ancestor of the sandbox. The clone inherits it, and a PDF exported at night would come out with white text on a transparent background. The fix is explicit overrides at matching specificity with !important — you cannot escape an ancestor selector by being more specific further down.

Duplicate identity. The clone is a deep copy of an element with an id, so for as long as it exists there are two id="editor-paper" nodes in the document. getElementById returns the first, which is usually but not reliably the original. The id, contenteditable, role and aria-* attributes are all stripped from the clone.

Transient UI. Find & Replace highlights are <mark> elements in the live DOM. They are unwrapped from the clone, because a search highlight is not part of the document and exporting a PDF with yellow bars behind every search hit is a bug report waiting to happen.

Never interleave DOM writes and geometry reads

This is the strictest rule in the whole codebase, and it is enforced by a comment at the top of the measuring module and two labelled phase boundaries inside it.

Reading a geometric property — getClientRects, offsetTop, getBoundingClientRect — forces the browser to flush any pending layout before it can answer. Writing to the DOM invalidates layout. Alternate the two and every read triggers a full reflow of the clone. The work is quadratic in a way that turns a sub-second export of a long note into a multi-second freeze of the main thread.

Interleaved writes and reads versus two separated phases Interleaving a DOM write with a geometry read forces a layout flush before every read, so a document with many elements pays a full reflow per element. Doing every write first and every read afterwards costs a single layout flush for the whole document. INTERLEAVED — ONE REFLOW PER READ write reflow read write reflow read write reflow read · · · × N elements PHASE-SEPARATED — ONE REFLOW TOTAL Phase 1 — every style read and DOM mutation reflow Phase 2 — every geometry read, no mutations The ordering is a correctness constraint too, not only a performance one. List markers are replaced with real spans in phase 1, and every list-style-type has to be read before any is written — because list-style-type inherits, and suppressing it on one item changes what a nested item reports.
The two phases are marked with banner comments in the source. It is the kind of invariant that is trivially easy to break with a one-line change six months later, and impossible to notice until somebody exports a long document.

Why list markers have to be rebuilt

A ::marker pseudo-element is not in the DOM. There is no node to select, no Range that can contain it, and no getClientRects() to call on it. It is generated content the layout engine draws, and the only way to find out where it landed is to replace it with something real.

So in phase 1 every list marker becomes an absolutely-positioned span containing the bullet or number, and the original list-style-type is set to none. In phase 2 that span is measured like any other text. The alternative — computing marker positions from the list's padding and the item's font metrics — would have to reimplement counter formatting, nested numbering and RTL marker placement, which is the reimplementation this whole approach exists to avoid.

Measuring word by word, and probing for the baseline

Text is measured one token at a time. A Range is placed around each non-whitespace run of characters and getClientRects() gives its box. Measuring whole lines would be cheaper, but a line is not addressable — the DOM has no handle for "the second visual line of this paragraph", because line boxes are a layout artefact with no node behind them.

Occasionally a token returns more than one rect, which means overflow-wrap: break-word split it mid-word across two lines. That case falls back to per-character rects. It is the only place in the walk where that cost is paid, and it is paid on the rare token rather than on all of them.

The harder measurement is vertical. A rect gives you a top and a height; PDF needs a baseline, and there is no API that reports one.

Finding the text baseline with a zero-width inline-block probe A line box has an ascent above the baseline and a descent below it, and the ratio between them is decided by the font's metrics, not by CSS. A zero-sized inline-block element with vertical-align baseline sits with its bottom edge exactly on the baseline, so measuring its rect reveals where the baseline is. Chrome and Firefox disagree about that rect's height, so it has to be measured rather than assumed. LINE BOX top of line box baseline bottom of line box Hxpg probe: 0×0 inline-block, vertical-align: baseline ascent descent The probe's bottom edge sits on the baseline by definition of vertical-align: baseline, so its rect is the answer. One probe is inserted per distinct font context, keyed on fontFamily | fontSize | lineHeight | fontWeight | fontStyle — so a document in one font pays for one probe, not one per word.
Deriving the baseline arithmetically from the line box would need the font's ascent and descent metrics, which the DOM does not expose. The probe rect's own height is engine-defined and Chrome and Firefox disagree about it, which is exactly why it is measured rather than assumed.

Underlines are drawn per line, not per word

function buildProbe(cs: CSSStyleDeclaration): HTMLElement {
  const wrap = document.createElement('div');
  wrap.style.cssText =
    'position:absolute;left:0;top:0;white-space:nowrap;visibility:hidden;';
  wrap.style.fontFamily = cs.fontFamily;
  wrap.style.fontSize   = cs.fontSize;
  wrap.style.lineHeight = cs.lineHeight;
  wrap.style.fontWeight = cs.fontWeight;
  wrap.style.fontStyle  = cs.fontStyle;

  const text = document.createElement('span');
  text.textContent = 'Hxg';
  // A zero-sized baseline-aligned inline-block sits exactly on the baseline.
  const base = document.createElement('span');
  base.style.cssText = 'display:inline-block;width:0;height:0;vertical-align:baseline;';

  wrap.appendChild(text);
  wrap.appendChild(base);
  return wrap;
}

Reading it back is the difference of two rectangles — the top of the text's own client rect, and the top of the probe, which is sitting on the baseline:

function readProbe(probe: HTMLElement): number {
  const textSpan = probe.firstElementChild as HTMLElement;
  const baseSpan = probe.lastElementChild as HTMLElement;
  const range = document.createRange();
  range.selectNodeContents(textSpan);
  const textRect = range.getClientRects()[0];
  const baseRect = baseSpan.getBoundingClientRect();
  if (!textRect) return 0;
  return baseRect.top - textRect.top;
}

Text decorations are rectangles in PDF, not text properties. The naive translation draws one under each measured word — and produces a visibly dashed underline, because the spaces between words are whitespace tokens that were never measured.

So decoration fills are grouped by visual line and merged into one rectangle per run. The grouping key is the baseline y-coordinate, rounded to the nearest half pixel, which is enough to tolerate sub-pixel differences between adjacent words in the same line without accidentally merging two lines that happen to be close together.

The encoding wall

PDF ships fourteen standard fonts that every reader is required to have — four weights each of Helvetica, Times and Courier, plus Symbol and ZapfDingbats. Using them means no font file has to be embedded, which keeps the exporter small and the output portable.

It also means /WinAnsiEncoding: every glyph has to map to a single byte in CP1252. That is roughly Latin-1 plus a block of typographic characters. Anything outside it — Arabic, Bengali, Devanagari, Greek, Cyrillic, CJK, emoji — cannot be drawn by this writer at all.

What the Standard-14 fonts can and cannot draw Bytes 0x20 to 0x7E are ASCII and map directly. Bytes 0x80 to 0x9F are where Latin-1 and CP1252 diverge, and are filled from an explicit table of typographic characters. Bytes 0xA0 to 0xFF map directly again. Everything outside that repertoire has no representation. A small set of characters is substituted silently instead: non-breaking and thin spaces become an ordinary space, and zero-width characters are dropped. THE WRITER'S REPERTOIRE ASCII 0x20 – 0x7E · direct CP1252 table 0x80 – 0x9F · lookup Latin-1 upper 0xA0 – 0xFF · direct everything else UNMAPPED € ‚ ƒ „ … † ‡ ˆ ‰ Š ‹ Œ Ž ‘ ’ “ ” • – — ˜ ™ š › œ ž Ÿ the one place Latin-1 and CP1252 disagree مرحبا · নমস্কার · CJK · Ελληνικά SUBSTITUTED, NOT REPORTED U+00A0 · U+2007 · U+2009 · U+202F → 0x20 U+200B · U+200C · U+200D · U+FEFF → dropped contenteditable produces these constantly. Reporting them would fire on every note.
The 0x80–0x9F range is the classic trap. Latin-1 defines those as control characters; CP1252 fills them with curly quotes, dashes and the euro sign — which are exactly the characters a rich-text editor produces most often.

The substitution set matters more than it looks. A contenteditable emits non-breaking spaces constantly, and the font-size tool injects U+200B to keep empty spans alive. If those counted as unrepresentable, virtually every note in the app would be diverted away from the direct writer, and the feature would effectively not exist.

One function decides, so the two callers cannot disagree

There are two questions to answer about any character: can we draw it? and what byte is it? Answering them in two places is how you end up with a pre-flight check that passes and an encoder that then substitutes a question mark.

So both go through one function, with two sentinel return values:

/** Sentinel: the character is intentionally dropped. */
const DROP = -1;
/** Sentinel: the character has no WinAnsi representation. */
const UNMAPPED = -2;

function winAnsiByte(code: number): number {
  if (code in SUBSTITUTE) {
    const sub = SUBSTITUTE[code];
    return sub === null ? DROP : sub;
  }
  // Control characters have no glyph and would corrupt the literal.
  if (code < 0x20) return code === 0x09 ? 0x20 : DROP;
  if (code < 0x80 || (code >= 0xa0 && code <= 0xff)) return code;
  const mapped = CP1252_HIGH[code];
  return mapped !== undefined ? mapped : UNMAPPED;
}

The pre-flight check walks the note's plain text, collects up to eight distinct unmappable characters, and if it finds any, the export is handed to the browser's print pipeline instead — with a toast naming the characters, so the outcome is explicable rather than mysterious:

const unsupported = findUnencodable(editor.getPlainText());
if (unsupported.length) {
  showToast(
    'This note uses characters (' + unsupported.slice(0, 4).join(' ') +
    ') that direct PDF export cannot render. Opening Print — choose ' +
    '"Save as PDF" there to keep them.',
    8000,
  );
  printAsPdf(safeFilename(tab.title));
  return;
}

That fallback is not a consolation prize. The browser's print pipeline embeds real Unicode fonts and — critically — shapes them. Arabic and Urdu need contextual forms, where a letter's glyph depends on its neighbours. Bengali and Devanagari need reordering and conjunct formation. No font file alone provides any of that; it needs a shaping engine, and the browser has one. A from-scratch writer that embedded a Unicode font would produce isolated, unjoined letterforms: technically Arabic characters, and unreadable as Arabic.

It also means the direct writer never needs bidirectional handling. Anything that would require it has already been routed away, and the print path honours the dir attributes the editor stamps — described here.

A second layer, on the assumption the first one has a hole

The encoder itself never throws. If it meets an unmapped character it emits ? and adds the character to a set, which the writer returns and the caller surfaces as a second toast. The pre-flight check should have caught it, and if it did not, the user gets a readable file and an explanation rather than a stack trace.

Both layers call winAnsiByte, so a gap can only ever be a gap in that one function.

There is one deliberate exception to the Latin-1 limit. The document title in the PDF's metadata is written as UTF-16BE with a byte-order mark, which PDF strings support, so a note titled in Bengali still reads correctly in a viewer's document-properties panel even when its content would have been routed to Print.

Writing the bytes

A PDF is a set of numbered objects followed by a cross-reference table that lists the exact byte offset of each one. Get an offset wrong by a single byte and the file is unopenable — the reader seeks to that position and does not find an object header.

Which makes the output encoder the most dangerous line in the module:

// Every byte here is either ASCII syntax or an already-WinAnsi-encoded
// string byte. A UTF-8 encoder would silently turn one 0xE9 into two bytes
// and desynchronise every xref offset after it.
const push = (s: string) => {
  for (let i = 0; i < s.length; i++) bytes.push(s.charCodeAt(i) & 0xff);
};

TextEncoder is the obvious tool and the wrong one. It produces UTF-8, so é — already encoded as the single byte 0xE9 by the WinAnsi step — becomes two bytes. The content stream is still readable; every byte offset recorded after that point is not. The failure appears only in documents containing accented characters, which is a test case it is easy not to have.

The coordinate flip

DOM coordinates versus PDF coordinates DOM geometry has its origin at the top left with y increasing downward. PDF has its origin at the bottom left with y increasing upward. Every measured y coordinate is converted through a single helper that subtracts it from the page height and offsets it by the page's start position and the top margin. DOM · origin top-left PDF · origin bottom-left 0,0 y grows down First line of the note y = 36 0,0 y grows up First line of the note y = 104 const py = (y: number) => (PAGE_H_PX - margins.top - (y - page.startPx)) * PT; One helper. Every vertical coordinate in the file goes through it, and nothing flips twice.
PT is 0.75 — CSS pixels to PDF points at 96 dpi. page.startPx is where the current page begins in the measured document, which is what turns one continuous stream of boxes into separate pages.

Every word gets its own text matrix

The efficient way to draw a line of text in PDF is one Tj operator for the whole line and let the font's advance widths space the glyphs. That is what a PDF produced by a layout engine does.

It is wrong here, because the font that laid the text out on screen and the Standard-14 font drawing it in the PDF are not the same font. Arial is metric-compatible with Helvetica by design, and Times New Roman with Times — but any other face is an approximation, and small per-character differences in advance width accumulate along a line until the last word is visibly out of place.

So each word is positioned with its own Tm, at the coordinate the browser actually measured. The difference between the screen font and the substitute stops being cumulative and becomes a sub-pixel variation in the gap between words.

Monospace gets one extra correction. Courier's advance is exactly 0.6 em by definition, and the monospace face the browser used almost certainly is not, so a horizontal scale operator (Tz) stretches or compresses the glyphs to match the measured width. That keeps code blocks aligned in a way per-word positioning alone would not.

Pagination has to respect what is atomic

while (pageTop < total && bounds.length < MAX_PAGES) {
  let candidate = pageTop + height;
  if (candidate < total) {
    for (const span of atomic) {
      if (span.top >= candidate) break;   // sorted — nothing later can straddle
      if (span.top > pageTop && span.bottom > candidate) {
        candidate = span.top;             // pull the break up to this line's top
        break;
      }
    }
  }
  // A single item taller than a whole page would otherwise stall the loop.
  if (candidate <= pageTop) candidate = pageTop + height;
  bounds.push([pageTop, candidate]);
  pageTop = candidate;
}

// A PDF with zero pages is invalid, so an empty note still gets one.
if (bounds.length === 0) bounds.push([0, height]);

Splitting a measured document into pages is mostly arithmetic, with two exceptions.

A visual line is atomic — a line that straddles a page boundary pushes the break up to that line's top rather than being sliced through the middle of its glyphs. But a tall background fill, like a code block's shading or a table cell, is not atomic: it is meant to continue across the break, so it is clipped per page band instead. Treating both the same way gives you either sliced text or backgrounds that stop abruptly at the bottom of a page.

There are two guards worth having. A PDF with zero pages is invalid, so an empty note still produces one blank page. And there is a hard cap of 500 pages, because the only way to reach it is a bug in the measurement — a document that reports an unbounded height will otherwise allocate until the tab dies.

What this approach is good and bad at

Good at: matching what the user sees, exactly, because it is the same layout. Staying small — the whole PDF pipeline is about 1,300 lines across seven files, with no dependencies and no font files. Producing real selectable, searchable text rather than an image. Working entirely offline, which for an app whose whole premise is that notes never leave the device is the point rather than a bonus.

Bad at: anything outside CP1252, which is a large fraction of the world's writing systems and is handled by handing the job back to the browser. Images, which are not implemented. Very long documents, where measuring thousands of tokens on the main thread is felt — the work is linear but it is synchronous. And anything requiring genuine typography: kerning pairs, ligatures, and hyphenation are whatever the browser decided and the Standard-14 substitute approximates.

The honest summary is that this is not a PDF library. It is a way to get a good PDF out of one specific document, for the common case, without shipping a PDF library — and a pre-flight check that recognises when it is out of its depth and says so instead of producing something wrong.

← All engineering write-ups Try the editor