[v3.0.2] List rendering robustness fixes (itemize / enumerate) - #422
Draft
OlgaRedozubova wants to merge 20 commits into
Draft
[v3.0.2] List rendering robustness fixes (itemize / enumerate)#422OlgaRedozubova wants to merge 20 commits into
OlgaRedozubova wants to merge 20 commits into
Conversation
These .d.ts files were emitted by an older tsc and only differ in tuple formatting (e.g. [number, number][]); regenerate them with the pinned 4.9.5 compiler so committed declarations match the toolchain.
A top-level itemize/enumerate only got its inline-start padding when the
widest \item[...] marker was measured. That measurement lived only in the
inline item path, so items whose content is a block environment
(\begin{figure}, \begin{tabular}, code fence) were skipped. When every
long-marker item held block content, the list lost its padding entirely.
Extract the marker-width calc into a shared computeMarkerPadding helper and
apply it in the block item path too (setTokenListItemOpenBlock now returns
the created token). Add regression tests for fenced-code and figure items.
Marker padding summed String.length, so a fullwidth/CJK marker like \item[11.] (U+FF0E, length 3) was undercounted versus its ASCII twin \item[11.42] (length 5) and fell under the padding threshold, leaving the list without indentation while a neighbouring ASCII list got it. Count East-Asian Wide/Fullwidth characters as width 2 (iterating by code point so surrogate pairs count once). ASCII markers are unaffected.
The Lists block rule parses speculatively into a buffered state, but that state shares env by prototype, so parsing mutates the real env.isBlock (and inheritedListType). On an unclosed list the rule returned false without restoring them; a leaked isBlock=true then let the inline list fallback fire on the following text, so an unclosed itemize before a tabular rendered a broken partial list with empty <> item bodies instead of plain text. Restore isBlock and inheritedListType when the speculative parse aborts, so the unclosed list degrades to text exactly as it does without a tabular.
The \footnote/\footnotetext block rules scan forward for their open tag; their
terminator set included the core markdown list rule but not the LaTeX list rule,
so a \begin{itemize} between a paragraph and a later footnote (no blank line) was
swallowed and rendered as literal text. Add the LaTeX list rule to the footnote
terminator set and use it for both footnote block rules.
Bumps version to 3.0.2; adds the list-rendering-robustness spec and changelog
entry covering this branch's list fixes.
Code-text styles were pinned to absolute values (`pre code` font-size 15px / line-height 24px / padding 1rem, `pre` font-size 85%), so a code block did not scale when a consumer sizes a rendered block via a single font-size on the container (e.g. image export): everything else scaled but the code stayed ~15px. Make them relative: `pre` font-size 0.9375em; `pre code` font-size inherit, line-height 1.6, padding 1em. Calibrated so a 16px base is pixel-identical except code padding (16px -> 15px). Styles only; no HTML change. Regenerate style snapshots, add a regression gate that rejects absolute font-size/line-height and rem padding for `pre code`; spec + changelog under 3.0.2.
…inators, perf guard - Restore env.isBlock/inheritedListType on both the abort and silent paths of the Lists block rule, so a silent terminator probe never mutates shared state. - \footnote uses a minimal fence+Lists terminator set (not the full set): fixes the list swallow with one cheap extra probe, avoiding a per-line cost regression. - Add a first-char fast bail to the Lists rule (no substring allocation) since it now runs as a per-line terminator in paragraph/footnote scans. - Measure markers by display width (East-Asian wide chars count as 2, BMP only). - Tests: silent env invariant, marker width edge cases (math/emoji/nested), footnote recognition after block constructs, and a scaling guard that rejects the O(N^2) footnote scan across unterminated lists. - Selector-scoped code-style gate; spec/changelog updated.
- Measure math markers by their rendered widthEx and wrapped markers (e.g. \textbf) by recursing into children, so such markers no longer get too small an indent (previously only top-level text tokens counted). - Extract the char-based width primitives (isWideChar, displayWidth, tokenDisplayWidth) into common/display-width.ts — the char-based counterpart of getTextWidthByTokens; computeMarkerPadding now delegates to them. - Round the padding px (marker width can be fractional once math width is included). - Tests: wide-math and bold marker padding, markdown list not swallowed before a \footnote; explicit non-empty guard on the code-style assertions. - Spec and changelog updated.
Empty `<>` item bodies from a leaked env.isBlock are fixed on both paths: the speculative list parse restores its transient env fields on every exit (abort/silent/commit/exception), and a tabular inside a list no longer carries those flags into its envToInline snapshot (which core-inline replays onto the shared env). Adds common/env-transient.ts as the single source of truth. Marker padding is emitted in ex, not px, so it scales with the container font-size like the marker's math SVG; this also fixes math markers being clipped (exact widthEx + gap, both ex). The marker gap (.li_level padding-right) and the default list indent move to ex/em. Marker width also trims edge whitespace and falls back to source length when widthEx is absent. Merges the two footnote terminator resolvers into resolveEnabledRuleFns. Adds tests/_display-width.js, an env-leak adversarial matrix, exact ex-value assertions, and updates fixtures/snapshot/specs/changelog.
Emit list marker padding-inline-start in em (was ex), converted from the ex
measurement via the default font metrics (EX_TO_EM = exDef/fonSizeDef). Custom
padding is emitted only when it exceeds the default (2.5em), so it never resolves
below it — no CSS max() or cross-unit comparison. The marker gap (0.625em) and
default indent (2.5em) move to em too; the attribute now carries its unit
("3.32em"). Math without a widthEx (non-SVG output) keeps the default indent
rather than a fabricated source-length estimate.
Text markers use 1.3 ex/char (reserve is per character cell, not per glyph);
this is ~5-14% tighter than 3.0.1 and wide all-caps markers can overlap —
documented in the spec Non-Goals and changelog.
Sync docs/comments with the code, add tests: EX_TO_EM vs fontMetrics, exact em
values, glyph-width limitation, MARKER_GAP_EM vs CSS.
A speculative (silent/aborted) list parse runs the body's \begin{figure} /
\begin{table}\caption and bumps the module-global caption counters, but its
tokens are discarded — so a \footnote after such a list, and the pre-existing
paragraph-terminator probe, shifted Figure N / Table N. Snapshot and restore the
counters when the parse is discarded: in the Lists block rule's finally and in
parseListEnvRawToTokens (inline path). A figure in a list is now Figure 1.
Extract the counters to the leaf module common/caption-counters.ts so the list
rule no longer imports the heavy begin-table module (removes an import-cycle
edge); the deep-import helpers Clear{Table,Figure}Numbers move there, renamed to
clear{Table,Figure}Numbers. Move the default font metrics (DEFAULT_FONT_SIZE_PX
/ DEFAULT_EX_PX) to consts as the single source for EX_TO_EM and FontMetrics.
Treat combining U+3099/U+309A as zero width in both isWideChar and displayWidth;
set committed before flush; drop the now-redundant CSS guard.
Bump 3.0.1 -> 3.1.0 (minor): the data-padding-inline-start attribute is now an
em value and caption numbering changes for existing docs (see changelog/spec).
Add regression tests: caption numbering (absolute, multi-float order, nested
list, \ref matches the caption number), attribute format contract, 3-char
threshold, combining marks in isWideChar and displayWidth.
Replace the flat per-character reserve with a per-glyph-class estimate in em (narrow / normal / wide / extra-wide W@% / East-Asian full-width; combining marks 0). Measured against real font metrics, the old flat 14px/char over-reserved narrow/digit markers by 36-70% and under-reserved all-caps by ~14%; the class estimate holds a +11..+27% margin — a tighter, correct indent for narrow/normal markers and a wider, safer one for all-caps. Math still uses the exact widthEx; short markers fall to the 2.5em default; the reservation is clamped at 20em so a pathological OCR marker can't blow out the content column. Only text-like leaves (text / code_inline / text_special) are measured, so a `code`-span marker now contributes its width while an html_inline marker's raw markup is not counted. Remove the now-unused displayWidth; dedupe the combining -mark predicate. Add a font-metrics lock test (Arial fixture): the rendered indent is never smaller than the marker's true glyph width + gap. Update fixtures, docs, and the attribute-format / clamp / mixed-marker tests.
A list's marker padding was accumulated as the max over all markers at every depth and written to the outermost list token, so a wide marker several levels deep drove the top-level list's indent (e.g. a deep [3.1.1.1] gave the outer list 4.31em while its own 1.–5. markers are narrow). Track a stack of open list tokens and attribute padding to the current (innermost) list instead. Nested lists now also emit and apply their own padding (drop the top-level-only gate in finalizeListItems and the level>1 suppression in the list renderers), so each level reserves for its own markers — a wide nested marker indents its own list rather than overflowing the container. The outer list reflects only its own markers (default when they're narrow). Add regression tests (per-level attribution; deep marker not bubbled up) and update changelog/spec.
Marker padding (on top of the per-nesting-level attribution): keep the default 2.5em indent on every list and emit a custom padding-inline-start only when a marker overflows the accumulated ancestor indent plus that default, reserving just the shortfall (need - ancestorIndent) so nested markers don't compound. Flat lists are unchanged (e.g. [11.33] -> 3.51em); ordinary nested lists (numbering, bullets) stay at the default; only a genuinely wide nested marker reserves extra, and less than its full width. Fix \item detection: the item-command regexes matched too loosely - LATEX_ITEM_COMMAND_INLINE_RE matched the bare word "item" (no backslash), and LATEX_ITEM_COMMAND_RE / LATEX_LIST_BOUNDARY_INLINE_RE matched any \item-prefixed command (\itemsep, \itemindent). A list-item body containing such text split into a spurious item, and a multiline \footnote/\footnotetext inside an item was broken mid-brace and rendered as literal text. All three now require \item plus a command boundary (\item[...] or \item(?![a-zA-Z])). Update changelog and the list-rendering pr-spec; add regression tests (_lists.js: \itemsep/\itemindent; _footnotes_latex.js: "item" in a footnote body; _list-marker-padding.js: per-level B2 attribution).
…ate padding style Resolve per-list padding in one top-down pass (resolveListPadding) after the whole environment is parsed, instead of emitting during finalize. finalizeListItems now only records each list's widest marker width; the pass computes the indent once every list's final width is known, so item order no longer skews a nested list (a wide parent item after the sublist resolves the same as before it). Clamp the total (ancestor + own), not just the added shortfall, to LIST_MAX_INDENT_EM, so cumulative indent never exceeds the clamp on a pathological nested marker. Validate data-padding-inline-start against /^\d+(\.\d+)?em$/ before inlining it as a style, so only a bare em value can reach the rendered CSS. Docs: the attribute is now emitted on nested lists (Migration); an empty parent item followed by a wide sublist marker overlaps as LaTeX does (Non-Goals, verified in Overleaf); image markers reserve by alt text; correct the code-block "16px" note for PreviewStyle's 17px base. Add regression tests (order independence, cumulative clamp, bold reserve lock).
Core-markdown <ul>/<ol> (no .itemize/.enumerate class) had no padding-inline-start, so browsers used their 40px default, which doesn't scale when a consumer sizes content by font-size. Add a scoped rule (#preview-content ul, #setText ul, ... ol) setting padding-inline-start: 2.5em, matching the LaTeX-list default and scaling with the container. Scoped only per 2026-03-mmd-css-scoping.md — a bare ul/ol would restyle the host page. LaTeX lists are unaffected (their class rules and inline padding win by specificity, same value); TOC is safe (separate #toc_container; the in-content .table-of-contents is display:none). Regenerate the listsStyles snapshot; update changelog and the list-rendering pr-spec.
…oped-only Follow-up to the markdown-list padding rule (bc57340): - The rule also covers the generated footnotes list (<ol class="footnotes-list">), which had no explicit padding — disclosed in changelog and the pr-spec/Migration. - Migration: the wrapped ul/ol padding floor rises from the UA default (specificity 0) to (1,0,1), so a consumer's class-specificity rule no longer overrides the indent. - Spec: the #toc_container-over-markdown precedence rests on bundle order (lists before toc), not specificity (both 1,0,1) — documented, with the specificity-boost option. - Test: assert listsStyles has no bare ul/ol selector, so a future edit can't quietly regenerate the snapshot with one.
The markdown-list rule (#setText ul { padding-inline-start: 2.5em }) and
#toc_container ul { padding: 0 } are both (1,0,1), so which wins depended on bundle
order (lists before toc). TocStyle is only emitted when toc is enabled, so a consumer
that renders a visible #toc_container inside the wrapper without TocStyle got 2.5em on
the TOC list.
Add a guard in the always-emitted lists module — #preview-content #toc_container ul/ol,
#setText #toc_container ul/ol { padding-inline-start: 0 }, specificity (2,0,1) — so the
TOC list stays at 0 regardless of order or whether TocStyle is present. It matches any
ul/ol inside a wrapper-nested #toc_container (so it also outranks .itemize (1,1,0); a
theoretical LaTeX-in-TOC would be zeroed too — accepted over a :not(.itemize) carve-out).
Also address review follow-ups: correct the clamp wording (the reserved indent is
clamped to 20em, the per-level default step adds on top); widen the no-bare-ul/ol
tripwire to catch post-comma and combinator forms; regenerate the listsStyles snapshot;
update the pr-spec.
… test fixtures Attribute marker padding through one shared registry (openTokens + allListTokens) seeded in ListOpen and threaded into ListItems/processListChildToken, so all list forms resolve identically — the block loop (own-line nested), the same-line form, and the fully single-line / in-cell form (which the block Lists rule bails on and ListOpen now handles, resolving on its single-line early-return). A wide nested marker no longer inflates the outer list regardless of how the list is written, and a single-line / in-cell list reserves for its marker too. Review follow-ups: remove the dead `padding` plumbing (ListItemsResult / ListInlineContext field, fallback branches, ListOpen local); make openTokens / allListTokens required (no silent-no-op default); fill skipped depth before the ancestor sum and clamp depth >= 0 in resolveListPadding (no RangeError / NaN); fix the stale "for top-level lists" render comments; note the inline-opened *_list_open prentLevel (0 -> depth) in Migration. Convert the rendered-HTML tests to full-HTML fixtures (catch visual regression the way the _data-driven tests do): marker padding, B2 nesting, empty-<>, \item detection and footnote-terminator cases -> tests/_data/_lists/_data.js; caption numbering -> _data/_captions/_data.js; "item"-word footnote body -> _data/_footnotes_latex/_data-footnotetext.js. Keep the non-renderable tests (display-width units, config-varying renders, env-state, the Arial glyph-width invariant, resolveListPadding unit, the scaling test) as targeted assertions.
Changelog: cut spec-level internals (per-glyph constants, px measurements, the 3.0.1 over-reservation percentages, internal probe/env mechanism) from the 3.1.0 list-fix bullets, keeping what changed plus the migration — matching the concise style of earlier entries. pr-spec: rewrite the Testing section, which described tests that were moved into full-HTML fixtures (the list-rendering, caption and footnote-body cases), to match the actual layout; drop an intermediate example value from a Non-Goal.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Two groups of renderer fixes, shipped together as one
3.0.2release:Lists (1–4): four unrelated inputs made LaTeX
itemize/enumeratelists render incorrectly; the fixes live in the list-env and footnote block rules.Code blocks (5): code-text styles were absolute (
px/rem), so a code block did not scale when a consumer sizes a rendered block by a singlefont-size; the styles are now relative.Version bump: 3.0.1 → 3.0.2
Specs:
pr-specs/2026-07-list-rendering-robustness.md,pr-specs/2026-07-code-block-font-scaling.mdChangelog:
doc/changelog.mdFull test suite green.
1. Marker padding for block-content items
A top-level list gets its
padding-inline-startfrom the widest custom\item[...]marker, but the width was only measured on the inline item path. Items whose content is a block environment (\begin{figure},\begin{tabular}, a code fence) were skipped, so a list whose long-marker items all hold block content lost its padding.MMD example
2. Marker width: fullwidth/CJK, math, and wrapped content
Marker width summed
String.lengthover top-leveltexttokens only, so several markers were undercounted and the list got too small an indent (marker overlaps the content):\item[11.], U+FF0E) counted as narrow ASCII — now East-Asian Wide/Fullwidth chars count as 2.\item[$x^4+x^4$]) contributed 0 — now uses the token's renderedwidthEx.\item[\textbf{…}]) contributed 0 — now measured through the wrapper's children.ASCII text markers are unchanged.
MMD examples
11.$x^4+x^4$\textbf{…}3. No env-state leak from an aborted list parse
The list block rule parses speculatively into a buffered state that shares
envby prototype. On abort (unclosed list) or a silent probe it returned without restoringenv.isBlock(andenv.inheritedListType); the leakedisBlock = truethen let the inline list fallback fire on the following content, so an uncloseditemizebefore atabularrendered a broken partial list with empty<>item bodies. Those transient fields are now restored on both paths, so the unclosed list degrades to plain text — exactly as it does without atabular.MMD example
4. Footnote block rules stop at a list start
The
\footnote/\footnotetextblock rules scan forward for their open tag, terminating at block boundaries so they don't swallow following blocks. The LaTeX list rule was not a terminator, so a\begin{itemize}between a paragraph and a later footnote (no blank line) did not stop the scan: the list was swallowed and rendered as literal text (a blank line masked the bug). The LaTeX list rule is now a terminator for both —\footnotetextvia its full set,\footnotevia a minimalfence+Lists(keeping its cheap scan).This is also a performance fix: on repeated paragraph + list-with-footnotetext units without blank separators the missing terminator made the scan run across every list into the rest of the document — O(N²), seconds on large inputs; terminating at the list makes it linear (guarded by a scaling test).
MMD example
5. Code-block styles scale with the em context
Code-text styles were pinned to absolute values, so when a consumer scales a rendered block by setting a single
font-sizeon the container (e.g. image export), everything scaled except the code —pxis fixed andremresolves against the root, not the block'sem. The four properties are now relative, calibrated so a 16px base is pixel-identical to before (only code padding moves 16px → 15px):#setText pre { font-size: 0.9375em; }(was85%)#setText pre code { font-size: inherit; }(was15px)#setText pre code { line-height: 1.6; }(was24px)#setText pre code { padding: 1em; }(was1rem)Styles only — no change to
lstlisting/ fenced-code markup.MMD example
Scale the rendered block by setting a large
font-sizeon the container (as image export does).Testing
List cases in
tests/_data/_lists/_data.jsandtests/_list-marker-padding.js: block-content markers (figure/fence), fullwidth11., math and\textbfmarkers, an unclosed list +tabular, a list after a paragraph with a multiline\footnote{}/\footnotetext{}, and a markdown list not swallowed before a\footnote.Guards: a scaling test (
tests/_footnotes_latex.js) rejecting the O(N²) footnote scan; a silent-Listsenv invariant (tests/_parse-isolation.js); selector-scoped code-style assertions (tests/_styles.js).Full suite green.
Non-goals
> 3threshold) is unchanged.code,prescroll/overflow, highlight colors, and table-cell padding are untouched.