Skip to content

fix(docx): write a list item the page reads as markdown as the page sets it - #869

Merged
DemchaAV merged 3 commits into
2.5-devfrom
fix/docx-write-list-markdown
Oct 7, 2026
Merged

DemchaAV merged 3 commits into
2.5-devfrom
fix/docx-write-list-markdown

Conversation

@DemchaAV

@DemchaAV DemchaAV commented Oct 7, 2026 •

Copy link
Copy Markdown
Owner

Why

A session reads markdown unless it is told not to (markdown(false)). The page reads a list item of plain text that holds *, _ or a backtick the way it reads a paragraph: the text its marks style is set bold or italic, and the marks are dropped. The DOCX export writes a paragraph that way since #868. It still wrote a list item as authored, so Word showed **Java** lead with its asterisks and none of the bold, and the report named it on the list's note.

What changed

  • DocxMarkdown.ItemReading is the text the page reads for an item. The page builds one paragraph per item, and DocxMarkdown.flatItem, nestedItem and items mirror what goes into it:

    • Flat list: the item's text with a typed marker taken off. Without hangingIndent, the page sets the list's marker before the first line as a prefix, apart from the text. The marker leads the line's letters, as a paragraph's bulletOffset does in fix(docx): write a paragraph the page reads as markdown as the page sets it #868.
    • Marker in a column of its own (hangingIndent): the text alone. The marker is a fragment of its own, and DocxLayoutMetrics.itemLines already takes the text's.
    • List built as a tree of items, no hangingIndent: at any level, the indent and marker, then the label. The page flattens the tree into labels and reads the indent and marker with the label. The indent is two no-break spaces a level (TextFlowSupport). The parser reads those as letters, so a deep item keeps its text, where four plain spaces would make a code block.
  • itemPieces returns an item's pieces where the lines the layout laid that item out in hold them.

    • It reads the item with DocxMarkdown.read and checks the pieces with DocxMarkdown.laidOutIn against the item's own lines. The export already matches items to the layout's one by one for the line gap.
    • The paragraph check moved into pagePieces, which paragraphs and items share: pieces that change nothing, no lines to read, and the comparison with the page's lines. markdownPieces calls it.
  • New DocxMarkdown.split takes a tree item's indent and marker off its pieces. They must be the leading characters, all in one style; otherwise the item is written as authored.

  • The marker stays where it was.

    • Word list: Word draws the marker. The page's parser sets the marker it reads with a tree's item regular, a bold list's too. Word draws such a list's bullets regular whatever the paragraph mark holds: measured in Word 16 on a bold Lato list, the bullets' ink matches the page's regular marker, not its bold one.
    • List written as a paragraph per item: the export writes the marker and indent as characters before the pieces, in the style the page sets them in.
  • writeItemText writes an item's text for every path:

    • writeListLine, which writes a Word list item, or the item text with its marker and indent as characters;
    • writeRichListLine's plain label, written after a drawn marker, a text marker or a tab to the marker column.

    It records each item in listItemsWritten for the note. The paragraph mark keeps the list's style.

  • The note names what a list's matched items lose (itemsMarkdownLost).

    • A heading written taller than its item's line is named. headingCut is extracted from markdownHeadingCut, which still names a paragraph's.
    • An item the parser reads into nothing, ***, is named as one the page sets none of the text of. The check counts letters past the marker prefix, because the page sets a flat list's marker even before an empty line. Before this, such an item was named as having its marks written as letters.
    • An item otherwise written as authored is named where its lines, less the prefix's own marks, hold fewer marks than it. A marker of ** holds marks of its own.
  • A list whose items are not matched keeps its note (itemsMarkdownUnmatched). That covers no layout, a list composed in a table cell, an item run onto the next page, and a hangingIndent list with a blank item drawn as a marker alone. It now reads its items with DocxMarkdown.items and subtracts the marks of the marker set before each flat item. A * marker no longer hides a dropped **.

  • DocxFontTable reads each item as the page does (DocxMarkdown.items) and ships the faces of its pieces. A tree's marker is read with its item, so a bold list ships the regular face the page sets it in. A typed marker is taken off first, so * Java in a bold list ships no regular face it does not use.

Verification

  • ./mvnw -B -ntp install -pl :graph-compose-render-docx → BUILD SUCCESS, 1218 tests, 0 failures, 1 skipped.
  • New DocxListMarkdownTest (14) covers:
    • a flat Word list's items in pieces, a wrapped item among them, with no note;
    • items whose markers stand in a column, and items after a drawn marker;
    • tree items, with and without hangingIndent, at three levels, a mark the parser keeps left whole;
    • a list written as a paragraph per item: its marker and indent before the pieces, and, with hangingIndent, the pieces after its marker;
    • a bold list: the tree's marker written as characters in the face the page sets it in, and a flat item the parser sets regular (node_js) written regular;
    • a bold Word list's tree items in pieces, a mark the parser keeps (Kot_lin) set regular, and an item the page does not read left bold;
    • a tree marker the parser reads (*a*), written as authored and named;
    • a heading in an item, named once for two;
    • items the page sets none of the text of, with a marker before them and in a column;
    • an item run onto the next page, a hangingIndent list with a blank item, and a * marker before items one of which runs on, each leaving its list's items written as authored and named;
    • markdown off, and a mark the parser keeps, written as they stand;
    • Arabic, written as authored and named, with a marker of marks before it;
    • the font table's faces for tree and flat items, a tree marker's face, and no face for a typed marker.
  • DocxMarkdownTest gains unit tests:
    • items: flat, in a column, markerless, a tree with a level of no marker and an item of runs, and a tree with hangingIndent and a typed marker;
    • split: a lead that ends inside a piece, spans pieces of one style, is absent, is read by the parser, sits in two styles, does not match a short or a longer piece, is all there is, or is longer than the pieces.
  • DocxMarkdownReportTest: a list's items are not named, flat, in a column, and as a tree.
  • DocxListParityTest.boldLeadIsNotMistakenForAMarker checks both readings: with markdown off the item is **bold** lead stays intact, and with it on, bold lead stays intact. In neither is a mark taken off as a typed marker.
  • Word 16, a bold Lato tree list exported and converted to PDF: • Lead item and ◦ Kot_lin stand in the faces the page sets them in, the bullets' ink within 1% of the page's regular marker.
  • 31 sabotages, each reverting a single decision in the change, are each caught by a test: every reading of an item, the indent and the lead, the lead's face, every write path, the record the note reads, each note, the unmatched list's reading, the font table's reading and each check in split.
  • Core doc guards: 166/0. qa doc and DOCX guards: 50/0.
  • DOCX fidelity corpus: no list item in it is written otherwise than before, and all 62 documents are byte-identical with 2.5-dev. Report notes: 954, unchanged.

Known limits

  • Written as authored, and named where the page drops a mark from the item: the items of a list not matched one by one to the layout's (no layout, a table cell, an item run onto the next page, a hangingIndent list with a blank item), an item set in other letters (Arabic), and a tree marker the parser reads with the item (*a*).
  • An item, or a paragraph, written as authored whose marks the page keeps all of is not named where the page changes only its face (node_js in a bold list in a table cell, set regular) or letters no mark is made of (an ordered item's 1.).
  • Word draws a bold Word list's bullets regular, where the page sets the marker of an item it does not read as markdown bold. That is not named, as before.

Lane: render-docx (DOCX semantic backend) — no public API change.

…ets it

A session reads markdown unless told not to, and the page reads a list
item of plain text holding a mark of emphasis or code as it reads a
paragraph: bold and italic text, the marks dropped. The DOCX export wrote
the item as authored, asterisks and all, and named it on the list's note.

Each item is now read as the page lays it out and matched to the lines
the page laid it out in: a flat item after the marker the page sets
before its first line, an item whose marker stands in its own column as
its text alone, and a nested item of a list without hangingIndent after
the indent and marker the page lays out in its text and reads with it
(DocxMarkdown.split takes them off). Where those lines hold the pieces,
the item is written one run a piece, in every path that writes an item;
Word, or the file, still draws the marker. A heading taller than its
item's line is named, and the font table ships the faces the pieces ask
for.

Written as authored and named, as before: a list whose items are not
matched one by one to the layout's, an item set in other letters, a
Word list's nested marker the page sets in another face than the list's.
An item the parser reads into nothing is now named as such.
… and read items for the font table as the page does

A Word list built as a tree of items, with no hangingIndent, has its
items' markers read with them; the parser sets the marker of a bold list
regular, and Word draws it in the list's face, so such an item is written
as authored. It was named only where the page dropped a mark from it:
`Kot_lin` in a bold list, written bold where the page sets it regular,
went unnamed. It is now named whatever marks the page keeps.

The font table read an item's raw label. It now reads each item as the
page lays it out (DocxMarkdown.items): a tree's marker with its item, so
a bold list ships the regular face its marker is written in, and a typed
marker taken off, so `* Java` in a bold list ships no face it does not
use. The item readings move into DocxMarkdown, with a unit test, and
writeItemText is told whether Word draws the marker.
…t an unmatched list's markers out of its note

A Word list built as a tree of items, with no hangingIndent, has its
items' markers read with them, and the page's parser sets the marker
regular. Such an item was written as authored on the premise that Word
draws the marker in the list's face. Measured in Word 16 on a bold Lato
list, Word draws such a list's bullets regular, whatever the paragraph
mark holds: the bullets' ink matches the page's regular marker. The item
is now written in its pieces, as any other, and the refusal is gone.

A list whose items are not matched to the layout's read its items the
old way: a tree's indent and marker were not read with the item, and the
marker set before each flat item's first line counted towards the marks
its lines hold, so a `*` marker could hide a dropped `**`. It now reads
its items with DocxMarkdown.items and subtracts the markers' marks.

The docs and the CHANGELOG say what is named: an item written as
authored is named where the page drops a mark from it, not where the page
changes only its face or letters no mark is made of.
@DemchaAV
DemchaAV merged commit eb89ee8 into 2.5-dev Oct 7, 2026
22 of 24 checks passed
@DemchaAV
DemchaAV deleted the fix/docx-write-list-markdown branch October 7, 2026 18:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant