diff --git a/CHANGELOG.md b/CHANGELOG.md
index cfac3121d..61ef6e2c0 100644
--- a/CHANGELOG.md
+++ b/CHANGELOG.md
@@ -8,6 +8,41 @@ follow semantic versioning; release dates are ISO 8601.
### Public API
+- **A DOCX export writes a list item the page reads as markdown as the page sets it.** The page
+ reads a list item of plain text holding a mark of emphasis or code as it reads a paragraph (the
+ entry below). The export wrote the item as authored, `**Java**` with its asterisks, and named it
+ on the list's note.
+ - **Each item is read as the page lays it out**, matched to the lines the page laid it out in,
+ and written in the pieces the page sets it in where those lines hold them, as a paragraph is:
+ - a flat list's item, after the marker the page sets before its first line;
+ - an item whose marker stands in a column of its own (`hangingIndent`);
+ - an item, at any level, of a list built as a tree of items
+ (`addItem(String, Consumer)`) with no `hangingIndent`, which the page lays out
+ after its indent and its marker and reads with them: the item's own text is written after
+ them.
+ - **The marker stays where it was.** Word draws a Word list's. The page's parser sets the marker
+ it reads with a tree's item regular, a bold list's too, and Word draws such a list's bullets
+ regular (measured in Word 16). A list written as a paragraph per item writes its marker and
+ nesting indent as characters before the pieces, in the face the page sets them in.
+ - **A heading written taller than its item's line is named**, as a paragraph's is.
+ - **The font table ships the faces the page sets an item's pieces in**, read off the text the
+ page reads: a tree's marker with its item, a marker typed before an item taken off.
+ - **Still written as authored**, and named where the page drops a mark from the item:
+ - the items of a list not matched one by one to the layout's: with no layout, composed in a
+ table cell, an item run onto the next page, or a `hangingIndent` list with a blank item the
+ page draws as a marker alone;
+ - an item of a tree whose marker the parser reads as markdown with it, as `*a*`;
+ - one the page sets in other letters than its text, as Arabic;
+ - an item the parser reads into nothing, as `***`: the page sets none of its text, and the
+ note now says so, where it named the item's marks.
+
+ An item written as authored whose marks the page keeps all of is not named where the page
+ changes only its face — `node_js` in a bold list composed in a table cell, set regular — or
+ letters no mark is made of, as an ordered item's number, `1.`; neither is a paragraph's. Across
+ the DOCX fidelity corpus no list item is written otherwise than before, and the 62 documents are
+ byte-identical. `DocxMarkdown` reads each item as the page lays it out, and `DocxMarkdown.split`
+ takes a tree's indent and marker off an item's pieces, with unit tests.
+
- **A DOCX export writes a paragraph the page reads as markdown as the page sets it.** A session
reads markdown unless it is told not to (`markdown(false)`). The page then sets a paragraph of
plain text holding a mark of emphasis or code: the text its marks style bold or italic, a
@@ -51,8 +86,7 @@ follow semantic versioning; release dates are ISO 8601.
- one the page sets in other letters than its text, as Arabic, which the page shapes before it
reads the marks;
- text the parser reads into nothing, which the page sets as nothing and which went unnamed:
- a lone `*`, an empty list item; `***`, a rule; a line set four spaces in, a block of code;
- - a list's items, as before.
+ a lone `*`, an empty list item; `***`, a rule; a line set four spaces in, a block of code.
- **With no layout**, a paragraph is read line by line for the note too: a list marker opening
a line, `* a_b`, is no longer counted as a mark the page drops.
@@ -198,8 +232,8 @@ follow semantic versioning; release dates are ISO 8601.
The DOCX export wrote the text as authored, so Word showed `**bold**` with its asterisks and
none of the bold, without a note. The note now says the markdown marks are written as letters,
- where they still are — a paragraph is written as the page sets it wherever the page's lines show
- how (the entry above):
+ where they still are — a paragraph or a list item is written as the page sets it wherever the
+ page's lines show how (the entries above):
- the paragraph's note (`ParagraphNode`);
- a zone paragraph's, on its `page zone` note;
- a list's, for its items (`ListNode`).
diff --git a/docs/architecture/backend-capability-matrix.md b/docs/architecture/backend-capability-matrix.md
index b3813f9a9..3d77a2b17 100644
--- a/docs/architecture/backend-capability-matrix.md
+++ b/docs/architecture/backend-capability-matrix.md
@@ -64,7 +64,7 @@ Payload records live in `core` under
| Capability (payload) | PDF (fixed) | PPTX (fixed) | DOCX (semantic) |
|---|---|---|---|
| Paragraph — pre-wrapped lines, runs, alignment (`ParagraphFragmentPayload`) | ✅ `PdfParagraphFragmentRenderHandler` | ✅ `PptxParagraphFragmentRenderHandler` (one absolute, wrap-disabled frame per measured line) | ⚠️ semantic paragraphs (`DocxSemanticBackend`) — each run keeps its own style, falling back to the paragraph's when it has none; a centred or right-aligned left-to-right line of its own, of text alone and untracked, that Word sets a point or more wider or narrower at its half-point size has its letters spaced by the difference (`w:spacing`) and its room reckoned from the page's width; a `linkTarget` becomes a `w:hyperlink`, with a relationship for an address or `w:anchor` for one of the document's own anchors, and a run's own link wins over the paragraph's; a paragraph seated off its baseline (`TextVerticalAlign`) has its runs raised or lowered in the line (`w:position`) by the PDF backend's own correction (`ParagraphSeating`), one shift for the paragraph where the page seats each line by its own; Word and LibreOffice stand an exact line's baseline four fifths of the way down it whatever the face, where the page sets it the face's ascent down, so a paragraph whose face puts the two half a point or more apart — Spectral's, not Lato's — has its text moved to the page's baseline in the same position, matched at its middle line (not yet a list item's or a table text cell's; a picture among it moves with it in Word and stays on its own baseline in LibreOffice); lines a container stacks over one another tighter than their face each end halfway between their letters and the next line's (Word draws an exact line's text on screen only inside the line; its PDF export does not cut it), and the last layer of a shape container on one page, where its line runs past the foot, ends at the foot or below its letters; letters two lines share are split halfway so the page does not move, and a stack that holds a picture keeps its lines' own heights; a `bulletOffset` of spaces becomes the paragraph's indent (`w:ind` left, hanging or first line, by `indentStrategy`) in the flow and in cells, not yet over the flow, in an overlay's left-and-right pair, as a badge's initials or in a header or footer; one with letters in it is not written, its wrapped lines still set after the spaces that cover it; an auto-sized paragraph's text is written at its style's size, not the one the page fits it to; a paragraph a session reads as markdown — the default, unless `markdown(false)` — is written as the page sets it, read through the page's own parser, one run a piece in the face, family, colour, tracking and size the page's laid-out lines hold (an auto-sized one's at its style's size), its marks dropped, wherever the lines hold the pieces' letters so; where they are not read, or hold other letters or none (text the parser reads into nothing, which the page sets as nothing), as authored, its marks as letters; a `bookmark(...)` is Word's `HeadingN`, which Word's outline lists by the text of its Word paragraph — an overlay's pair's whole line, one level for both sides — at no level past the ninth. Outside a header or footer, the paragraph's report note (`ParagraphNode`) names each of these where it moves or renames something: the prefix's letters, and the room a path that writes no prefix leaves out where it moves a line; the size written and the size the page fits the text to, to Word's half point; the marks of a paragraph the page read as markdown and the file holds as letters, where its laid-out lines hold fewer of them than its text, not measured where its lines are not read; a markdown heading written taller than the line the page sets it in, which Word cuts on screen; an outline title that is not the text Word lists, a level past the ninth that shares it with another, and the right side's entry where the left holds the line's level |
-| List hanging indent — a marker column and a content column (`ListBuilder.hangingIndent(true)`, `markerGap(...)`) | ✅ marker and content emitted as separate `ParagraphFragmentPayload` fragments at the resolved `markerX` / `contentX` | ✅ the same fragments — the fixed-layout pipeline resolves the geometry before either backend sees it | ⚠️ the top level only. `DocxSemanticBackend` exports a list as a real Word list — `numbering.xml`, `w:numPr` per item, the level carrying the marker — or, with rich items or a drawn marker, as paragraphs; content and nesting are unaffected. With the flag, the top level's marker column is the layout's — the marker's width and `markerGap`, the text and its wrapped lines where the page sets them — where the gap covers what Word may set the marker wider: a picture at its written size, its edges included, or text in the page's face (embedded, or a standard one Word sets in the same widths) grown to its half-point size, half a point clear. A Word list's level then indents and hangs by that column; a list of paragraphs writes the marker, a tab to a stop there, and hangs the item there. Word places content at absolute indents and has no relative-advance primitive, so without the layout's measure the gap could not be honoured; a Word list without the flag that the layout placed and that does not nest takes the page's column too, the spaces the page sets its wrapped lines after, its marker followed by a space (`w:suff`) and an item that wraps measured at Word's half-point size; a list that nests items, a list built as a tree of items (laid out flattened), and a marker the gap does not clear keep the stated column (180 twips, plus 120 per nesting level) — except, in a list of paragraphs, a nested rich item with no marker, which stands where the layout set its text, its measure weighed at Word's half-point sizes, where the layout's items are matched to the list's; a list that nests only such items sets its top level at the page's column too. The report counts, on the list, the items that stand at a stated column, a space past their marker or two spaces a level in, and names a centred or right-aligned list written flush left, a lineSpacing not written where the layout's items are not the list's own and one wraps (in a list composed in a table cell, its wrapping not measured), a continuationIndent not written where an item of a markerless list or a tree of items without the flag wraps or its wrapping is not measured, the rows the page draws as a marker alone for blank items of a flagged list, which are not written, and items the page reads as markdown, whose marks are written as letters, not yet as the page sets them |
+| List hanging indent — a marker column and a content column (`ListBuilder.hangingIndent(true)`, `markerGap(...)`) | ✅ marker and content emitted as separate `ParagraphFragmentPayload` fragments at the resolved `markerX` / `contentX` | ✅ the same fragments — the fixed-layout pipeline resolves the geometry before either backend sees it | ⚠️ the top level only. `DocxSemanticBackend` exports a list as a real Word list — `numbering.xml`, `w:numPr` per item, the level carrying the marker — or, with rich items or a drawn marker, as paragraphs; content and nesting are unaffected. With the flag, the top level's marker column is the layout's — the marker's width and `markerGap`, the text and its wrapped lines where the page sets them — where the gap covers what Word may set the marker wider: a picture at its written size, its edges included, or text in the page's face (embedded, or a standard one Word sets in the same widths) grown to its half-point size, half a point clear. A Word list's level then indents and hangs by that column; a list of paragraphs writes the marker, a tab to a stop there, and hangs the item there. Word places content at absolute indents and has no relative-advance primitive, so without the layout's measure the gap could not be honoured; a Word list without the flag that the layout placed and that does not nest takes the page's column too, the spaces the page sets its wrapped lines after, its marker followed by a space (`w:suff`) and an item that wraps measured at Word's half-point size; a list that nests items, a list built as a tree of items (laid out flattened), and a marker the gap does not clear keep the stated column (180 twips, plus 120 per nesting level) — except, in a list of paragraphs, a nested rich item with no marker, which stands where the layout set its text, its measure weighed at Word's half-point sizes, where the layout's items are matched to the list's; a list that nests only such items sets its top level at the page's column too. The report counts, on the list, the items that stand at a stated column, a space past their marker or two spaces a level in, and names a centred or right-aligned list written flush left, a lineSpacing not written where the layout's items are not the list's own and one wraps (in a list composed in a table cell, its wrapping not measured), a continuationIndent not written where an item of a markerless list or a tree of items without the flag wraps or its wrapping is not measured, the rows the page draws as a marker alone for blank items of a flagged list, which are not written, and what items the page reads as markdown lose. An item of plain text the page reads as markdown — the default, unless `markdown(false)` — is written as the page sets it, matched to the lines the page laid it out in and read as the page lays it out (a flat item after the marker the page sets before it; a flagged item as its text alone; an item of a list built as a tree of items without the flag after the indent and marker the page reads with it), one run a piece in the face, family, colour, tracking and size those lines hold, Word or the file drawing the marker where it did (Word draws a tree's bullets regular, as the parser sets the marker it reads), the font table shipping the faces the page sets the pieces in; it is written as authored, its marks as letters, where its list's items are not matched one by one to the layout's (no layout, composed in a table cell, an item run onto the next page, a flagged list with a blank item), where its lines hold other letters, where the parser reads a tree's marker as markdown with the item (`*a*`), and where the parser reads it into nothing, and named where the page drops a mark from it or sets none of its text — not where it changes only its face or letters no mark is made of (`1.`); a markdown heading written taller than its item's line is named |
| Inline code/badge chips (`InlineBackground` on text spans) | ✅ `PdfParagraphFragmentRenderHandler` | ✅ `PptxParagraphFragmentRenderHandler` | ⚠️ `DocxSemanticBackend` — the fill becomes the run's own `w:shd`, in a paragraph and in a list item alike, so a badge still reads as a badge. What Word has no way to say is the shape: shading covers the glyph box, so the corner radius and the padding above and below the letters are not in the file, and the export records them. The padding beside the letters is written as the room it takes (`spaceAfterTheLastLetter`): character spacing after the chip's last letter, shaded with it, and after the letter before the chip, unshaded. A chip opening its line or following a picture has no letter before it, so its left padding is not in the file; no space is written after right-to-left letters or after a symbol or emoji. The export records, chip by chip, how each side was written. LibreOffice sets no spacing after a line's last letter, so it does not apply the right padding of a chip that ends a line. A `w:shd` fill is opaque, so a translucent chip is flattened first against what Word paints underneath it — the paragraph's shading, the cell's, or else the colour the page paints under the paragraph, a page background included — so the chip agrees with the file it is in and shows the colour the PDF shows. It stops being translucent, and that is recorded with the rest |
| Inline images (`ParagraphImageSpan`) | ✅ `PdfParagraphFragmentRenderHandler` | ✅ `PptxParagraphFragmentRenderHandler` | ✅ `DocxSemanticBackend.writeInlinePicture` (a picture in its own run where it sits among the words, at its size, inside the run's or the paragraph's link; raised or lowered by `w:position` to where the page's alignment and `baselineOffset` put it, from the layout's measure of the paragraph's first line — in a list, the list's text on a line as tall as the item's own tallest picture; LibreOffice ignores `w:position` on a picture and stands it on the baseline, so a picture the export draws itself (icon, emoji, shape) that the page raises carries the rise as transparent rows and needs no `w:position`, while one the page lowers stands in LibreOffice higher than on the page by as much as the page lowers it — up to the text's descent for a centred icon as tall as its line; the editor clips a picture to an exact line height, so a paragraph holding a picture that leaves its text — past the ascent or the descent, in Word's placement or on the baseline — has its lines written at least the height the picture reaches, grown by the editor rather than clipped, every line of the paragraph since Word has one line height for it, and each as tall as the editor's font makes it — for 14pt text about 2.5pt taller than the page's in LibreOffice; a picture inside its text in both editors keeps the exact height; a paragraph of one line of text in a Word paragraph of its own, with room above for its pictures' reach, keeps an exact line at the page's height of it, the pictures set in it where the page puts them in Word and what their ink reaches past it taken from the gaps around it, and in LibreOffice a lowered picture there stands higher and loses what passes the line's top; its description is the text it stands for or empty) |
| Inline vector shapes (`ParagraphShapeSpan`) | ✅ `PdfParagraphFragmentRenderHandler` | ⚠️ `PptxParagraphFragmentRenderHandler` + `PptxInlineGeometry` (distinct per-corner radii render with the top-left radius — single-adjust preset) | ⚠️ `DocxSemanticBackend.writeInlinePicture` + `DocxShapePictures` (a transparent PNG drawn by the shared `InlineSvgRasters` from the outline, fill and stroke — every outline kind, each layer centred in the run's box — placed as an inline picture is; the picture takes as far as the stroked ink reaches past the outline — half the stroke on an edge, more at a sharp corner's miter — and a pixel on each side, measured side by side, and is lowered by what it takes below, so no edge is cut and a shape takes that much more room in the line; a list marker that draws a disc is its picture; at the top level of a `hangingIndent(true)` list that does not nest it is followed by a tab to where the layout starts the item's text, the item's lines hanging there, when the picture clears that stop, else by a space) |
diff --git a/docs/recipes/docx-export.md b/docs/recipes/docx-export.md
index 11382c7f6..a551e78d0 100644
--- a/docs/recipes/docx-export.md
+++ b/docs/recipes/docx-export.md
@@ -68,7 +68,7 @@ creation date is real metadata.
| Document node | DOCX output |
|---|---|
| Paragraphs | Word paragraphs with alignment, font, size, colour (a translucent one with its transparency, as Word's text fill — see "Translucent colours"), bold/italic/underline; inline runs preserved; a `\n` in the text is a line break (`w:br`), which Word would otherwise read as a space; a table cell's `text("a\nb")` stays one line, as the page sets it; a `bulletOffset` of spaces is an indent — the wrapped lines (`FROM_SECOND_LINE`), the first (`FIRST_LINE`) or all of them start that far in, measured in the paragraph's style as the page measures it; a prefix with letters in it is not written, but the wrapped lines still start after the spaces the page covers it with; no prefix is applied to a paragraph written over the flow, as one of an overlay's left-and-right pair, as a badge's initials, or in a header or footer; the report names a prefix's letters, and the room a prefix sets lines in by where a path that writes none leaves it out and it moves a line — on the paragraph, or in a header or footer on the zone's note. A session reads markdown unless told not to (`markdown(false)`): it sets a paragraph of plain text's emphasis marks as the style of the text they mark, a heading line bold (the first three levels larger), and drops its marks; the export writes the text as the page sets it — read line by line through the page's own parser, one run a piece in the face, family, colour, tracking and size the page sets it in, `**Java**` a bold run reading `Java`, a linked paragraph's pieces in one link, Word's outline listing a heading by the text written — wherever it writes a paragraph, a page zone's and a badge's included (a badge's initials counted as the page sets them, and written in the flow where they stand in two faces or as a heading, as initials in two runs' faces are). The page's laid-out lines decide it: the pieces are written where those lines hold their letters in their faces, families, colours and tracking, at their sizes or, for an auto-sized paragraph, at sizes in the same proportion; a session that reads no markdown has its marks written as they stand, and text the parser changes nothing of, an underscore inside a word, is written as it stands too. A paragraph composed in a table cell takes only lines that set its pieces so. The page sets a markdown heading in a line as tall as the paragraph's own and draws its letters past it; written in that exact line, Word cuts their tops on screen, and the report names a heading written taller than its line. The faces are the page's: where the session reads markdown its parser sets every piece in a face of its own, the paragraph's left aside, so a bold paragraph's `Senior_Engineer` is written regular, as the page sets it. A paragraph whose lines are not read — with no layout, composed in a table cell whose text, as authored or as the page reads it, no line of its table carries, or a page zone's the layout shows none of — or that the page sets in other letters than its text, as Arabic, which the page shapes before it reads the marks, or that the parser reads into nothing — a lone `*`, an empty list item; `***`, a rule; a line set four spaces in, a block of code — which the page sets as nothing, is written as authored, marks and all, and the report names it, saying where the lines are not read that whether the page reads its marks is not measured, where the page's parser drops one. An auto-sized paragraph (`autoSize`) is written at its style's size, not the one the page fits its text to, and the report names both sizes where Word, to its half point, holds them apart, on the paragraph or, in a header or footer, on the zone's note. A body paragraph the layout moves to a new page keeps the space the layout leaves above its text there — its own top edge, the edges of the containers opening with it, and the gap before it where the gap did not fit at the foot of the page above — as an empty line that tall before it, kept with it, since Word drops a paragraph's space above at the top of a page and keeps a line's height; the rest of the space owed stays above that line, so Word breaks the page where it did, and the gap between the paragraph's lines comes off the line as it would off its space above. A paragraph in a table cell, an overlay or a list is left to its container |
-| Lists | Real Word lists: a `numbering.xml` definition per list, `w:numPr` on each item, and the authored marker as the level's text. Nesting is a list level, so Enter continues the list and Tab demotes an item. See "What a list becomes" below for the kinds that stay plain paragraphs |
+| Lists | Real Word lists: a `numbering.xml` definition per list, `w:numPr` on each item, and the authored marker as the level's text. Nesting is a list level, so Enter continues the list and Tab demotes an item. An item the page reads as markdown is written as the page sets it, as a paragraph is. See "What a list becomes" below for the kinds that stay plain paragraphs |
| Tables | Word tables, one cell per cell. Each cell states its own padding, on all four sides, so a row is as tall as the page draws it: as `w:tcMar`, and above and below partly in its paragraphs. Word and LibreOffice give every cell of a row the largest top and bottom margin of any cell in it, so a row's cells are written with its smallest, and the rest of a cell's padding above and below is space above its first paragraph and below its last (measured: a row whose day cells were padded 5.5pt above and 10.25pt below beside a label padded 0.75pt stood 60.3pt tall in both editors, where its tallest cell came to 46). A cell opening with a table has no paragraph above it to hold its padding, and a cell in a vertical merge has its bottom edge in another row: these keep their margins, and the row's comes down no lower than the largest of them. Its padding above and below gives up the room Word makes for the table's horizontal rules — half of a rule between two rows to each, the lower row's rule where the two differ, the rule above the table and the one below it whole to their row — which the page does not (measured: a 0.75pt rule made each row 0.75pt taller). A row held at the page's height is written less its margins and those rules too — a rule and a half in the first row and the last, two in a table of one row — since both editors read a row's written height as its cells' content (measured: held less one rule, a table ruled at 0.75pt stood 0.46pt taller in its first row and 0.36pt in its last). A cell that holds nothing but an empty line — a row that is only a rule, its thickness the empty cell's font — has that line cut to the room its row leaves it, the page's row less the cell's own margins and border: the page draws the rule's borders across the line, and Word and LibreOffice keep them outside it and grow the row (measured: `CobaltRota`'s two rules under a 0.9pt border stood 0.9pt taller each). A line with letters, a picture or a paragraph border of its own keeps its height. A table that states no rule is written with the engine's default 1pt black rule, as the page draws it, not left on Word's thinner grid. Its `textAnchor` becomes `w:vAlign` and the paragraph's `w:jc`, with the engine's default — the vertical middle, on the left — where Word's is the top, so a line beside a taller neighbour sits where the page puts it and an amount column stays right-aligned. A cell with no style of its own is set in the engine's default cell face rather than the document's Normal. A cell's lines are one paragraph with line breaks; when its style's `lineSpacing(...)` is above zero and it has more than one line, they are a paragraph each, with the spacing after every line but the last, since Word has no space between the lines of one paragraph but a taller line. A column sized to its content gets a point more than the page gives it, so the editor's font substitute cannot wrap its widest cell. The width is written when the document states one or every column is fixed; otherwise Word sizes the table — see "What falls back". A table breaks across pages where the layout breaks it: every row the layout placed is kept whole (`w:cantSplit`), `repeatHeader(n)` rows repeat on each page (`w:tblHeader`) and stay with the row under them. Two tables in a row — rows included, since a row is carried as a table — are kept apart by a paragraph a tenth of a point tall, holding the rest of the gap between them: an editor joins two tables with nothing between them into one. A table or a row the layout moves to a new page keeps its own top edge there, as the page does — written as a line that tall, kept with it, since Word drops a paragraph's space above at the top of a page — while the gap between it and the block before stays at the foot of the page above, where it fits there; a gap the layout carries onto the new page, because it did not fit at the foot of the page above, is not yet held above a table (body paragraphs and spacers: see their rows). A table's margin is its indent and the space round it, and its padding on the sides holds its rows in as the page draws them: its left side is in the indent too, and both are out of the room its columns are given |
| Composed cells (`DocumentTableCell.node(...)`) | Written by the same writers that write that node anywhere else, so a cell built from an image, a list or a table carries it. A nested table is a real `w:tbl` followed by the paragraph Word requires a cell to end with — a hairline, which the paragraph written next in the cell takes over, so no empty line opens under the table, and whose mark is hidden where it is left at the cell's end holding nothing and no space, since LibreOffice lays it out — and takes the width of the column it sits in, less its own margins and padding — the column's, not the one the page gives it, because the layout reports a composed cell's content under the owner's path |
| Inline chips (`inlineCode(...)`, `inlineChip(...)`, `highlight(...)`) | The chip's fill becomes the run's own `w:shd`, in a paragraph and in a list item alike. Its shape does not travel — see "What a chip keeps and loses" below |
@@ -458,9 +458,30 @@ continuationIndent is not written in a list the page sets it in, a markerless li
items without `hangingIndent`, where an item wraps or its wrapping is not measured; it counts
the items that stand at a stated column, a space past their marker or two spaces a level in,
rather than where the page sets them; it names the rows a `hangingIndent` list draws as a
-marker alone for blank items, which the export does not write; and it names items of plain text
-the page reads as markdown, whose marks are written as letters. A paragraph's are written as the
-page sets them; a list item's not yet.
+marker alone for blank items, which the export does not write; and it names what items of plain
+text the page reads as markdown lose.
+
+An item of plain text the page reads as markdown is written as the page sets it, as a paragraph
+is: matched to the lines the page laid it out in, and written in the pieces the page sets it in
+where those lines hold them — its marks dropped, each piece in the page's face and size. A flat
+list's item is read after the marker the page sets before its first line; an item whose marker
+stands in a column of its own, as its text alone; an item, at any level, of a list built as a
+tree of items without `hangingIndent`, which the page lays out after its indent and its marker
+and reads with them, as its own text after them. The marker stays where it was: Word draws a
+Word list's — the parser sets the marker it reads with a tree's item regular, a bold list's too,
+and Word draws such a list's bullets regular (measured in Word 16) — and a list of paragraphs
+writes its marker and nesting indent as characters before the pieces, in the face the page sets
+them in. The font table ships the faces the page sets the pieces in. The report names a heading
+written taller than its item's line, which Word cuts on screen. An item is written as authored,
+its marks as letters, where its list's items are not matched one by one to the layout's — with
+no layout, composed in a table cell, an item run onto the next page, a `hangingIndent` list with
+a blank item the page draws as a marker alone — where its lines hold other letters than its
+pieces, as Arabic, or where the parser reads the marker of a tree's item as markdown with it, as
+`*a*`; the report names it where the page drops a mark from it. An item the parser reads into
+nothing, as `***`, is written as authored, and named as an item the page sets none of the text
+of. An item written as authored whose marks the page keeps all of is not named where the page
+changes only its face — `node_js` in a bold list in a table cell, set regular — or letters no
+mark is made of, as an ordered item's number, `1.`; nor is a paragraph's.
## What a panel keeps and loses
diff --git a/render-docx/README.md b/render-docx/README.md
index 34f9d75f4..a92d3694f 100644
--- a/render-docx/README.md
+++ b/render-docx/README.md
@@ -127,7 +127,15 @@ What is not written — each one is named in the export report
- the marker column and `markerGap` of an item at a stated column — a list that nests, or a gap
too narrow for its marker — while a flat hanging-indent list keeps the page's column;
- a row the page draws as a marker alone, for a blank item;
- - the marks of items the page reads as markdown, written as letters.
+ - the marks of items the page reads as markdown, where they are written as letters and the page
+ drops one: an item is written as the page sets it — its marks dropped, each piece in the
+ page's face and size — wherever its lines show how, and as authored where its list's items are
+ not matched one by one to the layout's (no layout, a cell, an item run onto the next page, a
+ hanging-indent list with a blank item), its lines hold other letters, the parser reads a
+ tree's marker as markdown with it (`*a*`), or the page sets none of its text, which is named
+ as such; and a markdown heading written taller than its line. An item written as authored
+ whose marks the page keeps is not named where the page changes only its face or letters no
+ mark is made of (`1.`).
- **What a paragraph's own fields set where Word cannot hold it**, named in the export report on
the paragraph outside a header or footer (a page zone's are named on the zone):
- the size an auto-sized paragraph's text is fitted to, where Word, to its half point, holds it
diff --git a/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxFontTable.java b/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxFontTable.java
index d4f707f98..8e3ab945d 100644
--- a/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxFontTable.java
+++ b/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxFontTable.java
@@ -315,10 +315,10 @@ private static void addFace(List faces,
}
/**
- * Collects every face the tree names, wherever a style can sit, and the faces a paragraph's
- * markdown sets pieces of it in ({@link DocxMarkdown}): the table is written before any
- * paragraph is, so a face the pieces ask for is shipped whether or not the page sets them —
- * a session that reads no markdown ships one it does not use.
+ * Collects every face the tree names, wherever a style can sit, and the faces the markdown of
+ * a paragraph or a list's item sets pieces of it in ({@link DocxMarkdown}): the table is
+ * written before any paragraph is, so a face the pieces ask for is shipped whether or not the
+ * page sets them — a session that reads no markdown ships one it does not use.
*/
private static void collectFonts(DocumentNode node, Map> into) {
if (node instanceof ParagraphNode paragraph) {
@@ -335,6 +335,15 @@ private static void collectFonts(DocumentNode node, Map> int
}
} else if (node instanceof ListNode list) {
add(list.textStyle(), into);
+ if (list.textStyle() != null) {
+ // Read as the page reads them: a typed marker taken off, a tree's indent and
+ // marker read with the item, in the face the page sets them in.
+ for (DocxMarkdown.ItemReading item : DocxMarkdown.items(list)) {
+ if (DocxMarkdown.holdsAMark(item.text())) {
+ DocxMarkdown.read(item.text(), list.textStyle()).forEach(piece -> add(piece.style(), into));
+ }
+ }
+ }
} else if (node instanceof TableNode table) {
add(styleOf(table.defaultCellStyle()), into);
table.rowStyles().values().forEach(style -> add(styleOf(style), into));
diff --git a/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxMarkdown.java b/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxMarkdown.java
index 5eefe2d8c..2e84639ae 100644
--- a/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxMarkdown.java
+++ b/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxMarkdown.java
@@ -3,6 +3,9 @@
import com.demcha.compose.document.layout.payloads.ParagraphLine;
import com.demcha.compose.document.layout.payloads.ParagraphSpan;
import com.demcha.compose.document.layout.payloads.ParagraphTextSpan;
+import com.demcha.compose.document.node.ListItem;
+import com.demcha.compose.document.node.ListMarker;
+import com.demcha.compose.document.node.ListNode;
import com.demcha.compose.document.node.ParagraphNode;
import com.demcha.compose.document.style.DocumentLetterSpacing;
import com.demcha.compose.document.style.DocumentTextDecoration;
@@ -20,8 +23,8 @@
import java.util.Objects;
/**
- * A paragraph's text as the page sets it where its session reads markdown: the pieces its marks
- * style, each in its face and size, the marks the parser reads dropped.
+ * A paragraph's or a list item's text as the page sets it where its session reads markdown: the
+ * pieces its marks style, each in its face and size, the marks the parser reads dropped.
*
* The page reads the text line by line, each through {@link MarkDownParser}, a line opening
* with {@code -}, {@code *} or {@code +} and a space keeping that marker and the space in the
@@ -43,6 +46,13 @@ final class DocxMarkdown {
/** How far apart two sizes may be and still be the same, in points. */
private static final double SIZE_CLEARANCE = 0.01;
+ /**
+ * The indent a list with no {@code hangingIndent} sets an item of a tree of items in, a level
+ * at a time: two no-break spaces ({@code TextFlowSupport}, where it flattens the tree into
+ * labels), which the parser reads as letters.
+ */
+ private static final String NESTED_ITEM_INDENT = Character.toString(0x00A0).repeat(2);
+
private DocxMarkdown() {
}
@@ -63,6 +73,76 @@ static boolean mayRead(ParagraphNode node) {
return (node.inlineRuns() == null || node.inlineRuns().isEmpty()) && holdsAMark(node.text());
}
+ /**
+ * A list item's text as the page lays it out, and reads it where its session reads markdown.
+ *
+ * @param text the text the page lays the item out in and reads
+ * @param lead the characters it opens with that the file writes apart from the item's own
+ * text: the indent and marker a list of a tree of items with no
+ * {@code hangingIndent} lays out in its text; empty for none
+ * @param prefix the prefix the page sets before its first line, apart from its text: a flat
+ * list's marker, where it has no {@code hangingIndent}; empty for none
+ */
+ record ItemReading(String text, String lead, String prefix) {
+ }
+
+ /**
+ * A flat list's item as the page reads it: its text, a marker typed before it taken off, after
+ * the list's marker where the page sets that before its first line.
+ *
+ * @param normalized the item's text, a typed marker taken off
+ */
+ static ItemReading flatItem(ListNode list, String normalized) {
+ return new ItemReading(normalized, "",
+ !list.hangingIndent() && list.marker().isVisible() ? list.marker().prefix() : "");
+ }
+
+ /**
+ * An item of a tree of items as the page reads it: with {@code hangingIndent}, its label, a
+ * marker typed before it taken off; without it, its label after its depth's indent and its
+ * marker, as the page flattens the tree into labels, nothing taken off.
+ *
+ * @param marker the marker the item takes, its own or its depth's
+ * @return its reading, {@code null} for an item of runs, which the page never reads
+ */
+ static ItemReading nestedItem(ListNode list, ListItem item, int depth, ListMarker marker) {
+ if (item.isRich()) {
+ return null;
+ }
+ if (list.hangingIndent()) {
+ return new ItemReading(ListMarker.normalizeItemText(item.label(), list.normalizeMarkers()), "", "");
+ }
+ String lead = NESTED_ITEM_INDENT.repeat(depth) + (marker.isVisible() ? marker.prefix() : "");
+ return new ItemReading(lead + item.label(), lead, "");
+ }
+
+ /**
+ * Every item of plain text a list writes, as the page reads it: its flat items, then its tree
+ * of items depth first, each taking its own marker or its depth's.
+ */
+ static List items(ListNode list) {
+ List readings = new ArrayList<>();
+ for (String item : list.items()) {
+ String normalized = ListMarker.normalizeItemText(item, list.normalizeMarkers());
+ if (!normalized.isBlank()) {
+ readings.add(flatItem(list, normalized));
+ }
+ }
+ addItems(list, list.nestedItems(), 0, readings);
+ return List.copyOf(readings);
+ }
+
+ private static void addItems(ListNode list, List items, int depth, List into) {
+ for (ListItem item : items) {
+ ItemReading reading = nestedItem(list, item,
+ depth, item.marker() != null ? item.marker() : ListMarker.defaultForDepth(depth));
+ if (reading != null) {
+ into.add(reading);
+ }
+ addItems(list, item.children(), depth + 1, into);
+ }
+ }
+
/**
* One piece of text the page sets in one style.
*
@@ -149,6 +229,60 @@ private static void add(List pieces, String text, DocumentTextStyle style
}
}
+ /**
+ * Pieces read off text that opens with a lead the file writes apart from them, split at the
+ * lead's end.
+ *
+ * @param leadStyle the style the page sets the lead in, {@code null} where there is no lead
+ * @param after the pieces after the lead
+ */
+ record Split(DocumentTextStyle leadStyle, List after) {
+ }
+
+ /**
+ * Splits pieces read off text that opens with a lead — a nested item's indent and marker, which
+ * a list with no {@code hangingIndent} lays out in the item's text — at the lead's end.
+ *
+ * @param pieces the pieces read off the text
+ * @param lead the characters the text opens with, empty for none
+ * @return the split, or {@code null} where the pieces do not open with the lead's characters,
+ * every one in one style, or hold nothing after it
+ */
+ static Split split(List pieces, String lead) {
+ if (lead.isEmpty()) {
+ return new Split(null, pieces);
+ }
+ DocumentTextStyle style = null;
+ List after = new ArrayList<>();
+ int taken = 0;
+ int index = 0;
+ for (; index < pieces.size() && taken < lead.length(); index++) {
+ Piece piece = pieces.get(index);
+ if (style != null && !piece.style().equals(style)) {
+ return null;
+ }
+ style = piece.style();
+ String left = lead.substring(taken);
+ if (piece.text().length() <= left.length()) {
+ if (!left.startsWith(piece.text())) {
+ return null;
+ }
+ taken += piece.text().length();
+ } else {
+ if (!piece.text().startsWith(left)) {
+ return null;
+ }
+ after.add(new Piece(piece.text().substring(left.length()), piece.style()));
+ taken = lead.length();
+ }
+ }
+ if (taken < lead.length()) {
+ return null;
+ }
+ after.addAll(pieces.subList(index, pieces.size()));
+ return after.isEmpty() ? null : new Split(style, List.copyOf(after));
+ }
+
/** The pieces' text, as Word's paragraph holds it. */
static String text(List pieces) {
StringBuilder text = new StringBuilder();
diff --git a/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxSemanticBackend.java b/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxSemanticBackend.java
index 4a27d43df..db52722bd 100644
--- a/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxSemanticBackend.java
+++ b/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxSemanticBackend.java
@@ -450,6 +450,8 @@ public final class DocxSemanticBackend implements SemanticBackend {
*/
private java.util.ArrayDeque listItemLines =
new java.util.ArrayDeque<>();
+ /** The items of plain text the list being written has written so far, as its note reads them. */
+ private List listItemsWritten = new ArrayList<>();
/** Whether the list being written has an item above the one about to be written. */
private boolean anItemWasWritten;
@@ -3915,12 +3917,13 @@ private static String levelText(com.demcha.compose.document.node.ListMarker mark
}
/**
- * Semantic list mapping: each item becomes a marker-prefixed paragraph in
- * the list's text style. Flat items run through the same
- * {@code ListMarker.normalizeItemText} step as fixed-layout rendering
- * (author-typed markers stripped, blank items skipped); nested items
- * indent two spaces per depth and use their own marker when one is set,
- * falling back to {@code ListMarker.defaultForDepth} otherwise.
+ * Semantic list mapping: each item becomes a paragraph, a Word list's or one
+ * with its marker as characters, in the list's text style, or in the pieces the
+ * page sets it in where it reads the item as markdown ({@link #writeItemText}).
+ * Flat items run through the same {@code ListMarker.normalizeItemText} step as
+ * fixed-layout rendering (author-typed markers stripped, blank items skipped);
+ * nested items indent two spaces per depth and use their own marker when one is
+ * set, falling back to {@code ListMarker.defaultForDepth} otherwise.
*/
private void writeList(XWPFDocument document,
com.demcha.compose.document.node.ListNode list) {
@@ -3948,6 +3951,9 @@ private void writeList(XWPFDocument document,
: new java.util.ArrayDeque<>();
boolean previousMatched = listItemsMatched;
listItemsMatched = !laidOut.isEmpty() && laidOut.size() == itemCount(list);
+ List previousWritten = listItemsWritten;
+ List written = new ArrayList<>();
+ listItemsWritten = written;
// Where the list's items start on the page, past its margin and padding: an item's text
// stands its own distance past it (see indentAsTheItemIs).
double previousTextLeft = listTextLeft;
@@ -3975,11 +3981,12 @@ private void writeList(XWPFDocument document,
listItemLines = previousItemLines;
listTextLeft = previousTextLeft;
listItemsMatched = previousMatched;
+ listItemsWritten = previousWritten;
}
owePendingSpacingAfter(list.margin().bottom() + list.padding().bottom());
reportWrittenWithout(list, itemCount(list) == 0 ? "writes no paragraph"
: numId != null ? "written as a Word list" : "written as a paragraph per item",
- listLost(list, laidOut, atTheirColumn));
+ listLost(list, laidOut, atTheirColumn, written));
}
/**
@@ -3997,16 +4004,20 @@ private void writeList(XWPFDocument document,
* the page sets them: with {@code hangingIndent}, its marker column and markerGap. A stated
* column is not measured against the page's, so one the page's happens to equal counts too.
* The rows the page draws as a marker alone, which the export does not write
- * ({@link #markerOnlyRows}). And the marks of items the page reads as markdown, which are
- * written as letters ({@link #itemsMarkdownLost}).
+ * ({@link #markerOnlyRows}). And what items the page reads as markdown lose
+ * ({@link #itemsMarkdownLost}).
*
* @param laidOut the items as the layout laid them out
* @param atTheirColumn how many of its items stand where the page sets them
+ * @param written its items of plain text as written
*/
private List listLost(com.demcha.compose.document.node.ListNode list,
- List laidOut, int atTheirColumn) {
+ List laidOut, int atTheirColumn,
+ List written) {
List lost = new ArrayList<>();
int items = itemCount(list);
+ // The layout's items matched to the list's one by one, as writeList matches them.
+ boolean matched = !laidOut.isEmpty() && laidOut.size() == items;
int markerRows = markerOnlyRows(list);
if (items > 0) {
lost.addAll(itemsLost(list, laidOut, items, markerRows, atTheirColumn));
@@ -4015,7 +4026,55 @@ private List listLost(com.demcha.compose.document.node.ListNode list,
lost.add(markerRows + (markerRows == 1 ? " row the page draws as a marker alone, for a blank item, is"
: " rows the page draws as a marker alone, for blank items, are") + " not written");
}
- String marks = itemsMarkdownLost(list);
+ lost.addAll(matched ? itemsMarkdownLost(list, items, written) : itemsMarkdownUnmatched(list));
+ return lost;
+ }
+
+ /**
+ * What a list's items lose where the page reads them as markdown, each item matched to the
+ * lines the page laid it out in. An item of plain text holding a mark of emphasis or code is
+ * written in the pieces the page sets it in where its lines say so ({@link #itemPieces}); a
+ * heading among them in a line Word cuts it in is named ({@link #headingCut}). One the page
+ * reads into nothing is written as authored and named; any other written as authored, its
+ * marks and all, is named where its lines, less the prefix the page sets before its first line,
+ * hold fewer marks than it ({@link #marksDropped}).
+ *
+ * @param items how many items the list writes
+ * @param written its items of plain text as written
+ */
+ private List itemsMarkdownLost(com.demcha.compose.document.node.ListNode list, int items,
+ List written) {
+ List lost = new ArrayList<>();
+ for (ItemWritten item : written) {
+ String cut = item.pieces() == null ? null
+ : headingCut(item.pieces(), list.textStyle(), item.lines(), "its items'", "item's");
+ if (cut != null) {
+ lost.add(cut);
+ break;
+ }
+ }
+ List asAuthored = written.stream()
+ .filter(item -> item.pieces() == null && DocxMarkdown.holdsAMark(item.reading().text())).toList();
+ List nothing = asAuthored.stream().filter(item -> setsNoneOfIt(item, list.textStyle())).toList();
+ if (!nothing.isEmpty()) {
+ boolean one = nothing.size() == 1;
+ lost.add(nothing.size() + " of its " + items + " items " + (one ? "is" : "are")
+ + " written as authored, where the page reads " + (one ? "it" : "them")
+ + " as markdown and sets none of " + (one ? "its" : "their") + " text");
+ }
+ List lines = new ArrayList<>();
+ int authored = 0;
+ int prefixMarks = 0;
+ List texts = new ArrayList<>();
+ for (ItemWritten item : asAuthored) {
+ if (!nothing.contains(item)) {
+ lines.addAll(item.lines());
+ authored += markdownMarksIn(item.reading().text());
+ prefixMarks += markdownMarksIn(item.reading().prefix());
+ texts.add(item.reading().text());
+ }
+ }
+ String marks = texts.isEmpty() ? null : marksDropped("its items'", lines, authored, prefixMarks, texts);
if (marks != null) {
lost.add(marks);
}
@@ -4023,52 +4082,130 @@ private List listLost(com.demcha.compose.document.node.ListNode list,
}
/**
- * What a list's items lose where the page reads them as markdown, as a paragraph's text is read
- * ({@link #markdownLost}): an item of plain text holding a mark of emphasis or code is set as
- * markdown, its marks dropped, and written as authored. Its text is read as the page lays it
- * out, a marker typed before it taken off
- * ({@link com.demcha.compose.document.node.ListMarker#normalizeItemText}); its laid-out lines
- * are the items', without a marker laid out on its own. A marker laid out in an item's line,
- * and a rich item's runs, hold every mark they had.
+ * Whether the page reads an item written as authored into nothing — a rule, a block of code it
+ * keeps no text of — and sets none of its text: its lines hold no letter past the prefix the
+ * page sets before them, a flat list's marker.
*/
- private String itemsMarkdownLost(com.demcha.compose.document.node.ListNode list) {
- List plain = new ArrayList<>();
- List all = new ArrayList<>();
- if (list.nestedItems().isEmpty()) {
- for (String item : list.items()) {
- String text = com.demcha.compose.document.node.ListMarker.normalizeItemText(item, list.normalizeMarkers());
- plain.add(text);
- all.add(text);
- }
- } else {
- collectItemTexts(list.nestedItems(), list.normalizeMarkers(), plain, all);
+ private static boolean setsNoneOfIt(ItemWritten item, DocumentTextStyle style) {
+ if (item.lines().isEmpty() || item.reading().text().isBlank()
+ || !DocxMarkdown.text(DocxMarkdown.read(item.reading().text(),
+ style == null ? DocumentTextStyle.DEFAULT : style)).isBlank()) {
+ return false;
+ }
+ long letters = 0;
+ for (com.demcha.compose.document.layout.payloads.ParagraphLine line : item.lines()) {
+ letters += lettersIn(line.text());
}
+ return letters <= lettersIn(item.reading().prefix());
+ }
+
+ private static long lettersIn(String text) {
+ return text.codePoints().filter(codePoint -> !Character.isWhitespace(codePoint)).count();
+ }
+
+ /**
+ * What a list's items lose where the page reads them as markdown, its items not matched to the
+ * layout's — with no layout, composed in a table cell, one run onto the next page, or a blank
+ * item of a {@code hangingIndent} list the page draws as a marker alone — and so each written
+ * as authored, as a paragraph's text is ({@link #markdownLost}): an item of plain text holding a
+ * mark of emphasis or code is set as markdown, its marks dropped. Its text is read as the page
+ * lays it out ({@link DocxMarkdown#items}): a marker typed before it taken off, a tree's indent
+ * and marker read with it. Its laid-out lines are the items', without a marker laid out on its
+ * own, and less the marks of the marker the page sets before a flat item's first line. A rich
+ * item's runs hold every mark they had.
+ */
+ private List itemsMarkdownUnmatched(com.demcha.compose.document.node.ListNode list) {
+ List readings = DocxMarkdown.items(list);
+ List plain = readings.stream().map(DocxMarkdown.ItemReading::text).toList();
if (plain.stream().noneMatch(DocxMarkdown::holdsAMark)) {
- return null;
+ return List.of();
}
+ List all = new ArrayList<>(plain);
+ collectRichTexts(list.nestedItems(), all);
List lines = new ArrayList<>();
for (DocxLayoutMetrics.ItemText item : layout.itemLines(list)) {
lines.addAll(item.lines());
}
- return marksDropped("its items'", lines, all.stream().mapToInt(DocxSemanticBackend::markdownMarksIn).sum(), 0,
- plain);
+ String marks = marksDropped("its items'", lines, all.stream().mapToInt(DocxSemanticBackend::markdownMarksIn).sum(),
+ readings.stream().mapToInt(reading -> markdownMarksIn(reading.prefix())).sum(), plain);
+ return marks == null ? List.of() : List.of(marks);
+ }
+
+ /**
+ * A list item of plain text as written.
+ *
+ * @param reading its text as the page reads it
+ * @param lines the lines the page laid it out in, empty where its list's items are not matched
+ * @param pieces the pieces it was written in, {@code null} where it was written as authored
+ */
+ private record ItemWritten(DocxMarkdown.ItemReading reading,
+ List lines,
+ List pieces) {
}
/**
- * Every item's text in a tree of items, as the page lays it out: a plain item's label, a marker
- * typed before it taken off, and a rich item's runs' text too.
+ * An item's text in the pieces the page sets it in, split at the end of the lead the file
+ * writes apart from them, where the page reads it as markdown and the lines it laid the item
+ * out in hold the pieces ({@link #pagePieces}); {@code null} where it is written as authored:
+ * its list's items are not matched to the layout's, its lines do not hold the pieces, or the
+ * pieces do not open with the lead's characters in one style ({@link DocxMarkdown#split}).
+ *
+ * @param laidOut how the layout set the item, {@code null} where its list's items are not matched
*/
- private static void collectItemTexts(List items, boolean normalizeMarkers,
- List plain, List all) {
+ private static DocxMarkdown.Split itemPieces(DocxMarkdown.ItemReading reading, DocumentTextStyle style,
+ DocxLayoutMetrics.ItemText laidOut) {
+ if (laidOut == null) {
+ return null;
+ }
+ List pieces = pagePieces(reading.text(), style, laidOut.lines(), reading.prefix(), false);
+ return pieces == null ? null : DocxMarkdown.split(pieces, reading.lead());
+ }
+
+ /**
+ * Writes an item's text: as authored in one run, or in the pieces the page sets it in, and
+ * records it for the list's note.
+ *
+ * Where Word draws the marker of a tree's item the page reads with it, it draws it as the
+ * page's parser sets it: a list built as a tree of items without {@code hangingIndent} states
+ * no face on its levels, and Word draws their bullets regular, a bold list's included
+ * (measured in Word 16), where the parser sets the marker it reads regular.
+ *
+ * @param before what the file writes before the item's own text — its indent and marker,
+ * where Word does not draw the marker — in the lead's style where the page sets
+ * the lead in the pieces
+ * @param label the item's text as authored
+ * @param laidOut how the layout set the item, {@code null} where its list's items are not matched
+ */
+ private void writeItemText(XWPFParagraph para, DocumentTextStyle style, String before, String label,
+ DocxMarkdown.ItemReading reading, DocxLayoutMetrics.ItemText laidOut) {
+ DocxMarkdown.Split split = itemPieces(reading, style, laidOut);
+ if (split == null) {
+ XWPFRun run = para.createRun();
+ applyStyle(run, style);
+ setTextBrokenAtLines(run, before + label);
+ } else {
+ if (!before.isEmpty()) {
+ XWPFRun lead = para.createRun();
+ applyStyle(lead, split.leadStyle() != null ? split.leadStyle() : style);
+ setTextBrokenAtLines(lead, before);
+ }
+ for (DocxMarkdown.Piece piece : split.after()) {
+ XWPFRun run = para.createRun();
+ applyStyle(run, piece.style());
+ setTextBrokenAtLines(run, piece.text());
+ }
+ }
+ listItemsWritten.add(new ItemWritten(reading, laidOut == null ? List.of() : laidOut.lines(),
+ split == null ? null : split.after()));
+ }
+
+ /** The text of every item of runs in a tree of items, which the page lays out as it stands. */
+ private static void collectRichTexts(List items, List into) {
for (com.demcha.compose.document.node.ListItem item : items) {
- if (item.runs().isEmpty()) {
- String text = com.demcha.compose.document.node.ListMarker.normalizeItemText(item.label(), normalizeMarkers);
- plain.add(text);
- all.add(text);
- } else {
- all.add(InlineRun.plainText(item.runs()));
+ if (item.isRich()) {
+ into.add(InlineRun.plainText(item.runs()));
}
- collectItemTexts(item.children(), normalizeMarkers, plain, all);
+ collectRichTexts(item.children(), into);
}
}
@@ -4237,26 +4374,27 @@ private int writeListItems(XWPFDocument document,
continue;
}
java.util.OptionalDouble lineHeight = layout.lineHeight(list);
+ DocxMarkdown.ItemReading reading = DocxMarkdown.flatItem(list, normalized);
boolean atItsColumn;
if (list.marker().isRich()) {
// A drawn marker's pieces are runs, so the row is written the way
// any row with runs in it is; its item is still just a label.
atItsColumn = writeRichListLine(document, list.textStyle(), list.marker(),
com.demcha.compose.document.node.ListItem.of(normalized), 0, lineHeight, layout.firstLine(list),
- layout.pathOf(list), list);
+ layout.pathOf(list), list, reading);
} else if (numId != null) {
// Word draws the marker, so the text is the item and nothing else.
- writeListLine(document, list.textStyle(), normalized, 0, numId, lineHeight);
+ writeListLine(document, list.textStyle(), "", normalized, reading, 0, numId, lineHeight);
atItsColumn = measuredColumns.containsKey(numId);
} else if (setInAColumn(list, list.marker())) {
// A list that is not a Word list, its marker in a column the layout set: the item
// is written as a rich one is, its text tabbed to that column.
atItsColumn = writeRichListLine(document, list.textStyle(), list.marker(),
com.demcha.compose.document.node.ListItem.of(normalized), 0, lineHeight, layout.firstLine(list),
- layout.pathOf(list), list);
+ layout.pathOf(list), list, reading);
} else {
- writeListLine(document, list.textStyle(),
- list.marker().prefix() + normalized, 0, null, lineHeight);
+ writeListLine(document, list.textStyle(), list.marker().prefix(), normalized, reading, 0, null,
+ lineHeight);
atItsColumn = standsAsItsLetters(list, 0, list.marker());
}
atTheirColumn += counted(atItsColumn);
@@ -4286,21 +4424,23 @@ private int writeNestedItem(XWPFDocument document,
? item.marker()
: com.demcha.compose.document.node.ListMarker.defaultForDepth(depth);
java.util.OptionalDouble lineHeight = layout.lineHeight(list);
+ DocxMarkdown.ItemReading reading = DocxMarkdown.nestedItem(list, item, depth, marker);
boolean atItsColumn;
if (item.isRich() || marker.isRich()) {
// A nested item stands after its depth's indent, and an item may carry a marker of
// its own: the list's measure of its first item's marker is only a top-level item's
// that carries the list's.
atItsColumn = writeRichListLine(document, list.textStyle(), marker, item, depth, lineHeight,
- layout.firstLine(list), layout.pathOf(list), depth == 0 && marker.equals(list.marker()) ? list : null);
+ layout.firstLine(list), layout.pathOf(list), depth == 0 && marker.equals(list.marker()) ? list : null,
+ reading);
} else if (numId != null) {
- writeListLine(document, list.textStyle(), item.label(), depth, numId, lineHeight);
+ writeListLine(document, list.textStyle(), "", item.label(), reading, depth, numId, lineHeight);
atItsColumn = depth == 0 && measuredColumns.containsKey(numId);
} else if (depth == 0 && marker.equals(list.marker()) && setInAColumn(list, marker)) {
atItsColumn = writeRichListLine(document, list.textStyle(), marker, item, depth, lineHeight,
- layout.firstLine(list), layout.pathOf(list), list);
+ layout.firstLine(list), layout.pathOf(list), list, reading);
} else {
- writeListLine(document, list.textStyle(), marker.prefix() + item.label(), depth, null,
+ writeListLine(document, list.textStyle(), marker.prefix(), item.label(), reading, depth, null,
lineHeight);
atItsColumn = standsAsItsLetters(list, depth, marker);
}
@@ -4315,11 +4455,20 @@ private int writeNestedItem(XWPFDocument document,
* Writes one item, either as a real Word list paragraph or as the marker-prefixed
* text the export used before Word numbering existed here.
*
- * @param numId the list definition to attach, or {@code null} to write the marker
- * and the nesting indent as characters
+ * Its text is written in the pieces the page sets it in where the page reads it as
+ * markdown ({@link #writeItemText}); the paragraph's mark keeps the list's style whatever
+ * piece ends the item.
+ *
+ * @param marker the marker written as characters before the item's text, after the nesting
+ * indent, where Word does not draw it
+ * @param label the item's text as authored
+ * @param reading the item's text as the page reads it
+ * @param numId the list definition to attach, or {@code null} to write the marker
+ * and the nesting indent as characters
*/
private void writeListLine(XWPFDocument document, DocumentTextStyle style,
- String text, int depth, BigInteger numId,
+ String marker, String label, DocxMarkdown.ItemReading reading, int depth,
+ BigInteger numId,
java.util.OptionalDouble lineHeight) {
spaceBeforeTheNextItem();
XWPFParagraph para = newBodyParagraph(document);
@@ -4334,9 +4483,7 @@ private void writeListLine(XWPFDocument document, DocumentTextStyle style,
measureTheItemAtWordsSize(para, style, lines, measuredColumns.get(numId));
}
}
- XWPFRun run = para.createRun();
- applyStyle(run, style);
- setTextBrokenAtLines(run, numId != null ? text : " ".repeat(depth) + text);
+ writeItemText(para, style, numId != null ? "" : " ".repeat(depth) + marker, label, reading, laidOut);
styleTheMark(para, style);
}
@@ -4436,6 +4583,9 @@ private void indentListItemInside(XWPFParagraph para, int depth, BigInteger numI
* @param measured the list the item's marker column is measured from, or {@code null}
* when it is not this item's — a nested item, or one with a marker of
* its own
+ * @param reading an item of plain text as the page reads it, its text written in the pieces
+ * the page sets it in where it reads it as markdown ({@link #writeItemText});
+ * {@code null} for an item of runs
* @return whether its text stands where the page sets it: tabbed to the marker column the
* layout set, at the layout's own place for it, or at the edge with no marker before it
*/
@@ -4445,7 +4595,8 @@ private boolean writeRichListLine(XWPFDocument document, DocumentTextStyle style
int depth,
java.util.OptionalDouble lineHeight,
java.util.Optional line,
- String path, com.demcha.compose.document.node.ListNode measured) {
+ String path, com.demcha.compose.document.node.ListNode measured,
+ DocxMarkdown.ItemReading reading) {
warnDroppedInlineRuns(marker.runs(), path);
warnDroppedInlineRuns(item.runs(), path);
line = line.map(listLine -> itemLine(listLine, item.runs()));
@@ -4519,9 +4670,7 @@ private boolean writeRichListLine(XWPFDocument document, DocumentTextStyle style
}
}
} else {
- XWPFRun label = para.createRun();
- applyStyle(label, style);
- setTextBrokenAtLines(label, item.label());
+ writeItemText(para, style, "", item.label(), reading, laidOut);
}
makeRoomForPictures(para, pictures);
styleTheMark(para, markStyle);
@@ -7531,18 +7680,35 @@ private static String autoSizeLost(ParagraphNode node, List markdownPieces(ParagraphNode node,
List lines) {
- if (!DocxMarkdown.mayRead(node) || node.textStyle() == null || lines.isEmpty()) {
+ if (!DocxMarkdown.mayRead(node)) {
+ return null;
+ }
+ return pagePieces(node.text(), node.textStyle(), lines,
+ setsAPrefixBeforeTheFirstLine(node) ? node.bulletOffset() : "", node.autoSize() != null);
+ }
+
+ /**
+ * Text in the pieces the page sets it in, where the page reads it as markdown and its lines say
+ * so ({@link DocxMarkdown#laidOutIn}); {@code null} where it is written as authored. Pieces
+ * that change nothing — every mark kept, every piece in the text's own style — leave the text
+ * to be written as it stands, its white space and all.
+ *
+ * @param lines the lines the page laid the text out in, empty where they are not read
+ * @param prefix the prefix the page sets before the first line, empty where it sets none
+ * @param fitted whether the page fits the text to a size of its own
+ */
+ private static List pagePieces(String text, DocumentTextStyle style,
+ List lines,
+ String prefix, boolean fitted) {
+ if (!DocxMarkdown.holdsAMark(text) || style == null || lines.isEmpty()) {
return null;
}
- List pieces = DocxMarkdown.read(node.text(), node.textStyle());
- // Pieces that change nothing — every mark kept, every piece in the paragraph's style —
- // leave the text to be written as it stands, its white space and all.
- if (pieces.stream().allMatch(piece -> piece.style().equals(node.textStyle()))
- && markdownMarksIn(DocxMarkdown.text(pieces)) == markdownMarksIn(node.text())) {
+ List pieces = DocxMarkdown.read(text, style);
+ if (pieces.stream().allMatch(piece -> piece.style().equals(style))
+ && markdownMarksIn(DocxMarkdown.text(pieces)) == markdownMarksIn(text)) {
return null;
}
- return DocxMarkdown.laidOutIn(pieces, lines, setsAPrefixBeforeTheFirstLine(node) ? node.bulletOffset() : "",
- node.autoSize() != null) ? pieces : null;
+ return DocxMarkdown.laidOutIn(pieces, lines, prefix, fitted) ? pieces : null;
}
/**
@@ -7559,17 +7725,33 @@ && markdownMarksIn(DocxMarkdown.text(pieces)) == markdownMarksIn(node.text())) {
*/
private String markdownHeadingCut(ParagraphNode node, List lines,
String whose) {
- List pieces = markdownWritten.get(node);
- if (pieces == null || node.textStyle() == null || lines.isEmpty()) {
+ return headingCut(markdownWritten.get(node), node.textStyle(), lines, whose, "paragraph's");
+ }
+
+ /**
+ * What text written in the pieces the page reads its markdown into loses of them: a heading
+ * written taller than the line the page sets it in (see {@link #markdownHeadingCut}).
+ *
+ * @param pieces the pieces written, {@code null} where the text was written as authored
+ * @param style the text's own style
+ * @param lines the lines the page laid the text out in
+ * @param whose whose heading the phrase names
+ * @param owner whose own line the page sets the heading in
+ * @return the phrase, or {@code null} where every piece fits its line
+ */
+ private String headingCut(List pieces, DocumentTextStyle style,
+ List lines,
+ String whose, String owner) {
+ if (pieces == null || style == null || lines.isEmpty()) {
return null;
}
double line = lines.stream().mapToDouble(com.demcha.compose.document.layout.payloads.ParagraphLine::lineHeight)
.max().orElse(0);
for (DocxMarkdown.Piece piece : pieces) {
- if (piece.style().size() > node.textStyle().size()
+ if (piece.style().size() > style.size()
&& styleLineHeight(piece.style()) > line + HEADING_CLEARANCE) {
return whose + " markdown heading is written at " + pointsOf(piece.style().size()) + "pt in a line "
- + pointsOf(line) + "pt tall, as tall as the paragraph's own line on the page: the page draws its "
+ + pointsOf(line) + "pt tall, as tall as the " + owner + " own line on the page: the page draws its "
+ "letters past the line, and Word cuts their tops on screen";
}
}
diff --git a/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxListMarkdownTest.java b/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxListMarkdownTest.java
new file mode 100644
index 000000000..31f93ac43
--- /dev/null
+++ b/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxListMarkdownTest.java
@@ -0,0 +1,298 @@
+package com.demcha.compose.document.backend.semantic.docx;
+
+import com.demcha.compose.GraphCompose;
+import com.demcha.compose.document.api.DocumentSession;
+import com.demcha.compose.document.dsl.ListBuilder;
+import com.demcha.compose.document.node.ListMarker;
+import com.demcha.compose.document.style.DocumentInsets;
+import com.demcha.compose.document.style.DocumentTextDecoration;
+import com.demcha.compose.document.style.DocumentTextStyle;
+import com.demcha.compose.font.FontName;
+import org.apache.poi.openxml4j.opc.PackagePart;
+import org.apache.poi.xwpf.usermodel.XWPFDocument;
+import org.apache.poi.xwpf.usermodel.XWPFParagraph;
+import org.apache.poi.xwpf.usermodel.XWPFRun;
+import org.junit.jupiter.api.Test;
+
+import java.io.ByteArrayInputStream;
+import java.io.InputStream;
+import java.nio.charset.StandardCharsets;
+import java.util.List;
+import java.util.concurrent.atomic.AtomicReference;
+import java.util.function.Consumer;
+
+import static org.assertj.core.api.Assertions.assertThat;
+
+/**
+ * A list item the page reads as markdown is written as the page sets it: its marks dropped and
+ * the text they mark in the face the page sets it in, the marker where Word or the file puts it,
+ * and nothing named but a heading the page draws past its line. Each item is matched to the lines
+ * the page laid it out in, which say whether it reads the item so; where they do not, or the
+ * items are not matched to the layout's, an item is written as authored and named.
+ */
+class DocxListMarkdownTest {
+
+ private static final String MARKS = "its items' markdown marks are written as letters, where the page sets the "
+ + "text they mark and drops them";
+ private static final DocumentTextStyle BOLD = DocumentTextStyle.builder().decoration(DocumentTextDecoration.BOLD).build();
+
+ @Test
+ void aListsItemsAreWrittenInThePiecesThePageSetsThem() throws Exception {
+ Export export = export(true, list -> list.items("**Java** lead", "Kotlin",
+ "a *long* item that runs on past the edge of the narrow page and wraps"));
+ XWPFParagraph java = export.paragraphWith("Java");
+ assertThat(java.getNumID()).as("a Word list, its marker Word's").isNotNull();
+ assertThat(java.getRuns()).extracting(XWPFRun::text).containsExactly("Java", " lead");
+ assertThat(java.getRuns()).extracting(XWPFRun::isBold).containsExactly(true, false);
+ XWPFParagraph wrapped = export.paragraphWith("wraps");
+ assertThat(wrapped.getText()).isEqualTo("a long item that runs on past the edge of the narrow page and wraps");
+ assertThat(wrapped.getRuns()).filteredOn(XWPFRun::isItalic).extracting(XWPFRun::text).containsExactly("long");
+ assertThat(export.paragraphWith("Kotlin").getRuns()).extracting(XWPFRun::text).containsExactly("Kotlin");
+ assertThat(export.notes()).isEmpty();
+ }
+
+ @Test
+ void itemsWhoseMarkersStandInAColumnOfTheirOwnAreWrittenSo() throws Exception {
+ Export column = export(true, list -> list.hangingIndent(true).dash().items("**Java** lead", "`code` here"));
+ assertThat(column.paragraphWith("Java").getRuns()).extracting(XWPFRun::isBold).containsExactly(true, false);
+ assertThat(column.paragraphWith("code").getText()).isEqualTo("code here");
+ assertThat(column.notes()).isEmpty();
+ // A marker that is runs: the item is written after it, as a rich one is.
+ Export drawn = export(true, list -> list.hangingIndent(true).marker(marker -> marker.plain("»"))
+ .items("**Java** lead"));
+ XWPFParagraph item = drawn.paragraphWith("Java");
+ assertThat(item.getText()).doesNotContain("*");
+ assertThat(item.getRuns()).filteredOn(XWPFRun::isBold).extracting(XWPFRun::text).containsExactly("Java");
+ assertThat(drawn.notes()).noneMatch(note -> note.contains("markdown"));
+ }
+
+ @Test
+ void nestedItemsAreWrittenAsThePageSetsThem() throws Exception {
+ // Without hangingIndent the page lays a nested item's indent and marker out in its text,
+ // and reads them with it; Word draws the marker, and the item's own text is written.
+ for (boolean hanging : List.of(false, true)) {
+ Export export = export(true, list -> list.hangingIndent(hanging).addItem("Languages **x**",
+ child -> child.addItem("**Java**").addItem("Kot_lin", grand -> grand.addItem("*deep* one"))));
+ assertThat(export.paragraphWith("Languages").getRuns()).extracting(XWPFRun::text)
+ .containsExactly("Languages ", "x");
+ XWPFParagraph java = export.paragraphWith("Java");
+ assertThat(java.getNumIlvl()).hasToString("1");
+ assertThat(java.getRuns()).extracting(XWPFRun::text).containsExactly("Java");
+ assertThat(java.getRuns()).extracting(XWPFRun::isBold).containsExactly(true);
+ assertThat(export.paragraphWith("Kot_lin").getText()).as("a mark the parser keeps").isEqualTo("Kot_lin");
+ XWPFParagraph deep = export.paragraphWith("deep");
+ assertThat(deep.getNumIlvl()).hasToString("2");
+ assertThat(deep.getRuns()).extracting(XWPFRun::text).containsExactly("deep", " one");
+ assertThat(deep.getRuns()).extracting(XWPFRun::isItalic).containsExactly(true, false);
+ assertThat(export.notes()).as("hangingIndent " + hanging).noneMatch(note -> note.contains("markdown"));
+ }
+ }
+
+ @Test
+ void aListWrittenAsAParagraphPerItemWritesItsMarkerAndIndentBeforeThePieces() throws Exception {
+ // A level with no marker makes the list no Word list: each item writes its marker and its
+ // nesting indent as characters, and the pieces after them.
+ Export export = export(true, list -> list.markerFor(1, ListMarker.none())
+ .addItem("**Lead** item", child -> child.addItem("**Java** sub")));
+ XWPFParagraph lead = export.paragraphWith("Lead");
+ assertThat(lead.getNumID()).isNull();
+ assertThat(lead.getRuns()).extracting(XWPFRun::text).containsExactly("• ", "Lead", " item");
+ assertThat(lead.getRuns()).extracting(XWPFRun::isBold).containsExactly(false, true, false);
+ assertThat(export.paragraphWith("Java").getRuns()).extracting(XWPFRun::text).containsExactly(" ", "Java", " sub");
+ assertThat(export.notes()).noneMatch(note -> note.contains("markdown"));
+
+ Export flat = export(true, list -> list.noMarker().continuationIndent(" ").items("**Java** lead", "Kotlin"));
+ assertThat(flat.paragraphWith("Java").getRuns()).extracting(XWPFRun::text).containsExactly("Java", " lead");
+ assertThat(flat.notes()).noneMatch(note -> note.contains("markdown"));
+ }
+
+ @Test
+ void aMarkerWrittenAsCharactersIsWrittenInTheFaceThePageSetsItIn() throws Exception {
+ // The page's parser sets the marker of a bold list's tree of items, read with the item,
+ // regular. Written as characters, the marker takes the face the page sets it in.
+ Export text = export(true, list -> list.textStyle(BOLD).markerFor(1, ListMarker.none())
+ .addItem("**Lead** item", child -> child.addItem("sub")));
+ assertThat(text.paragraphWith("Lead").getRuns()).extracting(XWPFRun::text).containsExactly("• ", "Lead", " item");
+ assertThat(text.paragraphWith("Lead").getRuns()).extracting(XWPFRun::isBold).containsExactly(false, true, false);
+ assertThat(text.notes()).noneMatch(note -> note.contains("markdown"));
+ // A flat list sets its marker before the item's text, in the list's face, as Word draws it;
+ // the text the parser keeps every mark of it sets regular, as it sets every piece.
+ Export flat = export(true, list -> list.textStyle(BOLD).items("**Lead** item", "node_js"));
+ assertThat(flat.paragraphWith("Lead").getRuns()).extracting(XWPFRun::isBold).containsExactly(true, false);
+ assertThat(flat.paragraphWith("node_js").getRuns()).extracting(XWPFRun::isBold).containsExactly(false);
+ assertThat(flat.notes()).isEmpty();
+ }
+
+ @Test
+ void aBoldWordListsTreeItemsAreWrittenInThePiecesThePageSetsThemIn() throws Exception {
+ // The page's parser sets the marker of a bold list's tree of items, read with the item,
+ // regular, and Word draws the bullets of such a list regular: the item's own text is
+ // written in its pieces, a mark the parser keeps set regular too.
+ Export word = export(true, list -> list.textStyle(BOLD).addItem("**Lead** item",
+ child -> child.addItem("Kot_lin").addItem("Plain")));
+ XWPFParagraph lead = word.paragraphWith("Lead");
+ assertThat(lead.getNumID()).isNotNull();
+ assertThat(lead.getRuns()).extracting(XWPFRun::text).containsExactly("Lead", " item");
+ assertThat(lead.getRuns()).extracting(XWPFRun::isBold).containsExactly(true, false);
+ assertThat(word.paragraphWith("Kot_lin").getRuns()).extracting(XWPFRun::isBold).containsExactly(false);
+ // An item the page does not read keeps the list's face.
+ assertThat(word.paragraphWith("Plain").getRuns()).extracting(XWPFRun::isBold).containsExactly(true);
+ assertThat(word.notes()).noneMatch(note -> note.contains("markdown"));
+ }
+
+ @Test
+ void aLeadThePageReadsAsMarkdownLeavesTheItemWrittenAsAuthoredAndNamed() throws Exception {
+ // A marker of marks, laid out in a nested item's text, is read with it: the page sets
+ // neither as the item's text after its marker.
+ Export export = export(true, list -> list.markerFor(0, ListMarker.custom("*a*"))
+ .addItem("**Java**", child -> child.addItem("sub")));
+ assertThat(export.paragraphWith("Java").getText()).isEqualTo("**Java**");
+ assertThat(export.notes()).singleElement().asString().endsWith("; " + MARKS);
+ }
+
+ @Test
+ void aHeadingInAnItemIsWrittenAndNamedWhereWordCutsIt() throws Exception {
+ Export export = export(true, list -> list.items("# Title *x*", "Kotlin", "# Other *y*"));
+ XWPFRun title = export.paragraphWith("Title").getRuns().get(0);
+ assertThat(title.text()).isEqualTo("Title *x*");
+ assertThat(title.isBold()).isTrue();
+ assertThat(title.getFontSizeAsDouble()).isEqualTo(28.0);
+ assertThat(export.notes()).singleElement().asString()
+ .startsWith("written as a Word list; its items' markdown heading is written at 28pt in a line ")
+ .endsWith("pt tall, as tall as the item's own line on the page: the page draws its letters past "
+ + "the line, and Word cuts their tops on screen");
+ }
+
+ @Test
+ void anItemThePageReadsIntoNothingIsWrittenAsItStandsAndNamed() throws Exception {
+ Export export = export(true, list -> list.items("***", "**Java**", "Kotlin"));
+ assertThat(export.document().getDocument().xmlText()).contains(">***<");
+ assertThat(export.paragraphWith("Java").getText()).isEqualTo("Java");
+ assertThat(export.notes()).containsExactly("written as a Word list; 1 of its 3 items is written as authored, "
+ + "where the page reads it as markdown and sets none of its text");
+ // With its marker in a column of its own, the item's line holds nothing at all.
+ Export column = export(true, list -> list.hangingIndent(true).items("***", "***", "Kotlin"));
+ assertThat(column.notes()).containsExactly("written as a Word list; 2 of its 3 items are written as authored, "
+ + "where the page reads them as markdown and sets none of their text");
+ }
+
+ @Test
+ void itemsNotMatchedToTheLayoutsAreWrittenAsAuthoredAndNamed() throws Exception {
+ // An item run onto the next page is laid out as a piece on each: the items no longer count
+ // as many as the layout's.
+ Export export = export(true, list -> list.items("Lead", "**Long** item that runs on and on. ".repeat(30)));
+ assertThat(export.paragraphWith("Long").getText()).startsWith("**Long**");
+ assertThat(export.notes()).singleElement().asString().endsWith("; " + MARKS);
+ // A blank item a hangingIndent list draws as a marker alone is a row the export writes none of.
+ Export blank = export(true, list -> list.hangingIndent(true).items("**Java**", "", "Kotlin"));
+ assertThat(blank.paragraphWith("Java").getText()).isEqualTo("**Java**");
+ assertThat(blank.notes()).singleElement().asString().endsWith("; " + MARKS);
+ // A marker of marks the page sets before each item's first line is not counted for the items.
+ Export marked = export(true, list -> list.marker("*").items("A", "B", "C",
+ "**Long** " + "item that runs on and on. ".repeat(40)));
+ assertThat(marked.notes()).singleElement().asString().endsWith("; " + MARKS);
+ // An item of runs keeps every mark it holds, on the page and in the count alike.
+ Export rich = export(true, list -> list.hangingIndent(true).addItem(runs -> runs.plain("x ** y ** z"))
+ .addItem("**Long** " + "item that runs on and on. ".repeat(40), child -> { }));
+ assertThat(rich.notes()).singleElement().asString().endsWith("; " + MARKS);
+ }
+
+ @Test
+ void aHangingIndentTreeWrittenAsAParagraphPerItemWritesThePiecesAfterItsMarker() throws Exception {
+ // A level with no marker makes it no Word list; its markers stand in a column of their own,
+ // and the pieces are the item's text alone.
+ Export export = export(true, list -> list.hangingIndent(true).markerFor(1, ListMarker.none())
+ .addItem("**Lead** item", child -> child.addItem("*Java* sub")));
+ XWPFParagraph lead = export.paragraphWith("Lead");
+ assertThat(lead.getNumID()).isNull();
+ assertThat(lead.getText()).doesNotContain("*").endsWith("Lead item");
+ assertThat(lead.getRuns()).filteredOn(XWPFRun::isBold).extracting(XWPFRun::text).containsExactly("Lead");
+ assertThat(export.paragraphWith("Java").getRuns()).filteredOn(XWPFRun::isItalic).extracting(XWPFRun::text)
+ .containsExactly("Java");
+ assertThat(export.notes()).noneMatch(note -> note.contains("markdown"));
+ }
+
+ @Test
+ void anItemThePageSetsAsAuthoredIsWrittenAsItStands() throws Exception {
+ Export off = export(false, list -> list.items("**Java** lead", "Kotlin"));
+ assertThat(off.paragraphWith("Java").getText()).as("markdown off").isEqualTo("**Java** lead");
+ assertThat(off.notes()).isEmpty();
+ Export kept = export(true, list -> list.items("node_js first", "Kotlin"));
+ assertThat(kept.paragraphWith("node_js").getRuns()).as("a mark the parser keeps, its white space and all")
+ .extracting(XWPFRun::text).containsExactly("node_js first");
+ assertThat(kept.notes()).isEmpty();
+ }
+
+ @Test
+ void anItemThePageSetsInOtherLettersIsWrittenAsAuthoredAndNamed() throws Exception {
+ Export export = export(true, list -> list.textStyle(DocumentTextStyle.builder().fontName(FontName.AMIRI).build())
+ .items("مرحبا **بالعالم**"));
+ assertThat(export.document().getDocument().xmlText()).contains("**");
+ assertThat(export.notes()).singleElement().asString().endsWith("; " + MARKS);
+ // The marker the page sets before the item's first line holds marks of its own, which the
+ // page keeps: its lines hold as many marks as the item, and none of them the item's.
+ Export marked = export(true, list -> list.marker("**")
+ .textStyle(DocumentTextStyle.builder().fontName(FontName.AMIRI).build()).items("مرحبا *بالعالم*"));
+ assertThat(marked.notes()).singleElement().asString().endsWith("; " + MARKS);
+ }
+
+ @Test
+ void theFacesTheItemsPiecesAreSetInTravelWithTheDocument() throws Exception {
+ Export export = export(true, list -> list.textStyle(DocumentTextStyle.builder().fontName(FontName.LATO).build())
+ .addItem("Languages", child -> child.addItem("**Java**").addItem("*Kotlin*")));
+ String table = partXml(export.document(), "/word/fontTable");
+ assertThat(table).contains(" list.textStyle(DocumentTextStyle.builder().fontName(FontName.LATO).build())
+ .items("**Java**", "Kotlin"));
+ assertThat(partXml(flat.document(), "/word/fontTable")).contains(" list.textStyle(boldLato).markerFor(1, ListMarker.none())
+ .addItem("**Lead**", child -> child.addItem("sub")));
+ assertThat(partXml(lead.document(), "/word/fontTable")).contains(" list.textStyle(boldLato).items("* Java", "* Kotlin"));
+ assertThat(partXml(typed.document(), "/word/fontTable")).contains(" paragraph.getText().contains(text))
+ .findFirst().orElseThrow(() -> new AssertionError("no paragraph holds " + text));
+ }
+
+ List notes() {
+ return report.bySubject().getOrDefault("ListNode", List.of()).stream()
+ .map(DocxExportReport.Note::detail).toList();
+ }
+ }
+
+ private static Export export(boolean markdown, Consumer list) throws Exception {
+ AtomicReference report = new AtomicReference<>();
+ byte[] docx;
+ try (DocumentSession session = GraphCompose.document().pageSize(300, 400).margin(DocumentInsets.of(30))
+ .markdown(markdown).create()) {
+ session.pageFlow(page -> page.addList(builder -> list.accept(builder.name("Skills"))));
+ docx = session.export(new DocxSemanticBackend(report::set));
+ }
+ return new Export(new XWPFDocument(new ByteArrayInputStream(docx)), report.get());
+ }
+
+ private static String partXml(XWPFDocument document, String prefix) throws Exception {
+ StringBuilder xml = new StringBuilder();
+ for (PackagePart part : document.getPackage().getParts()) {
+ if (part.getPartName().getName().startsWith(prefix)) {
+ try (InputStream in = part.getInputStream()) {
+ xml.append(new String(in.readAllBytes(), StandardCharsets.UTF_8));
+ }
+ }
+ }
+ assertThat(xml).as("the document holds " + prefix).isNotEmpty();
+ return xml.toString();
+ }
+}
diff --git a/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxListParityTest.java b/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxListParityTest.java
index 4135d4f11..962301a99 100644
--- a/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxListParityTest.java
+++ b/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxListParityTest.java
@@ -68,10 +68,15 @@ void flatItemsStripAuthorTypedMarkers() throws Exception {
@Test
void boldLeadIsNotMistakenForAMarker() throws Exception {
- List texts = exportTexts(flow -> flow
- .addList("**bold** lead stays intact"));
-
- assertThat(texts).contains("**bold** lead stays intact");
+ // A session that reads no markdown sets the item as authored, its bold lead whole.
+ try (XWPFDocument document = export(false, flow -> flow.addList("**bold** lead stays intact"))) {
+ assertThat(document.getParagraphs()).extracting(XWPFParagraph::getText)
+ .contains("**bold** lead stays intact");
+ }
+ // One that reads it sets the bold lead's text bold, its marks dropped, and the item is
+ // written so: not one of them was taken off as a typed marker.
+ assertThat(exportTexts(flow -> flow.addList("**bold** lead stays intact")))
+ .contains("bold lead stays intact");
}
@Test
@@ -137,10 +142,16 @@ private static long listItemCount(
private static XWPFDocument export(
Consumer author) throws Exception {
+ return export(true, author);
+ }
+
+ private static XWPFDocument export(boolean markdown,
+ Consumer author) throws Exception {
byte[] docxBytes;
try (DocumentSession session = GraphCompose.document()
.pageSize(595, 842)
.margin(DocumentInsets.of(36))
+ .markdown(markdown)
.create()) {
var flow = session.dsl().pageFlow().name("Flow");
author.accept(flow);
diff --git a/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxMarkdownReportTest.java b/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxMarkdownReportTest.java
index c7d1c2cdd..d473fb764 100644
--- a/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxMarkdownReportTest.java
+++ b/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxMarkdownReportTest.java
@@ -19,11 +19,11 @@
import static org.assertj.core.api.Assertions.assertThat;
/**
- * A paragraph the page reads as markdown is written as the page sets it, and not named
- * ({@link DocxSessionMarkdownTest}); where the page's lines do not tell how it sets it — with no
- * layout — it is written as authored and named. A list item the page reads as markdown is named in
- * the report: the page sets the text its marks style and drops the marks, and the Word file holds
- * the text as authored, marks and all — on the list's note.
+ * A paragraph or a list item the page reads as markdown is written as the page sets it, and not
+ * named ({@link DocxSessionMarkdownTest}, {@link DocxListMarkdownTest}); where the page's lines do
+ * not tell how it sets it — with no layout, or a list composed in a table cell — it is written as
+ * authored and named: the page sets the text its marks style and drops the marks, and the Word file
+ * holds the text as authored, marks and all — a list's items on the list's note.
*
* A session reads markdown unless it is told not to ({@code markdown(false)}), in a paragraph
* or a list item of plain text holding a mark of emphasis or code. Text the page sets as authored —
@@ -34,9 +34,6 @@ class DocxMarkdownReportTest {
private static final String UNMEASURED = "markdown marks are written as letters — whether the page reads them is "
+ "not measured";
- private static final String ITEMS = "its items' markdown marks are written as letters, where the page sets the "
- + "text they mark and drops them";
-
@Test
void aParagraphThePageReadsAsMarkdownIsWrittenSoAndNotNamed() throws Exception {
assertThat(paragraphNotes(true, page -> page.addParagraph("Some **bold** and `code` text"))).isEmpty();
@@ -79,15 +76,15 @@ void aParagraphThePageSetsAsAuthoredIsNotNamed() throws Exception {
}
@Test
- void aListsItemsThePageReadsAsMarkdownAreNamed() throws Exception {
+ void aListsItemsThePageReadsAsMarkdownAreNotNamed() throws Exception {
assertThat(listNotes(true, page -> page.addList(list -> list.name("Skills").items("**Java** lead", "Kotlin"))))
- .containsExactly("written as a Word list; " + ITEMS);
+ .isEmpty();
assertThat(listNotes(true, page -> page.addList(list -> list.name("Skills").hangingIndent(true)
.items("**Java** lead", "Kotlin"))))
- .as("markers in a column of their own").containsExactly("written as a Word list; " + ITEMS);
+ .as("markers in a column of their own").isEmpty();
assertThat(listNotes(true, page -> page.addList(list -> list.name("Skills")
.addItem("Languages", child -> child.addItem("**Java**").addItem("Kotlin")))))
- .as("a tree of items").singleElement().asString().endsWith("; " + ITEMS);
+ .as("a tree of items").noneMatch(note -> note.contains("markdown"));
// As the page lays an item out: a marker typed before it is taken off, and none is lost.
assertThat(listNotes(false, page -> page.addList(list -> list.name("Skills").items("* Java", "* Kotlin"))))
.as("a marker typed before an item, markdown off").isEmpty();
diff --git a/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxMarkdownTest.java b/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxMarkdownTest.java
index 4942a3002..275d68395 100644
--- a/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxMarkdownTest.java
+++ b/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxMarkdownTest.java
@@ -1,6 +1,8 @@
package com.demcha.compose.document.backend.semantic.docx;
+import com.demcha.compose.document.dsl.ListBuilder;
import com.demcha.compose.document.layout.payloads.ParagraphLine;
+import com.demcha.compose.document.node.ListMarker;
import com.demcha.compose.document.layout.payloads.ParagraphShapeSpan;
import com.demcha.compose.document.layout.payloads.ParagraphSpan;
import com.demcha.compose.document.layout.payloads.ParagraphTextSpan;
@@ -134,6 +136,60 @@ void thePiecesAreThePagesWhereItsLinesHoldThemSoAndNotOtherwise() {
.as("anything but text").isFalse();
}
+ @Test
+ void piecesAreSplitAtTheEndOfTheLeadTheyOpenWith() {
+ // A nested item's indent and marker, laid out in its text: no-break spaces are letters.
+ String indent = Character.toString(0x00A0).repeat(2);
+ List pieces = DocxMarkdown.read(indent + "◦ **Java** lead", BODY);
+ DocxMarkdown.Split split = DocxMarkdown.split(pieces, indent + "◦ ");
+ assertThat(split.leadStyle()).isEqualTo(BODY);
+ assertThat(split.after()).containsExactly(piece("Java", DocumentTextDecoration.BOLD, 10),
+ piece(" lead", DocumentTextDecoration.DEFAULT, 10));
+ // The lead may end inside a piece and span pieces of one style.
+ assertThat(DocxMarkdown.split(List.of(piece("- ", DocumentTextDecoration.DEFAULT, 10),
+ piece("a b", DocumentTextDecoration.DEFAULT, 10)), "- a").after())
+ .containsExactly(piece(" b", DocumentTextDecoration.DEFAULT, 10));
+ // No lead leaves the pieces whole.
+ assertThat(DocxMarkdown.split(pieces, "")).isEqualTo(new DocxMarkdown.Split(null, pieces));
+
+ assertThat(DocxMarkdown.split(DocxMarkdown.read("*a* **Java**", BODY), "*a* "))
+ .as("a lead the parser reads, its marks dropped").isNull();
+ assertThat(DocxMarkdown.split(List.of(piece("-", DocumentTextDecoration.DEFAULT, 10),
+ piece(" x", DocumentTextDecoration.BOLD, 10), piece(" y", DocumentTextDecoration.DEFAULT, 10)), "- x"))
+ .as("a lead in two styles").isNull();
+ assertThat(DocxMarkdown.split(pieces, indent + "▪ ")).as("another lead").isNull();
+ assertThat(DocxMarkdown.split(List.of(piece("◦ ab", DocumentTextDecoration.DEFAULT, 10),
+ piece("c", DocumentTextDecoration.BOLD, 10)), "▪ ")).as("another lead, in a longer piece").isNull();
+ assertThat(DocxMarkdown.split(List.of(piece("◦ ", DocumentTextDecoration.DEFAULT, 10)), "◦ "))
+ .as("nothing after the lead").isNull();
+ assertThat(DocxMarkdown.split(List.of(piece("◦", DocumentTextDecoration.DEFAULT, 10)), "◦ "))
+ .as("pieces shorter than the lead").isNull();
+ }
+
+ @Test
+ void aListsItemsAreReadAsThePageLaysThemOut() {
+ String indent = Character.toString(0x00A0).repeat(2);
+ // A flat list sets its marker before the item's first line, apart from its text, a typed
+ // marker taken off; with hangingIndent the marker stands in a column of its own.
+ assertThat(DocxMarkdown.items(new ListBuilder().items("* **Java**", "", "Kotlin").build())).containsExactly(
+ new DocxMarkdown.ItemReading("**Java**", "", "• "), new DocxMarkdown.ItemReading("Kotlin", "", "• "));
+ assertThat(DocxMarkdown.items(new ListBuilder().hangingIndent(true).items("**Java**").build()))
+ .containsExactly(new DocxMarkdown.ItemReading("**Java**", "", ""));
+ assertThat(DocxMarkdown.items(new ListBuilder().noMarker().items("**Java**").build()))
+ .containsExactly(new DocxMarkdown.ItemReading("**Java**", "", ""));
+ // A tree of items is laid out in labels, each after its depth's indent and its marker.
+ assertThat(DocxMarkdown.items(new ListBuilder().markerFor(1, ListMarker.none())
+ .addItem("**A**", child -> child.addItem("b_c", grand -> grand.addItem("*d*")))
+ .addItem(rich -> rich.bold("runs")).build())).containsExactly(
+ new DocxMarkdown.ItemReading("• **A**", "• ", ""),
+ new DocxMarkdown.ItemReading(indent + "b_c", indent, ""),
+ new DocxMarkdown.ItemReading(indent + indent + "▪ *d*", indent + indent + "▪ ", ""));
+ // With hangingIndent, its label alone, a typed marker taken off.
+ assertThat(DocxMarkdown.items(new ListBuilder().hangingIndent(true)
+ .addItem("- **A**", child -> child.addItem("b")).build())).containsExactly(
+ new DocxMarkdown.ItemReading("**A**", "", ""), new DocxMarkdown.ItemReading("b", "", ""));
+ }
+
@Test
void whatThePageMayReadAsMarkdownHoldsAMarkOfEmphasisOrCode() {
assertThat(DocxMarkdown.holdsAMark("a *b*")).isTrue();