diff --git a/CHANGELOG.md b/CHANGELOG.md index 676fc8653..d982ec536 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -8,6 +8,39 @@ follow semantic versioning; release dates are ISO 8601. ### Public API +- **A DOCX export's report names a paragraph or a list item the page reads as markdown.** A + session reads markdown unless it is told not to (`markdown(false)`). It reads a paragraph, or a + list item, of plain text holding a mark of emphasis or code: + - text its emphasis marks style is set bold or italic; + - a heading is set bold, the first three levels larger; + - the marks its parser reads as syntax are dropped, a code span's backticks and a link's + address among them. + + The DOCX export writes the text as authored, so Word showed `**bold**` with its asterisks and + none of the bold, without a note. It still writes it so; the note now says the markdown marks + are written as letters: + - the paragraph's note (`ParagraphNode`); + - a zone paragraph's, on its `page zone` note; + - a list's, for its items (`ListNode`). + + It is read from the laid-out lines, which hold fewer marks than the text where the page read + it. The text is read as the page lays it out: less a paragraph's prefix's own marks, and a list + item with a marker typed before it taken off, as the page and the file both take it off. So + text the page sets as authored is not named: + - markdown off; + - no mark; + - an underscore inside a word, which markdown keeps; + - a marker typed before a list item; + - runs, which the page never reads. + + Where the lines are not read — with no layout, or composed in a table cell, whose paragraphs are + matched to their lines by text and whose lists not at all — the note says whether the page reads + the marks is not measured, where the page's own parser drops a mark from the text. + Across the DOCX fidelity corpus the report names one paragraph: `TimelineMinimal`'s open-source + project line, whose `*(Open source)*` the page sets in italic and Word showed with its + asterisks. None of this changes what is written: the 62 documents of the corpus export to the + same bytes. In `DocxNodeFieldLedgerTest` a paragraph's `text` and a list's `nestedItems` move + from `WRITTEN` to `REPORTED`, and a list's `items` name it too. - **A DOCX export's report names where a page zone's parts stand, and what its paragraphs lose.** A page zone is written as one Word line. Word sets its parts one after another from the page's left margin, and those after the first spacer against its right margin, on one diff --git a/docs/architecture/backend-capability-matrix.md b/docs/architecture/backend-capability-matrix.md index 633e20b7e..8b5e5149a 100644 --- a/docs/architecture/backend-capability-matrix.md +++ b/docs/architecture/backend-capability-matrix.md @@ -63,8 +63,8 @@ Payload records live in `core` under | Capability (payload) | PDF (fixed) | PPTX (fixed) | DOCX (semantic) | |---|---|---|---| -| Paragraph — pre-wrapped lines, runs, alignment (`ParagraphFragmentPayload`) | ✅ `PdfParagraphFragmentRenderHandler` | ✅ `PptxParagraphFragmentRenderHandler` (one absolute, wrap-disabled frame per measured line) | ⚠️ semantic paragraphs (`DocxSemanticBackend`) — each run keeps its own style, falling back to the paragraph's when it has none; a centred or right-aligned left-to-right line of its own, of text alone and untracked, that Word sets a point or more wider or narrower at its half-point size has its letters spaced by the difference (`w:spacing`) and its room reckoned from the page's width; a `linkTarget` becomes a `w:hyperlink`, with a relationship for an address or `w:anchor` for one of the document's own anchors, and a run's own link wins over the paragraph's; a paragraph seated off its baseline (`TextVerticalAlign`) has its runs raised or lowered in the line (`w:position`) by the PDF backend's own correction (`ParagraphSeating`), one shift for the paragraph where the page seats each line by its own; Word and LibreOffice stand an exact line's baseline four fifths of the way down it whatever the face, where the page sets it the face's ascent down, so a paragraph whose face puts the two half a point or more apart — Spectral's, not Lato's — has its text moved to the page's baseline in the same position, matched at its middle line (not yet a list item's or a table text cell's; a picture among it moves with it in Word and stays on its own baseline in LibreOffice); lines a container stacks over one another tighter than their face each end halfway between their letters and the next line's (Word draws an exact line's text on screen only inside the line; its PDF export does not cut it), and the last layer of a shape container on one page, where its line runs past the foot, ends at the foot or below its letters; letters two lines share are split halfway so the page does not move, and a stack that holds a picture keeps its lines' own heights; a `bulletOffset` of spaces becomes the paragraph's indent (`w:ind` left, hanging or first line, by `indentStrategy`) in the flow and in cells, not yet over the flow, in an overlay's left-and-right pair, as a badge's initials or in a header or footer; one with letters in it is not written, its wrapped lines still set after the spaces that cover it; an auto-sized paragraph's text is written at its style's size, not the one the page fits it to; a `bookmark(...)` is Word's `HeadingN`, which Word's outline lists by the text of its Word paragraph — an overlay's pair's whole line, one level for both sides — at no level past the ninth. Outside a header or footer, the paragraph's report note (`ParagraphNode`) names each of these where it moves or renames something: the prefix's letters, and the room a path that writes no prefix leaves out where it moves a line; the size written and the size the page fits the text to, to Word's half point; an outline title that is not the text Word lists, a level past the ninth that shares it with another, and the right side's entry where the left holds the line's level | -| List hanging indent — a marker column and a content column (`ListBuilder.hangingIndent(true)`, `markerGap(...)`) | ✅ marker and content emitted as separate `ParagraphFragmentPayload` fragments at the resolved `markerX` / `contentX` | ✅ the same fragments — the fixed-layout pipeline resolves the geometry before either backend sees it | ⚠️ the top level only. `DocxSemanticBackend` exports a list as a real Word list — `numbering.xml`, `w:numPr` per item, the level carrying the marker — or, with rich items or a drawn marker, as paragraphs; content and nesting are unaffected. With the flag, the top level's marker column is the layout's — the marker's width and `markerGap`, the text and its wrapped lines where the page sets them — where the gap covers what Word may set the marker wider: a picture at its written size, its edges included, or text in the page's face (embedded, or a standard one Word sets in the same widths) grown to its half-point size, half a point clear. A Word list's level then indents and hangs by that column; a list of paragraphs writes the marker, a tab to a stop there, and hangs the item there. Word places content at absolute indents and has no relative-advance primitive, so without the layout's measure the gap could not be honoured; a Word list without the flag that the layout placed and that does not nest takes the page's column too, the spaces the page sets its wrapped lines after, its marker followed by a space (`w:suff`) and an item that wraps measured at Word's half-point size; a list that nests items, a list built as a tree of items (laid out flattened), and a marker the gap does not clear keep the stated column (180 twips, plus 120 per nesting level) — except, in a list of paragraphs, a nested rich item with no marker, which stands where the layout set its text, its measure weighed at Word's half-point sizes, where the layout's items are matched to the list's; a list that nests only such items sets its top level at the page's column too. The report counts, on the list, the items that stand at a stated column, a space past their marker or two spaces a level in, and names a centred or right-aligned list written flush left, a lineSpacing not written where the layout's items are not the list's own and one wraps (in a list composed in a table cell, its wrapping not measured), a continuationIndent not written where an item of a markerless list or a tree of items without the flag wraps or its wrapping is not measured, and the rows the page draws as a marker alone for blank items of a flagged list, which are not written | +| Paragraph — pre-wrapped lines, runs, alignment (`ParagraphFragmentPayload`) | ✅ `PdfParagraphFragmentRenderHandler` | ✅ `PptxParagraphFragmentRenderHandler` (one absolute, wrap-disabled frame per measured line) | ⚠️ semantic paragraphs (`DocxSemanticBackend`) — each run keeps its own style, falling back to the paragraph's when it has none; a centred or right-aligned left-to-right line of its own, of text alone and untracked, that Word sets a point or more wider or narrower at its half-point size has its letters spaced by the difference (`w:spacing`) and its room reckoned from the page's width; a `linkTarget` becomes a `w:hyperlink`, with a relationship for an address or `w:anchor` for one of the document's own anchors, and a run's own link wins over the paragraph's; a paragraph seated off its baseline (`TextVerticalAlign`) has its runs raised or lowered in the line (`w:position`) by the PDF backend's own correction (`ParagraphSeating`), one shift for the paragraph where the page seats each line by its own; Word and LibreOffice stand an exact line's baseline four fifths of the way down it whatever the face, where the page sets it the face's ascent down, so a paragraph whose face puts the two half a point or more apart — Spectral's, not Lato's — has its text moved to the page's baseline in the same position, matched at its middle line (not yet a list item's or a table text cell's; a picture among it moves with it in Word and stays on its own baseline in LibreOffice); lines a container stacks over one another tighter than their face each end halfway between their letters and the next line's (Word draws an exact line's text on screen only inside the line; its PDF export does not cut it), and the last layer of a shape container on one page, where its line runs past the foot, ends at the foot or below its letters; letters two lines share are split halfway so the page does not move, and a stack that holds a picture keeps its lines' own heights; a `bulletOffset` of spaces becomes the paragraph's indent (`w:ind` left, hanging or first line, by `indentStrategy`) in the flow and in cells, not yet over the flow, in an overlay's left-and-right pair, as a badge's initials or in a header or footer; one with letters in it is not written, its wrapped lines still set after the spaces that cover it; an auto-sized paragraph's text is written at its style's size, not the one the page fits it to; a paragraph a session reads as markdown — the default, unless `markdown(false)` — is written as authored, its marks as letters and none of their style, not yet as the page sets it; a `bookmark(...)` is Word's `HeadingN`, which Word's outline lists by the text of its Word paragraph — an overlay's pair's whole line, one level for both sides — at no level past the ninth. Outside a header or footer, the paragraph's report note (`ParagraphNode`) names each of these where it moves or renames something: the prefix's letters, and the room a path that writes no prefix leaves out where it moves a line; the size written and the size the page fits the text to, to Word's half point; the marks of a paragraph the page read as markdown, where its laid-out lines hold fewer of them than its text, not measured where its lines are not read; an outline title that is not the text Word lists, a level past the ninth that shares it with another, and the right side's entry where the left holds the line's level | +| List hanging indent — a marker column and a content column (`ListBuilder.hangingIndent(true)`, `markerGap(...)`) | ✅ marker and content emitted as separate `ParagraphFragmentPayload` fragments at the resolved `markerX` / `contentX` | ✅ the same fragments — the fixed-layout pipeline resolves the geometry before either backend sees it | ⚠️ the top level only. `DocxSemanticBackend` exports a list as a real Word list — `numbering.xml`, `w:numPr` per item, the level carrying the marker — or, with rich items or a drawn marker, as paragraphs; content and nesting are unaffected. With the flag, the top level's marker column is the layout's — the marker's width and `markerGap`, the text and its wrapped lines where the page sets them — where the gap covers what Word may set the marker wider: a picture at its written size, its edges included, or text in the page's face (embedded, or a standard one Word sets in the same widths) grown to its half-point size, half a point clear. A Word list's level then indents and hangs by that column; a list of paragraphs writes the marker, a tab to a stop there, and hangs the item there. Word places content at absolute indents and has no relative-advance primitive, so without the layout's measure the gap could not be honoured; a Word list without the flag that the layout placed and that does not nest takes the page's column too, the spaces the page sets its wrapped lines after, its marker followed by a space (`w:suff`) and an item that wraps measured at Word's half-point size; a list that nests items, a list built as a tree of items (laid out flattened), and a marker the gap does not clear keep the stated column (180 twips, plus 120 per nesting level) — except, in a list of paragraphs, a nested rich item with no marker, which stands where the layout set its text, its measure weighed at Word's half-point sizes, where the layout's items are matched to the list's; a list that nests only such items sets its top level at the page's column too. The report counts, on the list, the items that stand at a stated column, a space past their marker or two spaces a level in, and names a centred or right-aligned list written flush left, a lineSpacing not written where the layout's items are not the list's own and one wraps (in a list composed in a table cell, its wrapping not measured), a continuationIndent not written where an item of a markerless list or a tree of items without the flag wraps or its wrapping is not measured, the rows the page draws as a marker alone for blank items of a flagged list, which are not written, and items the page reads as markdown, whose marks are written as letters | | Inline code/badge chips (`InlineBackground` on text spans) | ✅ `PdfParagraphFragmentRenderHandler` | ✅ `PptxParagraphFragmentRenderHandler` | ⚠️ `DocxSemanticBackend` — the fill becomes the run's own `w:shd`, in a paragraph and in a list item alike, so a badge still reads as a badge. What Word has no way to say is the shape: shading covers the glyph box, so the corner radius and the padding above and below the letters are not in the file, and the export records them. The padding beside the letters is written as the room it takes (`spaceAfterTheLastLetter`): character spacing after the chip's last letter, shaded with it, and after the letter before the chip, unshaded. A chip opening its line or following a picture has no letter before it, so its left padding is not in the file; no space is written after right-to-left letters or after a symbol or emoji. The export records, chip by chip, how each side was written. LibreOffice sets no spacing after a line's last letter, so it does not apply the right padding of a chip that ends a line. A `w:shd` fill is opaque, so a translucent chip is flattened first against what this export wrote underneath it — the paragraph's shading, the cell's, or the page — so the chip agrees with the file it is in, which on a white page is the colour the PDF shows. It stops being translucent, and that is recorded with the rest | | Inline images (`ParagraphImageSpan`) | ✅ `PdfParagraphFragmentRenderHandler` | ✅ `PptxParagraphFragmentRenderHandler` | ✅ `DocxSemanticBackend.writeInlinePicture` (a picture in its own run where it sits among the words, at its size, inside the run's or the paragraph's link; raised or lowered by `w:position` to where the page's alignment and `baselineOffset` put it, from the layout's measure of the paragraph's first line — in a list, the list's text on a line as tall as the item's own tallest picture; LibreOffice ignores `w:position` on a picture and stands it on the baseline, so a picture the export draws itself (icon, emoji, shape) that the page raises carries the rise as transparent rows and needs no `w:position`, while one the page lowers stands in LibreOffice higher than on the page by as much as the page lowers it — up to the text's descent for a centred icon as tall as its line; the editor clips a picture to an exact line height, so a paragraph holding a picture that leaves its text — past the ascent or the descent, in Word's placement or on the baseline — has its lines written at least the height the picture reaches, grown by the editor rather than clipped, every line of the paragraph since Word has one line height for it, and each as tall as the editor's font makes it — for 14pt text about 2.5pt taller than the page's in LibreOffice; a picture inside its text in both editors keeps the exact height; a paragraph of one line of text in a Word paragraph of its own, with room above for its pictures' reach, keeps an exact line at the page's height of it, the pictures set in it where the page puts them in Word and what their ink reaches past it taken from the gaps around it, and in LibreOffice a lowered picture there stands higher and loses what passes the line's top; its description is the text it stands for or empty) | | Inline vector shapes (`ParagraphShapeSpan`) | ✅ `PdfParagraphFragmentRenderHandler` | ⚠️ `PptxParagraphFragmentRenderHandler` + `PptxInlineGeometry` (distinct per-corner radii render with the top-left radius — single-adjust preset) | ⚠️ `DocxSemanticBackend.writeInlinePicture` + `DocxShapePictures` (a transparent PNG drawn by the shared `InlineSvgRasters` from the outline, fill and stroke — every outline kind, each layer centred in the run's box — placed as an inline picture is; the picture takes as far as the stroked ink reaches past the outline — half the stroke on an edge, more at a sharp corner's miter — and a pixel on each side, measured side by side, and is lowered by what it takes below, so no edge is cut and a shape takes that much more room in the line; a list marker that draws a disc is its picture; at the top level of a `hangingIndent(true)` list that does not nest it is followed by a tab to where the layout starts the item's text, the item's lines hanging there, when the picture clears that stop, else by a space) | diff --git a/docs/recipes/docx-export.md b/docs/recipes/docx-export.md index ac3c48b0f..0a333a590 100644 --- a/docs/recipes/docx-export.md +++ b/docs/recipes/docx-export.md @@ -67,7 +67,7 @@ creation date is real metadata. | Document node | DOCX output | |---|---| -| Paragraphs | Word paragraphs with alignment, font, size, colour, bold/italic/underline; inline runs preserved; a `\n` in the text is a line break (`w:br`), which Word would otherwise read as a space; a table cell's `text("a\nb")` stays one line, as the page sets it; a `bulletOffset` of spaces is an indent — the wrapped lines (`FROM_SECOND_LINE`), the first (`FIRST_LINE`) or all of them start that far in, measured in the paragraph's style as the page measures it; a prefix with letters in it is not written, but the wrapped lines still start after the spaces the page covers it with; no prefix is applied to a paragraph written over the flow, as one of an overlay's left-and-right pair, as a badge's initials, or in a header or footer; the report names a prefix's letters, and the room a prefix sets lines in by where a path that writes none leaves it out and it moves a line — on the paragraph, or in a header or footer on the zone's note. An auto-sized paragraph (`autoSize`) is written at its style's size, not the one the page fits its text to, and the report names both sizes where Word, to its half point, holds them apart, on the paragraph or, in a header or footer, on the zone's note. A body paragraph the layout moves to a new page keeps the space the layout leaves above its text there — its own top edge, the edges of the containers opening with it, and the gap before it where the gap did not fit at the foot of the page above — as an empty line that tall before it, kept with it, since Word drops a paragraph's space above at the top of a page and keeps a line's height; the rest of the space owed stays above that line, so Word breaks the page where it did, and the gap between the paragraph's lines comes off the line as it would off its space above. A paragraph in a table cell, an overlay or a list is left to its container | +| Paragraphs | Word paragraphs with alignment, font, size, colour, bold/italic/underline; inline runs preserved; a `\n` in the text is a line break (`w:br`), which Word would otherwise read as a space; a table cell's `text("a\nb")` stays one line, as the page sets it; a `bulletOffset` of spaces is an indent — the wrapped lines (`FROM_SECOND_LINE`), the first (`FIRST_LINE`) or all of them start that far in, measured in the paragraph's style as the page measures it; a prefix with letters in it is not written, but the wrapped lines still start after the spaces the page covers it with; no prefix is applied to a paragraph written over the flow, as one of an overlay's left-and-right pair, as a badge's initials, or in a header or footer; the report names a prefix's letters, and the room a prefix sets lines in by where a path that writes none leaves it out and it moves a line — on the paragraph, or in a header or footer on the zone's note. A session reads markdown unless told not to (`markdown(false)`): it sets a paragraph of plain text's emphasis marks as the style of the text they mark, and drops its marks; the export writes the text as authored, marks and all, which it does not yet write as the page sets it, and the report names it — and, where the paragraph's lines are not read, composed in a table cell whose text the page set otherwise or with no layout, says whether the page reads its marks is not measured, where the page's parser drops one. An auto-sized paragraph (`autoSize`) is written at its style's size, not the one the page fits its text to, and the report names both sizes where Word, to its half point, holds them apart, on the paragraph or, in a header or footer, on the zone's note. A body paragraph the layout moves to a new page keeps the space the layout leaves above its text there — its own top edge, the edges of the containers opening with it, and the gap before it where the gap did not fit at the foot of the page above — as an empty line that tall before it, kept with it, since Word drops a paragraph's space above at the top of a page and keeps a line's height; the rest of the space owed stays above that line, so Word breaks the page where it did, and the gap between the paragraph's lines comes off the line as it would off its space above. A paragraph in a table cell, an overlay or a list is left to its container | | Lists | Real Word lists: a `numbering.xml` definition per list, `w:numPr` on each item, and the authored marker as the level's text. Nesting is a list level, so Enter continues the list and Tab demotes an item. See "What a list becomes" below for the kinds that stay plain paragraphs | | Tables | Word tables, one cell per cell. Each cell states its own padding, on all four sides, so a row is as tall as the page draws it: as `w:tcMar`, and above and below partly in its paragraphs. Word and LibreOffice give every cell of a row the largest top and bottom margin of any cell in it, so a row's cells are written with its smallest, and the rest of a cell's padding above and below is space above its first paragraph and below its last (measured: a row whose day cells were padded 5.5pt above and 10.25pt below beside a label padded 0.75pt stood 60.3pt tall in both editors, where its tallest cell came to 46). A cell opening with a table has no paragraph above it to hold its padding, and a cell in a vertical merge has its bottom edge in another row: these keep their margins, and the row's comes down no lower than the largest of them. Its padding above and below gives up the room Word makes for the table's horizontal rules — half of a rule between two rows to each, the lower row's rule where the two differ, the rule above the table and the one below it whole to their row — which the page does not (measured: a 0.75pt rule made each row 0.75pt taller). A row held at the page's height is written less its margins and those rules too — a rule and a half in the first row and the last, two in a table of one row — since both editors read a row's written height as its cells' content (measured: held less one rule, a table ruled at 0.75pt stood 0.46pt taller in its first row and 0.36pt in its last). A cell that holds nothing but an empty line — a row that is only a rule, its thickness the empty cell's font — has that line cut to the room its row leaves it, the page's row less the cell's own margins and border: the page draws the rule's borders across the line, and Word and LibreOffice keep them outside it and grow the row (measured: `CobaltRota`'s two rules under a 0.9pt border stood 0.9pt taller each). A line with letters, a picture or a paragraph border of its own keeps its height. A table that states no rule is written with the engine's default 1pt black rule, as the page draws it, not left on Word's thinner grid. Its `textAnchor` becomes `w:vAlign` and the paragraph's `w:jc`, with the engine's default — the vertical middle, on the left — where Word's is the top, so a line beside a taller neighbour sits where the page puts it and an amount column stays right-aligned. A cell with no style of its own is set in the engine's default cell face rather than the document's Normal. A cell's lines are one paragraph with line breaks; when its style's `lineSpacing(...)` is above zero and it has more than one line, they are a paragraph each, with the spacing after every line but the last, since Word has no space between the lines of one paragraph but a taller line. A column sized to its content gets a point more than the page gives it, so the editor's font substitute cannot wrap its widest cell. The width is written when the document states one or every column is fixed; otherwise Word sizes the table — see "What falls back". A table breaks across pages where the layout breaks it: every row the layout placed is kept whole (`w:cantSplit`), `repeatHeader(n)` rows repeat on each page (`w:tblHeader`) and stay with the row under them. Two tables in a row — rows included, since a row is carried as a table — are kept apart by a paragraph a tenth of a point tall, holding the rest of the gap between them: an editor joins two tables with nothing between them into one. A table or a row the layout moves to a new page keeps its own top edge there, as the page does — written as a line that tall, kept with it, since Word drops a paragraph's space above at the top of a page — while the gap between it and the block before stays at the foot of the page above, where it fits there; a gap the layout carries onto the new page, because it did not fit at the foot of the page above, is not yet held above a table (body paragraphs and spacers: see their rows). A table's margin is its indent and the space round it, and its padding on the sides holds its rows in as the page draws them: its left side is in the indent too, and both are out of the room its columns are given | | Composed cells (`DocumentTableCell.node(...)`) | Written by the same writers that write that node anywhere else, so a cell built from an image, a list or a table carries it. A nested table is a real `w:tbl` followed by the paragraph Word requires a cell to end with — a hairline, which the paragraph written next in the cell takes over, so no empty line opens under the table, and whose mark is hidden where it is left at the cell's end holding nothing and no space, since LibreOffice lays it out — and takes the width of the column it sits in, less its own margins and padding — the column's, not the one the page gives it, because the layout reports a composed cell's content under the owner's path | @@ -457,8 +457,9 @@ a line apiece — nor in a list composed in a table cell, whose wrapping is not continuationIndent is not written in a list the page sets it in, a markerless list or a tree of items without `hangingIndent`, where an item wraps or its wrapping is not measured; it counts the items that stand at a stated column, a space past their marker or two spaces a level in, -rather than where the page sets them; and it names the rows a `hangingIndent` list draws as a -marker alone for blank items, which the export does not write. +rather than where the page sets them; it names the rows a `hangingIndent` list draws as a +marker alone for blank items, which the export does not write; and it names items of plain text +the page reads as markdown, whose marks are written as letters, as a paragraph's are. ## What a panel keeps and loses diff --git a/render-docx/README.md b/render-docx/README.md index c357a7b8e..b460e530d 100644 --- a/render-docx/README.md +++ b/render-docx/README.md @@ -121,11 +121,14 @@ What is not written — each one is named in the export report - its `continuationIndent`; - the marker column and `markerGap` of an item at a stated column — a list that nests, or a gap too narrow for its marker — while a flat hanging-indent list keeps the page's column; - - a row the page draws as a marker alone, for a blank item. + - a row the page draws as a marker alone, for a blank item; + - the marks of items the page reads as markdown, written as letters. - **What a paragraph's own fields set where Word cannot hold it**, named in the export report on the paragraph outside a header or footer (a page zone's are named on the zone): - the size an auto-sized paragraph's text is fitted to, where Word, to its half point, holds it apart from its style's; + - the marks of a paragraph the session reads as markdown (the default; `markdown(false)` turns + it off), written as letters with none of their style; - the letters of a `bulletOffset` prefix; and, where it moves a line, the room a prefix sets lines in by in a paragraph written over the flow, as a side of an overlay's left-and-right pair or as a badge's initials; diff --git a/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxSemanticBackend.java b/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxSemanticBackend.java index ca904947b..5e30cf301 100644 --- a/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxSemanticBackend.java +++ b/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxSemanticBackend.java @@ -1890,7 +1890,7 @@ private void reportZoneLine(int zoneIndex, boolean header, DocumentNode content) /** * What a paragraph of a page zone loses of its own on the zone's line: its direction, its - * prefix's letters, the size its text is fitted to, and its outline entry. + * prefix's letters, the size its text is fitted to, its markdown marks, and its outline entry. * * @param lines the lines the page laid it out in, empty where they are not read */ @@ -1907,6 +1907,10 @@ private List zoneParagraphLost(ParagraphNode node, List + * The rows the page draws as a marker alone, which the export does not write + * ({@link #markerOnlyRows}). And the marks of items the page reads as markdown, which are + * written as letters ({@link #itemsMarkdownLost}).

* * @param laidOut the items as the layout laid them out * @param atTheirColumn how many of its items stand where the page sets them @@ -3725,9 +3730,63 @@ private List listLost(com.demcha.compose.document.node.ListNode list, lost.add(markerRows + (markerRows == 1 ? " row the page draws as a marker alone, for a blank item, is" : " rows the page draws as a marker alone, for blank items, are") + " not written"); } + String marks = itemsMarkdownLost(list); + if (marks != null) { + lost.add(marks); + } return lost; } + /** + * What a list's items lose where the page reads them as markdown, as a paragraph's text is read + * ({@link #markdownLost}): an item of plain text holding a mark of emphasis or code is set as + * markdown, its marks dropped, and written as authored. Its text is read as the page lays it + * out, a marker typed before it taken off + * ({@link com.demcha.compose.document.node.ListMarker#normalizeItemText}); its laid-out lines + * are the items', without a marker laid out on its own. A marker laid out in an item's line, + * and a rich item's runs, hold every mark they had. + */ + private String itemsMarkdownLost(com.demcha.compose.document.node.ListNode list) { + List plain = new ArrayList<>(); + List all = new ArrayList<>(); + if (list.nestedItems().isEmpty()) { + for (String item : list.items()) { + String text = com.demcha.compose.document.node.ListMarker.normalizeItemText(item, list.normalizeMarkers()); + plain.add(text); + all.add(text); + } + } else { + collectItemTexts(list.nestedItems(), list.normalizeMarkers(), plain, all); + } + if (plain.stream().noneMatch(DocxSemanticBackend::holdsAMarkdownTrigger)) { + return null; + } + List lines = new ArrayList<>(); + for (DocxLayoutMetrics.ItemText item : layout.itemLines(list)) { + lines.addAll(item.lines()); + } + return marksDropped("its items'", lines, all.stream().mapToInt(DocxSemanticBackend::markdownMarksIn).sum(), 0, + plain); + } + + /** + * Every item's text in a tree of items, as the page lays it out: a plain item's label, a marker + * typed before it taken off, and a rich item's runs' text too. + */ + private static void collectItemTexts(List items, boolean normalizeMarkers, + List plain, List all) { + for (com.demcha.compose.document.node.ListItem item : items) { + if (item.runs().isEmpty()) { + String text = com.demcha.compose.document.node.ListMarker.normalizeItemText(item.label(), normalizeMarkers); + plain.add(text); + all.add(text); + } else { + all.add(InlineRun.plainText(item.runs())); + } + collectItemTexts(item.children(), normalizeMarkers, plain, all); + } + } + /** What the items a list writes lose (see {@link #listLost}). */ private List itemsLost(com.demcha.compose.document.node.ListNode list, List laidOut, int items, int markerRows, @@ -7000,11 +7059,12 @@ private void writeParagraph(XWPFDocument document, ParagraphNode node) { /** * What a paragraph's own fields lose on the way to Word, on any path that writes it in the - * body: the size an auto-sized paragraph's text is fitted to, its {@code bulletOffset} and its - * outline entry. A page zone's paragraphs are written apart, and named on the zone's note - * ({@link #zoneParagraphLost}). + * body: the size an auto-sized paragraph's text is fitted to, the marks of a paragraph the page + * reads as markdown, its {@code bulletOffset} and its outline entry. A page zone's paragraphs + * are written apart, and named on the zone's note ({@link #zoneParagraphLost}). * - *

Its text is written at its style's size, not the one the page fits it to. A prefix's + *

Its text is written at its style's size, not the one the page fits it to, and as authored, + * its markdown marks and all ({@link #markdownLost}). A prefix's * letters are never written. A path that does not write a prefix as a distance * ({@link #indentAsThePrefixDoes}) may leave out the room it sets lines in by, too; which * path does is the caller's to say. Word's outline lists a heading by the text of its Word @@ -7023,6 +7083,10 @@ private List paragraphLost(ParagraphNode node, boolean roomLost, String if (size != null) { lost.add(size); } + String marks = readsAsMarkdown(node) ? markdownLost(node, layout.lines(node), "its") : null; + if (marks != null) { + lost.add(marks); + } String prefix = node.bulletOffset(); if (setsAPrefixBeforeTheFirstLine(node) && !prefix.isBlank()) { lost.add("its bulletOffset's letters, \"" + prefix.strip() + "\", are not written before its first line"); @@ -7087,6 +7151,108 @@ private static String autoSizeLost(ParagraphNode node, List= 0 || text.indexOf('_') >= 0 || text.indexOf('`') >= 0; + } + + /** + * Whether the page may read a paragraph as markdown, where its session asks it to: plain text, + * with no runs, holding a mark it reads markdown on ({@link #holdsAMarkdownTrigger}). Any other + * paragraph is laid out with every mark it holds, so its lines need not be read. + */ + private static boolean readsAsMarkdown(ParagraphNode node) { + return node.inlineRuns().isEmpty() && holdsAMarkdownTrigger(node.text()); + } + + /** + * What a paragraph read as markdown loses: the page sets the text its marks style, and drops + * the marks, where its session reads markdown; the file writes the text as authored, marks and + * all, and none of their style. Read off the laid-out lines ({@link #marksDropped}), less the + * marks of a prefix the page sets before the first line. + * + * @param lines the lines the page laid the paragraph out in, empty where they are not read + * @param whose whose text the phrase names + */ + private static String markdownLost(ParagraphNode node, List lines, + String whose) { + if (!readsAsMarkdown(node)) { + return null; + } + return marksDropped(whose, lines, markdownMarksIn(node.text()), + setsAPrefixBeforeTheFirstLine(node) ? markdownMarksIn(node.bulletOffset()) : 0, List.of(node.text())); + } + + /** + * Whether the page dropped markdown marks from text it laid out, and the phrase that names it: + * where its lines, less a prefix's marks, hold fewer than the text did; {@code null} where they + * hold as many, or hold no text at all, as the lines of text given no width do. Where they are + * not read — with no layout, or composed in a table cell whose text the page set otherwise than + * authored — a session reads markdown unless told not to, and the phrase says whether the page + * read them is not measured, where the page's parser drops a mark from the text. + * + * @param whose whose text the phrase names + * @param lines the lines the page laid the text out in, empty where they are not read + * @param authored how many marks the text holds as the page lays it out + * @param prefixMarks how many marks a prefix the page sets before the first line holds + * @param texts the text the page may read as markdown + */ + private static String marksDropped(String whose, List lines, + int authored, int prefixMarks, List texts) { + if (lines.isEmpty()) { + return texts.stream().anyMatch(DocxSemanticBackend::parserDropsAMark) + ? whose + " markdown marks are written as letters — whether the page reads them is not measured" + : null; + } + if (lines.stream().allMatch(line -> line.text().isBlank())) { + return null; + } + int laidOut = -prefixMarks; + for (com.demcha.compose.document.layout.payloads.ParagraphLine line : lines) { + laidOut += markdownMarksIn(line.text()); + } + return laidOut < authored + ? whose + " markdown marks are written as letters, where the page sets the text they mark and drops them" + : null; + } + + /** + * Whether the page's markdown parser drops a mark from text, reading it as the page does where + * its session reads markdown: an underscore inside a word, which it keeps, drops none. + */ + private static boolean parserDropsAMark(String text) { + if (!holdsAMarkdownTrigger(text)) { + return false; + } + StringBuilder parsed = new StringBuilder(); + for (com.demcha.compose.engine.components.content.text.TextDataBody body + : new com.demcha.compose.engine.text.markdown.MarkDownParser().getBody(text, + com.demcha.compose.engine.components.content.text.TextStyle.DEFAULT_STYLE)) { + parsed.append(body.text()); + } + return markdownMarksIn(parsed.toString()) < markdownMarksIn(text); + } + + private static int markdownMarksIn(String text) { + int marks = 0; + for (int index = 0; index < text.length(); index++) { + if (MARKDOWN_MARKS.indexOf(text.charAt(index)) >= 0) { + marks++; + } + } + return marks; + } + /** Whether any of a paragraph's text takes the paragraph's style: its plain text, or a run with no style of its own. */ private static boolean holdsTextInTheParagraphsStyle(ParagraphNode node) { if (node.inlineRuns().isEmpty()) { diff --git a/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxMarkdownReportTest.java b/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxMarkdownReportTest.java new file mode 100644 index 000000000..9de5f2cd3 --- /dev/null +++ b/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxMarkdownReportTest.java @@ -0,0 +1,169 @@ +package com.demcha.compose.document.backend.semantic.docx; + +import com.demcha.compose.GraphCompose; +import com.demcha.compose.document.api.DocumentSession; +import com.demcha.compose.document.dsl.PageFlowBuilder; +import com.demcha.compose.document.dsl.ParagraphBuilder; +import com.demcha.compose.document.output.DocumentPageZone; +import com.demcha.compose.document.style.DocumentInsets; +import com.demcha.compose.document.style.DocumentTextIndent; +import com.demcha.compose.document.style.DocumentTextStyle; +import com.demcha.compose.document.table.DocumentTableCell; +import com.demcha.compose.document.table.DocumentTableColumn; +import org.junit.jupiter.api.Test; + +import java.util.List; +import java.util.concurrent.atomic.AtomicReference; +import java.util.function.Consumer; + +import static org.assertj.core.api.Assertions.assertThat; + +/** + * A paragraph or a list item the page reads as markdown is named in the report: the page sets the + * text its marks style and drops the marks, and the Word file holds the text as authored, marks and + * all — on the paragraph's note, a zone paragraph's on the zone's, a list's items' on the list's. + * + *

A session reads markdown unless it is told not to ({@code markdown(false)}), in a paragraph + * or a list item of plain text holding a mark of emphasis or code. Text the page sets as authored — + * markdown off, no mark, a mark it keeps, a marker typed before an item, or runs, which it never + * reads — is not named.

+ */ +class DocxMarkdownReportTest { + + private static final String MARKS = "its markdown marks are written as letters, where the page sets the text " + + "they mark and drops them"; + private static final String UNMEASURED = "markdown marks are written as letters — whether the page reads them is " + + "not measured"; + private static final String ITEMS = "its items' markdown marks are written as letters, where the page sets the " + + "text they mark and drops them"; + + @Test + void aParagraphThePageReadsAsMarkdownIsNamed() throws Exception { + assertThat(paragraphNotes(true, page -> page.addParagraph("Some **bold** and `code` text"))) + .containsExactly("written as a paragraph; " + MARKS); + assertThat(paragraphNotes(true, page -> page.addParagraph("# Title *now*"))).as("a heading") + .containsExactly("written as a paragraph; " + MARKS); + assertThat(paragraphNotes(true, page -> page.addParagraph("`x`"))).as("a code span alone") + .containsExactly("written as a paragraph; " + MARKS); + // A prefix with a mark of its own is laid out with it; the paragraph's marks still go. + assertThat(paragraphNotes(true, page -> page.addParagraph(p -> p.text("Some *emphasis* here") + .bulletOffset("* ").indentStrategy(DocumentTextIndent.FIRST_LINE)))) + .containsExactly("written as a paragraph; " + MARKS + "; its bulletOffset's letters, \"*\", are not " + + "written before its first line"); + assertThat(paragraphNotes(true, page -> page.addParagraph(p -> p.text("*x*") + .bulletOffset("** ").indentStrategy(DocumentTextIndent.FIRST_LINE)))) + .as("as many marks in the prefix as the text drops") + .containsExactly("written as a paragraph; " + MARKS + "; its bulletOffset's letters, \"**\", are not " + + "written before its first line"); + } + + @Test + void aParagraphThePageSetsAsAuthoredIsNotNamed() throws Exception { + assertThat(paragraphNotes(false, page -> page.addParagraph("Some **bold** and `code` text"))) + .as("markdown off").isEmpty(); + assertThat(paragraphNotes(true, page -> page.addParagraph("Plain text, no marks"))) + .as("no mark").isEmpty(); + assertThat(paragraphNotes(true, page -> page.addParagraph("Read the file_name field"))) + .as("an underscore inside a word, which markdown keeps").isEmpty(); + assertThat(paragraphNotes(true, page -> page.addParagraph(p -> p.inlineText("Some **bold** text")))) + .as("runs, which the page never reads as markdown").isEmpty(); + assertThat(paragraphNotes(false, page -> page.addParagraph(p -> p.text("Some *emphasis* here") + .padding(new DocumentInsets(0, 0, 0, 240))))) + .as("given no width, the page lays out lines with no text").isEmpty(); + assertThat(paragraphNotes(false, page -> page.addParagraph(p -> p.text("Some *emphasis* here") + .bulletOffset("* ").indentStrategy(DocumentTextIndent.FIRST_LINE)))) + .as("markdown off, the prefix's mark aside") + .containsExactly("written as a paragraph; its bulletOffset's letters, \"*\", are not written before its " + + "first line"); + } + + @Test + void aListsItemsThePageReadsAsMarkdownAreNamed() throws Exception { + assertThat(listNotes(true, page -> page.addList(list -> list.name("Skills").items("**Java** lead", "Kotlin")))) + .containsExactly("written as a Word list; " + ITEMS); + assertThat(listNotes(true, page -> page.addList(list -> list.name("Skills").hangingIndent(true) + .items("**Java** lead", "Kotlin")))) + .as("markers in a column of their own").containsExactly("written as a Word list; " + ITEMS); + assertThat(listNotes(true, page -> page.addList(list -> list.name("Skills") + .addItem("Languages", child -> child.addItem("**Java**").addItem("Kotlin"))))) + .as("a tree of items").singleElement().asString().endsWith("; " + ITEMS); + // As the page lays an item out: a marker typed before it is taken off, and none is lost. + assertThat(listNotes(false, page -> page.addList(list -> list.name("Skills").items("* Java", "* Kotlin")))) + .as("a marker typed before an item, markdown off").isEmpty(); + assertThat(listNotes(true, page -> page.addList(list -> list.name("Skills").items("* Java", "* Kotlin")))) + .as("a marker typed before an item").isEmpty(); + assertThat(listNotes(false, page -> page.addList(list -> list.name("Skills").items("**Java** lead", "Kotlin")))) + .as("markdown off").isEmpty(); + assertThat(listNotes(true, page -> page.addList(list -> list.name("Skills").items("Java", "Kotlin")))) + .as("no mark").isEmpty(); + } + + @Test + void whereTheLinesAreNotReadWhetherThePageReadsTheMarksIsNotMeasured() throws Exception { + // With no layout, nothing tells; a session reads markdown unless told not to. + assertThat(DocxExports.reportWithoutLayout(300, 400, 30, page -> page.addParagraph("Some **bold** text")) + .bySubject().get("ParagraphNode")).extracting(DocxExportReport.Note::detail) + .containsExactly("written as a paragraph; its " + UNMEASURED); + assertThat(DocxExports.reportWithoutLayout(300, 400, 30, page -> page.addParagraph("Plain text")) + .bySubject()).as("no mark").doesNotContainKey("ParagraphNode"); + // Composed in a table cell, a paragraph is matched to its lines by its text, which the page + // set otherwise than authored; a list's lines are not matched at all. + assertThat(paragraphNotes(true, page -> page.add(cell(new ParagraphBuilder().name("Note") + .text("Some **bold** text").build())))) + .containsExactly("written as a paragraph; its " + UNMEASURED); + assertThat(paragraphNotes(true, page -> page.add(cell(new ParagraphBuilder().name("Note") + .text("Install node_js first").build())))).as("a mark the parser keeps").isEmpty(); + assertThat(listNotes(true, page -> page.add(cell(new com.demcha.compose.document.dsl.ListBuilder() + .name("Skills").items("**Java** lead", "Kotlin").build())))) + .singleElement().asString().endsWith("; its items' " + UNMEASURED); + assertThat(listNotes(true, page -> page.add(cell(new com.demcha.compose.document.dsl.ListBuilder() + .name("Skills").items("node_js", "Kotlin").build())))).as("a mark the parser keeps") + .noneMatch(note -> note.contains("markdown")); + } + + @Test + void aZoneParagraphThePageReadsAsMarkdownIsNamedOnTheZone() throws Exception { + assertThat(zoneNotes(true)).containsExactly("a footer written as one line of Word's footer; a paragraph's " + + "markdown marks are written as letters, where the page sets the " + + "text they mark and drops them"); + assertThat(zoneNotes(false)).as("markdown off").isEmpty(); + } + + private static com.demcha.compose.document.node.DocumentNode cell(com.demcha.compose.document.node.DocumentNode content) { + return new com.demcha.compose.document.dsl.TableBuilder().name("Rota").columns(DocumentTableColumn.fixed(200)) + .rowCells(DocumentTableCell.node(content)).build(); + } + + private static List zoneNotes(boolean markdown) throws Exception { + AtomicReference captured = new AtomicReference<>(); + try (DocumentSession session = GraphCompose.document().pageSize(300, 400).margin(DocumentInsets.of(36)) + .markdown(markdown).create()) { + session.chrome().zone(DocumentPageZone.footer(30, page -> new ParagraphBuilder().name("ZoneLine") + .text("**Confidential**").textStyle(DocumentTextStyle.DEFAULT.withSize(8)).build())); + session.pageFlow(page -> page.addParagraph("Body")); + session.export(new DocxSemanticBackend(captured::set)); + } + return captured.get().bySubject().getOrDefault("page zone", List.of()).stream() + .map(DocxExportReport.Note::detail).toList(); + } + + private static List paragraphNotes(boolean markdown, Consumer content) throws Exception { + return notes(markdown, content, "ParagraphNode"); + } + + private static List listNotes(boolean markdown, Consumer content) throws Exception { + return notes(markdown, content, "ListNode"); + } + + private static List notes(boolean markdown, Consumer content, String subject) throws Exception { + AtomicReference captured = new AtomicReference<>(); + try (DocumentSession session = GraphCompose.document().pageSize(300, 400).margin(DocumentInsets.of(30)) + .markdown(markdown).create()) { + session.pageFlow(content::accept); + session.export(new DocxSemanticBackend(captured::set)); + } + assertThat(captured.get()).as("the sink is called once the bytes exist").isNotNull(); + return captured.get().bySubject().getOrDefault(subject, List.of()).stream() + .map(DocxExportReport.Note::detail).toList(); + } +} diff --git a/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxNodeFieldLedgerTest.java b/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxNodeFieldLedgerTest.java index e01dab967..444799e62 100644 --- a/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxNodeFieldLedgerTest.java +++ b/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxNodeFieldLedgerTest.java @@ -141,8 +141,10 @@ private record Entry(Fate fate, String note) { "anchor:REPORTED:drawn; a rule in the flow is bookmarked", "lineCap:REPORTED:a cap other than butt; whether an editor ends a drawn butt line flat is not measured", "fillWidth:WRITTEN", "keepWithNext:REPORTED:of a line drawn in the flow the page keeps blocks together in"); node(ListNode.class, "name:INERT", - "items:REPORTED:a blank one a hangingIndent list draws as its marker alone; any other is written", - "nestedItems:WRITTEN", "marker:WRITTEN", + "items:REPORTED:a blank one a hangingIndent list draws as its marker alone, and the marks of one " + + "the page reads as markdown; any other is written", + "nestedItems:REPORTED:the marks of one the page reads as markdown; any other is written", + "marker:WRITTEN", "textStyle:WRITTEN", "align:REPORTED", "lineSpacing:REPORTED:where the layout's items are not its own and one wraps, an item run " + "onto the next page among them; composed in a table cell, its wrapping not measured; " @@ -170,7 +172,10 @@ private record Entry(Fate fate, String note) { "textStyle:WRITTEN", "align:WRITTEN", "placeholderText:WRITTEN", "padding:REPORTED:the sides its alignment sets it from", "margin:REPORTED:the sides its alignment sets it from"); - node(ParagraphNode.class, "name:INERT", "text:WRITTEN", "inlineRuns:WRITTEN", "textStyle:WRITTEN", + node(ParagraphNode.class, "name:INERT", + "text:REPORTED:where the page reads it as markdown, its marks written as letters; where its " + + "lines are not read and the page's parser drops a mark, not measured; any other is written", + "inlineRuns:WRITTEN", "textStyle:WRITTEN", "align:WRITTEN", "lineSpacing:WRITTEN", "bulletOffset:REPORTED:its letters before the first line; over the flow, as a side of an " + "overlay's pair or as a badge's text, the room it sets lines in by where that moves one",