diff --git a/CHANGELOG.md b/CHANGELOG.md index 4b0957f62..15573a362 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -8,6 +8,55 @@ follow semantic versioning; release dates are ISO 8601. ### Public API +- **A DOCX page zone's text stands on the page's baseline.** A page zone is written as one line + of a Word header or footer. The line was Word's single line for its face. Its top sat at the + zone content's top in a header, and its foot at the content's foot in a footer. So the line's + height, and the room a lone part's padding holds above and below its text, were Word's, and + nothing named them. In LibreOffice a header's text stood up to a point higher than on the + page. Two zones of one kind — a cover's header on the first page, the running header on the + rest — shared one distance from the edge, the last zone's, and the cover's header stood 16pt + high in both editors. A zone paragraph's anchor had no bookmark, without a note. + - **The zone's line is an exact line**, as tall as its tallest part's line on the page. It + stands as far from its page edge as puts that part's baseline where the page has it, since + both editors stand an exact line's baseline four fifths of the way down it (measured on + exact lines). A lone part's padding and margin above and below are in that distance. + - **The line is taller where it needs to be:** + - for a picture in a zone paragraph, which Word stands on the baseline, so the exact line + does not cut its top; + - for a part the page fits smaller, which Word writes at its style's size. + - **A tallest part the page sets in more lines than one** is written as as many exact lines. + Word grows a footer up from its distance, so a footer of two lines stands a line further from + the edge, its first line on the page's first baseline. + - **A zone the layout measures that shares its kind with another page zone** stands in a frame + (`w:framePr`) at its own height, as a text band does. The frame is at least its lines tall, + so a part Word sets in more lines goes on below them. + - **Measured** in Word 16.0.20430 and LibreOffice 26.8 on 8pt and 18pt Lato headers and 8pt + Lato footers. A lone part's text, and a line's tallest part's, stood within 0.1pt of the + page's baseline in each editor, the cover's header included. A smaller part beside a taller + one stays on Word's one baseline, and the note counts it. + - **Text the page sets right against the edge** stands off the page's baseline, as the line + stops at the edge: lower in a header, by what its ascent falls short of four fifths of its + line, and higher in a footer, by what its descent falls short of a fifth. In the default face + only a header's is, about 0.4pt at 18pt. Past a point and a half, the note names it. + - **The `page zone` note:** + - counts a part off the baseline Word sets the line on, and says where the parts after a + part of more lines than one stand is not measured, since Word sets them on a later line; + - names a zone paragraph's anchor, which has no bookmark, since a zone is written into a part + each kind of page repeats; + - names a picture the page sets anywhere but on the baseline; + - names a zone's lines that reach past the page margin by more than half a point. The margin + is then written negative, as for a text band, so Word holds the body at it; LibreOffice + moves the body clear. Within half a point — text under about 22pt set against the margin in + the default face — Word moves the body by as much, unnamed; + - says where the zone's text stands is not measured where the zone's content is built + otherwise for the first page it is drawn on than as written — other text, another face or + size, other pictures. The line is Word's there; + - names a zone whose content is none for no page in particular: it is not written. + - No document of the DOCX fidelity corpus has a page zone; its bytes are unchanged. + - Ledger: the `zones` option moves from a gap to `REPORTED`, and a lone or tallest page field's + padding and margin above and below are written. No entry is a gap any more, and `GAP` is no + longer a fate an entry can take. + - **A DOCX canvas holds its height.** The page gives a canvas its height whatever it holds. The export writes what it holds one block after another, and dropped the room under it, so what followed stood that much higher. The report named the loss only where something followed the @@ -135,7 +184,9 @@ follow semantic versioning; release dates are ISO 8601. page sets it when it meets all of these: - its line starts within a point and a half of Word's, or, against the right margin, ends there; - - it sits on Word's baseline: its tallest part's, with the line standing at the zone's edge; + - it sits on Word's baseline: its tallest part's, with the line standing at the zone's edge + (since placed by that part's baseline: see "A DOCX page zone's text stands on the page's + baseline"); - it is one line; - no prefix stands before it, on the line's left side. @@ -149,7 +200,9 @@ follow semantic versioning; release dates are ISO 8601. zone. In `DocxNodeFieldLedgerTest` a page field's `padding` and `margin` move from a gap to `REPORTED`, and its `align` to `INERT`: the page sets a field in a box a point wider than its number. The `zones` option keeps a gap for the line's height and the room its parts hold above - and below, and a zone paragraph's anchor. 1 node-field gap remains, `CanvasLayerNode.height`. + and below, and a zone paragraph's anchor (since closed: see "A DOCX page zone's text stands on + the page's baseline"). 1 node-field gap remains, `CanvasLayerNode.height` (since closed: see + "A DOCX canvas holds its height"). - **A DOCX export's report names what a paragraph's own fields lose.** Several of a paragraph's own fields were lost without a note: - an auto-sized paragraph's text was written at its style's size, not the one the page fits it @@ -359,7 +412,8 @@ follow semantic versioning; release dates are ISO 8601. lost: 29 fields have nothing to carry, 55 node fields are gaps, and so are page zones, where a paragraph's alignment, spacing and direction and a row's columns are not written. A node class that is not a record, or a kind or field added to the engine, fails it until someone - decides. + decides. (The gaps have since closed, and a gap is no longer a fate: see "A DOCX page zone's + text stands on the page's baseline".) - **A DOCX timeline's dots, and drawings centred on a line, move with their text too.** A drawing in a table cell went into that cell only when the cell held it across, and a dot set in a column of its own beside its entry's text — `CharcoalGold`'s timeline, a row of diff --git a/docs/architecture/backend-capability-matrix.md b/docs/architecture/backend-capability-matrix.md index 1f24eacfb..57e26ee9a 100644 --- a/docs/architecture/backend-capability-matrix.md +++ b/docs/architecture/backend-capability-matrix.md @@ -115,7 +115,7 @@ honour an option ignores it (documented contract). | Page backgrounds (`DocumentSession.pageBackgrounds`, `PageBackgroundFill` — full page, columns, bands) | ✅ `DocumentPageBackgrounds` adds each fill as a shape fragment under every page's content, drawn by the ordinary shape handler | ✅ the same fragments, drawn as shapes on every slide | ⚠️ `DocxPageBackgrounds` — each fill is a rectangle anchored to the page, behind the text. A section of more than one page, or with a header or footer, carries them in every header part it has (default, first page, even pages), so they are drawn on every page; a section without a header gets an empty one against the page edge to carry them, and a later section without fills gets an empty header of its own rather than inheriting them. On a page with no top margin LibreOffice still sets the first line about 3pt lower under that header. A section of one page with neither draws them from the body instead — in the first cell's paragraph when the page opens with a table — and its lines stand where the page sets them in both editors (the seven sidebar CVs stood 2.6 to 3.3pt low in LibreOffice); they are on that page alone, so a page an editor's text runs onto has none. A fill's alpha is carried as the shape's. A two-column layout still flows its columns one after the other, so a column fill can stand beside text that is not its column's | | Watermark (front/back layers) | ✅ `PdfWatermarkRenderer` | ✅ `PptxChromeRenderer` (per-slide shape at the PDF placement math; behind-content applies before fragments, so no z-order surgery) | ❌ (not written; reported `DROPPED`, `watermark`) | | Repeating headers / footers | ✅ `PdfHeaderFooterRenderer` — the zone's `fontName` is resolved through the document's own `FontLibrary`, so a zone draws in the family the author named; unnamed means standard-14 Helvetica, and a code point that family cannot encode is substituted with `?` exactly as body text is | ✅ `PptxChromeRenderer` (positioned per-slide text boxes; `{page}` / `{pages}` / `{date}` tokens with the numbering window rules). The named family reaches the slide run through `PptxFontMapping.familyFor`, and the same family measures the slots — a run measured against one face and typeset in another lands off-centre | ✅ `DocxSemanticBackend.writeBand` (`DocxTextBands`) — one line of a Word header or footer part: the left slot, the centre slot at a centre tab and the right slot at a right tab against the margins; `{page}` / `{pages}` as `PAGE` / `NUMPAGES` (`SECTIONPAGES` per section) fields with the roman or alphabetic switch; `{date}` as the date of the export; the separator as the paragraph's border, a translucent one flattened against white and reported (`translucency`); the header or footer distance from the band's geometry, baseline within 0.1pt in LibreOffice; a band sharing its kind with another band or a page zone stands in a frame (`w:framePr`) at its own height; `showOnFirstPage(false)` or counting from page 2 → an empty first-page part. A band starting after page 2, numbers not counting from 1 on page 1, and a band alone of its kind reaching past the page margin (written as a negative margin, so that Word holds the body at it as the page does; LibreOffice moves the body clear of it) are reported | -| Page zones (node subtree in the band) | ✅ Spliced into the layout graph by `DocumentPageZones`, so the ordinary fragment handlers draw it — no zone-specific code in the backend | ✅ Same splice, same reason: `PptxFixedLayoutBackend.renderGraph` draws every fragment of the graph | ✅ Written into a real `w:ftr` / `w:hdr` part. The band's children become runs on one Word line: a paragraph contributes its runs, a flex spacer becomes the right tab stop, and `PageContext.pageNumber()` / `pageTotal()` become live `PAGE` / `NUMPAGES` fields. Other node kinds are skipped, logged once a kind and reported once each (`DROPPED`, `page zone content`); a zone row's own fill, outline and side borders are reported as `row paint`. Word sets the line's parts one after another from the left margin and, after the first spacer, against the right margin, on one baseline, its tallest part's; the page sets each by the zone's padding, a row's columns and gap, a paragraph's alignment and a part's own sides (a page field's alignment moves nothing: its box is a point wider than its number). The `page zone` note counts the parts the page sets elsewhere — read from the layout's zone fragments: a part stands where the page sets it when its line starts, or against the right margin ends, within a point and a half of Word's, on Word's baseline with the line at the zone's edge, in one line and with no prefix before it on the left side; past a part whose width Word does not keep, or a zone whose nodes the page names or nests otherwise, where a part stands is said to be not measured — and names a zone paragraph's right-to-left direction, prefix letters, fitted size and outline entry, which the line does not carry; the line's height and the room its parts hold above and below are Word's, and a zone paragraph's anchor has no bookmark, none of these yet named. Because Word paginates, `PageContext.number()` refuses here rather than baking a number that would be wrong on every page but one. A zone's `appliesTo` predicate is asked over sample pages (`DocxPageClasses`) and, when it follows Word's first / even / other pages, becomes the matching part — `w:titlePg` for the first page, `w:evenAndOddHeaders` for even pages — with an empty part on the pages it skips; a predicate that picks pages within a kind (the last page) is written on every page and reported | +| Page zones (node subtree in the band) | ✅ Spliced into the layout graph by `DocumentPageZones`, so the ordinary fragment handlers draw it — no zone-specific code in the backend | ✅ Same splice, same reason: `PptxFixedLayoutBackend.renderGraph` draws every fragment of the graph | ✅ Written into a real `w:ftr` / `w:hdr` part. The band's children become runs on one Word line: a paragraph contributes its runs, a flex spacer becomes the right tab stop, and `PageContext.pageNumber()` / `pageTotal()` become live `PAGE` / `NUMPAGES` fields. Other node kinds are skipped, logged once a kind and reported once each (`DROPPED`, `page zone content`); a zone row's own fill, outline and side borders are reported as `row paint`. Word sets the line's parts one after another from the left margin and, after the first spacer, against the right margin, on one baseline, its tallest part's; the page sets each by the zone's padding, a row's columns and gap, a paragraph's alignment and a part's own sides (a page field's alignment moves nothing: its box is a point wider than its number). The `page zone` note counts the parts the page sets elsewhere — read from the layout's zone fragments: a part stands where the page sets it when its line starts, or against the right margin ends, within a point and a half of Word's, on Word's baseline, in one line and with no prefix before it on the left side; past a part whose width Word does not keep, or a zone whose nodes the page names or nests otherwise, where a part stands is said to be not measured — and names a zone paragraph's right-to-left direction, prefix letters, fitted size, markdown marks, outline entry, anchor (no bookmark) and a picture set off the baseline, which the line does not carry. The line is an exact line as tall as the tallest part's line on the page — taller for a picture in it or a part written at a larger size than the page's — standing as far from its edge as puts that part's baseline where the page has it (a lone or tallest part within 0.1pt in Word and LibreOffice, measured on 8pt and 18pt Lato headers and 8pt Lato footers), a lone part's padding and margin above and below included; a tallest part of more lines than one is as many exact lines, a footer's standing as many lines further from its edge; lines reaching past the page margin by more than half a point are held there with a negative margin and named; a zone the layout measures that shares its kind with another page zone stands in a frame (`w:framePr`) at its own height, at least its lines tall; a zone whose content is built otherwise for its first page — other text, face, size or pictures — is not measured, its line Word's, and a zone whose content is none for no page in particular is not written and named. Because Word paginates, `PageContext.number()` refuses here rather than baking a number that would be wrong on every page but one. A zone's `appliesTo` predicate is asked over sample pages (`DocxPageClasses`) and, when it follows Word's first / even / other pages, becomes the matching part — `w:titlePg` for the first page, `w:evenAndOddHeaders` for even pages — with an empty part on the pages it skips; a predicate that picks pages within a kind (the last page) is written on every page and reported | | Protection / encryption | ✅ `PdfDocumentPostProcessor` | ❌ (ignored with a one-time warning — no OOXML encryption support planned) | ❌ (not written, so the file opens unprotected; reported `DROPPED`, `protection`) | | Viewer preferences | ✅ `applyViewerPreferences` in `PdfFixedLayoutBackend` | ❌ (ignored with a one-time warning — PDF-viewer concept) | n/a (not written; reported `DROPPED`, `viewer preferences`) | | Debug guide lines / node labels | ✅ `PdfGuideLinesRenderer`, `PdfNodeLabelRenderer` | ❌ (ignored with a one-time warning — render through the PDF backend to see overlays) | n/a | diff --git a/docs/recipes/docx-export.md b/docs/recipes/docx-export.md index f9c92b748..9526e39fb 100644 --- a/docs/recipes/docx-export.md +++ b/docs/recipes/docx-export.md @@ -1048,17 +1048,44 @@ deterministic export byte-identical only within one day unless `-Dgraphcompose.renderDate` pins it, as for the PDF. A page zone (`session.chrome().zone(...)`) exports as a real -Word header or footer part, with the page number as a live field, and it -sits as far from its page edge as the page puts it — the distance is read -from where the zone's content landed in the resolved layout and written as -`w:pgMar/@w:header` or `@w:footer`, rather than left to Word's 36pt. -A zone is written from its paragraphs, page fields and spacers: anything else in it — a -logo, a panel, a table — is not written, and the export report names it once -(`page zone content`). - -The zone's line is Word's: its parts one after another from the page's left margin, and -those after the first spacer against its right margin, at the right tab the line holds, all -on one baseline, its tallest part's. Where the page sets a part elsewhere — by the zone's +Word header or footer part, with the page number as a live field, and its text stands on the +baseline the page sets it on. The zone's line is an exact line as tall as its tallest part's +line on the page, and it stands as far from its page edge — `w:pgMar/@w:header` or `@w:footer`, +rather than Word's 36pt — as puts that part's baseline where the page has it, both editors +standing an exact line's baseline four fifths of the way down it. So what a lone part's padding +and margin hold above and below its text is in that distance. The line is taller where a +picture in it needs the room above the baseline, as Word stands a zone's picture on it, and +where a part the page fits smaller is written at its style's size. A tallest part the page sets +in more lines than one is as many exact lines; Word grows a footer up from its distance, so a +footer of two lines stands a line further from the edge, its first line on the page's first +baseline. Measured in Word 16.0.20430 and LibreOffice 26.8 on 8pt and 18pt Lato headers and 8pt +Lato footers, a lone part's text, and a line's tallest part's, stood within 0.1pt of the page's +baseline; a smaller part beside it stays on Word's one baseline, and is counted. Text the page +sets right against the edge stands off its baseline, the line stopping at the edge: lower in a +header, by what its ascent falls short of four fifths of its line, and higher in a footer, by +what its descent falls short of a fifth — in the default face only a header's is, about 0.4pt at +18pt; past a point and a half, the report names it. Lines reaching past the page margin, into the +body, by more than half a point are held there as a text band is: the margin is written +negative, and the report names it. Within half a point, as text under about 22pt set against the +margin reaches in the default face, Word moves the body down by as much, unnamed; a framed zone +moves nothing. + +A zone the layout measures that shares its kind with another page zone — a cover's header on +the first page, the running header on the rest — stands in a frame (`w:framePr`) at its own +height, at least its lines tall, since Word holds one distance from the edge for a kind; written +in the flow, the cover's header stood 16pt high. Beside a text band of its kind, the band is +framed. Where the layout shows no text of the zone, or the zone's content is built otherwise for +the first page it is drawn on than it is written — other text, another face or size, other +pictures — the line is Word's, the distance is read from where its content landed, and the +report says where its text stands is not measured. Such a zone is not framed: two of one kind +share Word's one distance, the last one's, and stand one under the other. A zone whose content +is none for no page in particular is not written, and the report names it. A zone is written +from its paragraphs, page fields and spacers: anything else in it — a logo, a panel, a table — +is not written, and the export report names it once (`page zone content`). + +Word sets the line's parts one after another from the page's left margin, and those after +the first spacer against its right margin, at the right tab the line holds, all on one +baseline, its tallest part's. Where the page sets a part elsewhere — by the zone's padding, a row's columns and gap, a paragraph's alignment or a part's own sides — off that baseline, over more than one line or after a prefix, the export report counts it (`page zone`): "1 of its 3 parts stands off where the page sets them". A page field's alignment @@ -1067,8 +1094,9 @@ at another width — after a prefix, auto-sized to a size the file does not hold lines than one — or in a zone whose nodes the page names or nests otherwise than the file, it says where a part stands is not measured. It names what a paragraph in the zone loses of its own too: its right-to-left direction, a prefix's letters, the size an auto-sized one is fitted -to, its outline entry. The line's height and the room its parts hold above and below are -Word's, and an anchor in a zone has no bookmark; none of these is named yet. +to, the markdown marks the page reads, its outline entry, its anchor, which has no bookmark — a zone is written into a part each +kind of page repeats, no one place a bookmark could mark, and a link to it points at none — and +a picture the page sets anywhere but on the baseline, where Word stands it. A zone drawn on some pages only (`appliesTo(...)`) lands on the same pages when Word can say so. Word has a header and footer for the first page, for even pages and for the rest, diff --git a/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxLayoutMetrics.java b/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxLayoutMetrics.java index c40f10902..0fd1fb9fc 100644 --- a/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxLayoutMetrics.java +++ b/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxLayoutMetrics.java @@ -513,18 +513,8 @@ OptionalDouble zoneDistanceFromEdge(int zoneIndex, boolean header, double pageHe * @return the fragments by path, empty when the layout carries no such zone */ Map zoneText(int zoneIndex) { - java.util.regex.Pattern zone = - java.util.regex.Pattern.compile("^@page-zone\\[\\d+]\\[" + zoneIndex + "]"); - // The page a zone's fragment is drawn on is the one its path names. - int first = Integer.MAX_VALUE; - for (Map.Entry> entry : fragments.entrySet()) { - if (zone.matcher(entry.getKey()).find()) { - for (PlacedFragment fragment : entry.getValue()) { - first = Math.min(first, fragment.pageIndex()); - } - } - } - if (first == Integer.MAX_VALUE) { + int first = zoneFirstPage(zoneIndex); + if (first < 0) { return Map.of(); } String prefix = "@page-zone[" + first + "][" + zoneIndex + "]"; @@ -543,6 +533,27 @@ Map zoneText(int zoneIndex) { return text; } + /** + * The first page a page zone is drawn on, which {@link #zoneText} reads it from. + * + * @param zoneIndex the zone's position in the section's zone list + * @return the page's index, counted from 0, or -1 when the layout carries no such zone + */ + int zoneFirstPage(int zoneIndex) { + java.util.regex.Pattern zone = + java.util.regex.Pattern.compile("^@page-zone\\[\\d+]\\[" + zoneIndex + "]"); + // The page a zone's fragment is drawn on is the one its path names. + int first = Integer.MAX_VALUE; + for (Map.Entry> entry : fragments.entrySet()) { + if (zone.matcher(entry.getKey()).find()) { + for (PlacedFragment fragment : entry.getValue()) { + first = Math.min(first, fragment.pageIndex()); + } + } + } + return first == Integer.MAX_VALUE ? -1 : first; + } + /** * The paths the layout gives a tree compiled as a document of its own root: a page zone's * content, which the layout lays out apart from the body. diff --git a/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxSemanticBackend.java b/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxSemanticBackend.java index 0c4946824..46fc59467 100644 --- a/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxSemanticBackend.java +++ b/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxSemanticBackend.java @@ -1261,14 +1261,30 @@ private java.util.Set applyPageZones( ? null : zone.getContent().apply(PageContext.unpaginated()); if (content == null) { + // Built for no page in particular, it builds nothing to write: lost, unless the + // layout shows it drawn on no page either. + if (zone.getContent() != null && (layout.isEmpty() || layout.zoneFirstPage(index) >= 0)) { + String kind = zone.getZone() == DocumentHeaderFooterZone.HEADER ? "header" : "footer"; + report.add(DocxExportReport.Severity.DROPPED, "page zone", + sectioned ? "section " + (sectionIndex + 1) : null, + "a " + kind + " is not written: built for no page in particular, its content is " + + "none; a zone absent from some pages says so through appliesTo"); + } continue; } int at = index; boolean header = zone.getZone() == DocumentHeaderFooterZone.HEADER; + ZonePlacement placement = zonePlacement(at, header, zone, content); + // A zone that shares its kind with another page zone stands at its own height on the + // page, as a band does (writeBand): Word holds one distance from the edge for a kind, + // and lines stacked in one part would keep the order the zones were added in. Beside a + // band of its kind it stays in the flow, the band framed. + boolean framed = placement.measured() && sectionZones.stream() + .anyMatch(other -> other != zone && other.getZone() == zone.getZone() && other.getContent() != null); writers.add(new ZoneWriter(zone.getZone(), pageClassesOf(zone), - part -> writeZoneLine(part, content), () -> { - placeZone(document, zone, at, header); - reportZoneLine(at, header, content); + part -> writeZoneLine(part, content, placement, framed, header), () -> { + placeZone(document, zone, at, header, placement, framed); + reportZoneLine(at, header, content, placement); })); } boolean titlePage = false; @@ -1423,23 +1439,14 @@ private void writeBand(XWPFHeaderFooter part, DocumentHeaderFooter band, boolean long frameTop = toTwips(band.getZone() == DocumentHeaderFooterZone.HEADER ? DocxTextBands.distanceFromEdge(band) : canvasHeight - DocxTextBands.distanceFromEdge(band) - line - border); - List before = part.getParagraphs(); - if (inFrame && !before.isEmpty() && sameFrameHeight(before.get(before.size() - 1), frameTop)) { - // Word takes adjacent paragraphs with the same frame for one frame. - collapsed(part.createParagraph()); + if (inFrame) { + separateFromAnEqualFrame(part, frameTop); } XWPFParagraph para = part.createParagraph(); CTPPr properties = para.getCTP().isSetPPr() ? para.getCTP().getPPr() : para.getCTP().addNewPPr(); if (inFrame) { - org.openxmlformats.schemas.wordprocessingml.x2006.main.CTFramePr frame = properties.addNewFramePr(); - frame.setW(BigInteger.valueOf(toTwips(contentWidth))); - frame.setH(BigInteger.valueOf(toTwips(line + border))); - frame.setHRule(org.openxmlformats.schemas.wordprocessingml.x2006.main.STHeightRule.EXACT); - frame.setHAnchor(org.openxmlformats.schemas.wordprocessingml.x2006.main.STHAnchor.MARGIN); - frame.setX(BigInteger.ZERO); - frame.setVAnchor(org.openxmlformats.schemas.wordprocessingml.x2006.main.STVAnchor.PAGE); - frame.setY(BigInteger.valueOf(frameTop)); - frame.setWrap(org.openxmlformats.schemas.wordprocessingml.x2006.main.STWrap.THROUGH); + frameAcrossTheMargins(properties, frameTop, line + border, + org.openxmlformats.schemas.wordprocessingml.x2006.main.STHeightRule.EXACT); } CTSpacing spacing = properties.isSetSpacing() ? properties.getSpacing() : properties.addNewSpacing(); spacing.setBefore(BigInteger.ZERO); @@ -1573,12 +1580,7 @@ private void placeBand(XWPFDocument document, DocumentHeaderFooter band, boolean // footer taller than its margin — MerchantInvoice's footer row, set down to its 3.4pt // margin, went to a second page under a footer reaching 9.8pt — unless the margin is // written negative, which holds the body at it whatever the band reaches. - // No margin has no negative: the least one stands for it. - if (header) { - margin.setTop(BigInteger.valueOf(-Math.max(1, twipsOf(margin.getTop())))); - } else { - margin.setBottom(BigInteger.valueOf(-Math.max(1, twipsOf(margin.getBottom())))); - } + holdTheBodyAtTheMargin(margin, header); report.add(DocxExportReport.Severity.APPROXIMATED, "page " + zoneName(band), sectioned ? "section " + (sectionIndex + 1) : null, "it reaches " + Math.round(reachOf(band) * 10) / 10.0 + "pt from the page edge, past the " @@ -1703,6 +1705,12 @@ private static DocxPageClasses.PageClass pageClassOf( : DocxPageClasses.PageClass.LATER_ODD; } + /** + * How far past the page margin a zone's line may reach unnamed, in points: Word moves the + * body down by as much, no more than half a point. + */ + private static final double ZONE_REACH_CLEARANCE = 0.5; + /** * Puts a header or footer as far from its page edge as the page puts it. * @@ -1710,19 +1718,40 @@ private static DocxPageClasses.PageClass pageClassOf( * sat 14.5pt higher than the page draws it, on every page. Word holds the distance as * {@code w:pgMar/@w:header} and {@code @w:footer}, so this is a mapping.

* - *

The distance is where the zone's content landed in the resolved layout, which is - * the number the page was drawn with. Without a layout it falls back to the zone's own - * padding on that edge — a band's content is laid from its top, so for a footer that is - * the nearer estimate rather than the exact one, and it is only reached when the - * document could not be laid out at all.

- */ - private void placeZone(XWPFDocument document, DocumentPageZone zone, int index, boolean header) { + *

The distance is the one that stands the zone's line on its tallest part's baseline + * ({@link #zonePlacement}). Where the layout shows none of the zone's text, it is where the + * zone's content landed in the resolved layout, the highest edge of a header's and the + * lowest of a footer's. Without a layout it falls back to the zone's own padding on that + * edge — a band's content is laid from its top, so for a footer that is the nearer estimate + * rather than the exact one, and it is only reached when the document could not be laid + * out at all.

+ * + *

A framed zone stands at its own height whatever the distance, and leaves a distance + * already written to what the part's flow holds: a zone of its kind in the flow, or a text + * band. With none written, it writes its own, where the paragraph the frame is placed from + * stands.

+ * + *

Its paragraph reaching past a positive page margin, into the body, further than + * {@link #ZONE_REACH_CLEARANCE}, the margin is written negative and the reach named, as for + * a text band ({@link #placeBand}); a paragraph of more lines than one reaches as far as all + * of them.

+ * + * @param placement where the zone's line stands in Word + * @param framed whether the zone stands in a frame (see {@link #writeZoneLine}) + */ + private void placeZone(XWPFDocument document, DocumentPageZone zone, int index, boolean header, + ZonePlacement placement, boolean framed) { CTSectPr sectPr = document.getDocument().getBody().isSetSectPr() ? document.getDocument().getBody().getSectPr() : document.getDocument().getBody().addNewSectPr(); CTPageMar margin = sectPr.isSetPgMar() ? sectPr.getPgMar() : sectPr.addNewPgMar(); + if (framed && (header ? margin.getHeader() : margin.getFooter()) != null) { + return; + } double pageHeight = canvasHeight; - OptionalDouble measured = Double.isNaN(pageHeight) + OptionalDouble measured = placement.measured() + ? OptionalDouble.of(placement.distance()) + : Double.isNaN(pageHeight) ? OptionalDouble.empty() : layout.zoneDistanceFromEdge(index, header, pageHeight); DocumentInsets padding = zone.getPadding() == null ? DocumentInsets.zero() : zone.getPadding(); @@ -1732,6 +1761,46 @@ private void placeZone(XWPFDocument document, DocumentPageZone zone, int index, } else { margin.setFooter(BigInteger.valueOf(toTwips(distance))); } + // A paragraph of exact lines reaches as far from the edge as it is placed and tall; past a + // positive margin Word moves the body clear of it, as of a band (placeBand). A frame moves + // nothing. Text set against the margin reaches a fraction of a point past it — its line's + // top four fifths above the baseline rather than its ascent — which moves the body by no + // more. + double reach = placement.measured() ? placement.distance() + placement.height() : Double.NaN; + Object edge = header ? margin.getTop() : margin.getBottom(); + if (!framed && !Double.isNaN(reach) && edge instanceof Number pageMargin && pageMargin.longValue() >= 0 + && toTwips(reach) > pageMargin.longValue() + toTwips(ZONE_REACH_CLEARANCE)) { + holdTheBodyAtTheMargin(margin, header); + String kind = header ? "header" : "footer"; + report.add(DocxExportReport.Severity.APPROXIMATED, "page zone", sectioned ? "section " + (sectionIndex + 1) : null, + "a " + kind + " whose line reaches " + Math.round(reach * 10) / 10.0 + "pt from the page edge, past " + + "the page margin, which is written negative so that Word holds the body at the margin, as the " + + "page does; LibreOffice moves the body clear of it"); + } + } + + /** + * Writes a header's or a footer's page margin negative, which holds the body at the margin + * whatever the header or footer reaches: Word otherwise moves the body clear of it. No margin + * has no negative, so the least one stands for it. + */ + private static void holdTheBodyAtTheMargin(CTPageMar margin, boolean header) { + if (header) { + margin.setTop(BigInteger.valueOf(-Math.max(1, twipsOf(margin.getTop())))); + } else { + margin.setBottom(BigInteger.valueOf(-Math.max(1, twipsOf(margin.getBottom())))); + } + } + + /** + * Parts a framed paragraph from the frame at the same height before it with a hairline + * paragraph: Word takes adjacent paragraphs with the same frame for one frame. + */ + private void separateFromAnEqualFrame(XWPFHeaderFooter part, long frameTop) { + List before = part.getParagraphs(); + if (!before.isEmpty() && sameFrameHeight(before.get(before.size() - 1), frameTop)) { + collapsed(part.createParagraph()); + } } /** @@ -1741,8 +1810,20 @@ private void placeZone(XWPFDocument document, DocumentPageZone zone, int index, * than stacked paragraphs: a row's children become runs in document order, and * a flex spacer becomes the tab that carries the rest to the right margin — * which is how a Word footer is built by hand anyway.

+ * + *

Where the layout shows its text, the line is exact and as tall as {@link #zonePlacement} + * says, which with the distance {@link #placeZone} writes stands it on the tallest part's + * baseline; framed, where the zone shares its kind with another page zone, the frame stands + * at the line's own height on the page, at least the line tall.

*/ - private void writeZoneLine(XWPFHeaderFooter target, DocumentNode content) { + private void writeZoneLine(XWPFHeaderFooter target, DocumentNode content, ZonePlacement placement, + boolean framed, boolean header) { + long frameTop = framed + ? toTwips(header ? placement.distance() : canvasHeight - placement.distance() - placement.height()) + : 0; + if (framed) { + separateFromAnEqualFrame(target, frameTop); + } XWPFParagraph para = target.createParagraph(); para.setSpacingBefore(0); para.setSpacingAfter(0); @@ -1753,11 +1834,24 @@ private void writeZoneLine(XWPFHeaderFooter target, DocumentNode content) { CTPPr properties = para.getCTP().isSetPPr() ? para.getCTP().getPPr() : para.getCTP().addNewPPr(); + if (placement.measured()) { + // An exact line, its baseline four fifths down it as in either editor, stands the + // zone's text on the baseline the page sets it on (zonePlacement). + if (framed) { + // At least its lines: a part Word sets in more lines than the page goes on below + // them, not hidden. + frameAcrossTheMargins(properties, frameTop, placement.height(), + org.openxmlformats.schemas.wordprocessingml.x2006.main.STHeightRule.AT_LEAST); + } + CTSpacing spacing = properties.getSpacing(); + spacing.setLineRule(STLineSpacingRule.EXACT); + spacing.setLine(BigInteger.valueOf(Math.round(placement.line() * POINT_TO_TWIP))); + } CTTabStop tab = properties.addNewTabs().addNewTab(); tab.setVal(STTabJc.RIGHT); tab.setPos(java.math.BigInteger.valueOf(Math.round(zoneRightTab() * TWIPS_PER_POINT))); - List parts = zoneParts(content); + List parts = DocxZoneParts.of(content); // A zone's row is its children on one line; its paint is lost as a body row's is, and // said once, though the line is written into each kind of header the zone is given. if (content instanceof RowNode row && zonePartsReported.add(row)) { @@ -1768,9 +1862,126 @@ private void writeZoneLine(XWPFHeaderFooter target, DocumentNode content) { } } - /** The parts a zone's line is written from, in order: a row's children, or the zone's one node. */ - private static List zoneParts(DocumentNode content) { - return content instanceof RowNode row ? row.children() : List.of(content); + /** + * Where a page zone's line stands in Word. + * + *

The page sets each part of a zone on a baseline of its own. Word's line has one, its + * tallest part's, and in an exact line both editors stand it four fifths of the way down + * (measured on exact lines: {@link DocxTextBands#BASELINE_SHARE}). So the line is exact, as + * tall as its tallest part's line on the page — and as a picture in it, which Word stands on + * the baseline, and a part written at its style's larger size, need — and it stands as far + * from its page edge as puts that part's baseline where the page has it: what a lone part's + * padding and margin hold above and below its text is in that distance. A line the page sets + * past the edge stops at it, its baseline as much further in: lower in a header, higher in + * a footer. A line placed by its parts' edges instead, its top at the highest part's top in + * a header, stands a header's text its padding above it high.

+ * + *

A part the page sets in more lines than one makes the zone's paragraph as many exact + * lines, which Word grows down from a header's distance and up from a footer's: a footer's + * stands the lines below its first further from its edge, so that the first is the one on + * the page's baseline.

+ * + *

Where the page set the zone on the first page it draws it otherwise than the zone is + * written — other text, another face or size, other pictures ({@link DocxZoneParts#readAlike}) + * — the line is not measured, and Word's own.

+ * + * @param line the exact line's height in points, NaN where no part's text is laid out + * @param lines how many lines the part of the most takes on the page + * @param distance from the page's top edge to the paragraph's top in a header, from its foot + * to the paragraph's foot in a footer, in points + * @param baseline the baseline Word sets the first line on, measured up from the page's foot + */ + private record ZonePlacement(double line, int lines, double distance, double baseline) { + + private static final ZonePlacement UNMEASURED = new ZonePlacement(Double.NaN, 0, Double.NaN, Double.NaN); + + /** Whether the layout laid out the text the line is placed by. */ + boolean measured() { + return !Double.isNaN(line); + } + + /** How tall the zone's paragraph stands in Word: its lines, each the exact line. */ + double height() { + return line * lines; + } + } + + /** Where a zone's line stands in Word (see {@link ZonePlacement}). */ + private ZonePlacement zonePlacement(int zoneIndex, boolean header, DocumentPageZone zone, DocumentNode content) { + int firstPage = layout.zoneFirstPage(zoneIndex); + if (Double.isNaN(canvasHeight) || firstPage < 0) { + return ZonePlacement.UNMEASURED; + } + // The page set the zone otherwise there than it is written: its line is not measured by it. + if (!DocxZoneParts.readAlike(content, zone.getContent().apply( + PageContext.paginated(firstPage + 1, Math.max(firstPage + 1, layout.pageCount()))))) { + return ZonePlacement.UNMEASURED; + } + java.util.Map paths = DocxLayoutMetrics.pathsWithin(content); + java.util.Map laid = layout.zoneText(zoneIndex); + ZoneLine tallest = null; + int mostLines = 1; + double picture = 0; + double styled = 0; + for (DocumentNode part : DocxZoneParts.of(content)) { + com.demcha.compose.document.layout.PlacedFragment fragment = + part instanceof ParagraphNode || part instanceof PageFieldNode ? laid.get(paths.get(part)) : null; + if (fragment == null) { + continue; + } + ZoneLine first = firstLineOnThePage(fragment); + if (tallest == null || first.above() + first.below() > tallest.above() + tallest.below()) { + tallest = first; + } + mostLines = Math.max(mostLines, first.lines()); + if (part instanceof ParagraphNode paragraph) { + List lines = + ((com.demcha.compose.document.layout.payloads.ParagraphFragmentPayload) fragment.payload()).lines(); + picture = Math.max(picture, DocxZoneParts.tallestPicture(paragraph)); + // Written at its style's size where the page fits its text smaller, it needs the + // style's line, or its letters' tops are cut. + if (autoSizeLost(paragraph, lines, "its") != null) { + styled = Math.max(styled, styleLineHeight(paragraph.textStyle())); + } + } + } + // As tall as its tallest part's line on the page, as a body paragraph's exact line is, + // and tall enough for a picture in it, which Word stands on the baseline and an exact + // line cuts off at its top, and for a part written at a larger size than the page's. + double line = tallest == null ? 0 + : Math.max(Math.max(tallest.above() + tallest.below(), styled), + picture / DocxTextBands.BASELINE_SHARE); + if (!(line > 0)) { + return ZonePlacement.UNMEASURED; + } + double above = DocxTextBands.BASELINE_SHARE * line; + // A footer's paragraph grows up from its distance: its last line's baseline is the one + // that distance places, the lines above it each an exact line higher. + double distance = DocxTextBands.distanceFromEdge(header, + header ? canvasHeight - tallest.baseline() : tallest.baseline() - (mostLines - 1) * line, line); + double baseline = header ? canvasHeight - distance - above : distance + (mostLines - 1) * line + (line - above); + return new ZonePlacement(line, mostLines, distance, baseline); + } + + /** + * Frames a header's or footer's paragraph across the margins at a height on the page, where + * the part's flow does not move it and it moves nothing in the body. + * + * @param top the frame's top, from the page's top edge, in twips + * @param height the frame's height in points + * @param rule whether the frame is that height exactly, or at least, growing down + */ + private void frameAcrossTheMargins(CTPPr properties, long top, double height, + org.openxmlformats.schemas.wordprocessingml.x2006.main.STHeightRule.Enum rule) { + org.openxmlformats.schemas.wordprocessingml.x2006.main.CTFramePr frame = properties.addNewFramePr(); + frame.setW(BigInteger.valueOf(toTwips(contentWidth))); + frame.setH(BigInteger.valueOf(toTwips(height))); + frame.setHRule(rule); + frame.setHAnchor(org.openxmlformats.schemas.wordprocessingml.x2006.main.STHAnchor.MARGIN); + frame.setX(BigInteger.ZERO); + frame.setVAnchor(org.openxmlformats.schemas.wordprocessingml.x2006.main.STVAnchor.PAGE); + frame.setY(BigInteger.valueOf(top)); + frame.setWrap(org.openxmlformats.schemas.wordprocessingml.x2006.main.STWrap.THROUGH); } /** Where a zone's line holds its right tab, from the left margin: at the right margin, in points. */ @@ -1793,8 +2004,9 @@ private double zoneRightTab() { * after the first spacer against its right margin, at the right tab the line holds; a part * after a second spacer goes to no tab the line holds. The page sets each part where the * zone's padding, a row's columns and gap and the part's own alignment and sides put it, on - * a baseline of its own. Word's line has one baseline, its tallest part's, and stands at the - * zone's edge — the foot of a footer, the head of a header ({@link #placeZone}). A part + * a baseline of its own. Word's line has one baseline, the one it is placed on + * ({@link #zonePlacement}): its tallest part's, where the page sets that part, or off it + * where the line stops at the page's edge. A part * stands where the page sets it when it is one line — the zone's line is written with none * of what holds a body paragraph's breaks where the page sets them — on Word's baseline, and * its line starts within {@link #ZONE_PLACE_CLEARANCE} of Word's start, or, against the @@ -1802,23 +2014,28 @@ private double zoneRightTab() { * the prefix being unwritten. Word's place for a part follows from the widths of those * before it on its side of the line: past a part whose width Word does not keep — of more * lines than one, after a prefix, or auto-sized to a size the file does not hold — or one - * the layout does not show, where it stands across the line is not measured. A zone whose - * nodes the page names or nests otherwise than the file, built for no page in particular, - * shows none of its parts.

+ * the layout does not show, where it stands across the line is not measured; past a part of + * more lines than one — Word sets the parts after it on a later line — where they stand is + * not measured at all. A zone whose nodes the page names or nests otherwise than the file, + * built for no page in particular, shows none of its parts.

* - *

A paragraph's own losses are named too: its right-to-left text is written left to - * right, its prefix's letters are not written, its text is written at its style's size, - * and its outline entry is not written.

+ *

A paragraph's own losses are named too ({@link #zoneParagraphLost}): its right-to-left + * text is written left to right, its prefix's letters are not written, its text is written + * at its style's size or with the markdown marks the page reads, its outline entry is not + * written, its anchor has no bookmark, and a picture the page sets off the baseline stands + * on it.

* * @param zoneIndex the zone's position in the section's zone list * @param header whether the zone is a header * @param content the zone's content, as written + * @param placement where the line stands in Word */ - private void reportZoneLine(int zoneIndex, boolean header, DocumentNode content) { - List parts = zoneParts(content); + private void reportZoneLine(int zoneIndex, boolean header, DocumentNode content, ZonePlacement placement) { + List parts = DocxZoneParts.of(content); java.util.Map paths = DocxLayoutMetrics.pathsWithin(content); java.util.Map laid = layout.zoneText(zoneIndex); - boolean measured = !laid.isEmpty() && !Double.isNaN(canvasLeftMargin); + // Read where the line is placed by what the page set: not where the page set other text. + boolean measured = placement.measured() && !laid.isEmpty() && !Double.isNaN(canvasLeftMargin); // Each text part's side of the line: 0 from the left margin, 1 against the right, 2 // after a second spacer, where the line holds no tab for it. java.util.Map side = new java.util.IdentityHashMap<>(); @@ -1870,16 +2087,21 @@ private void reportZoneLine(int zoneIndex, boolean header, DocumentNode content) onThePage.put(part, firstLineOnThePage(fragment)); } } - double baseline = wordsBaseline(onThePage.values(), header); + // Read from the same text the line is placed by: where none is, no part is read either. + double baseline = placement.baseline(); int off = 0; int unread = 0; + // Past a part of more lines than one, Word sets the rest on a later line of its own. + boolean broken = false; for (DocumentNode part : texts) { ZoneLine line = onThePage.get(part); Double wordStart = wordsStart.get(part); Double wordEnd = wordsEnd.get(part); boolean prefixed = side.get(part) == 0 && part instanceof ParagraphNode paragraph && setsAPrefixBeforeTheFirstLine(paragraph); - if (line == null) { + boolean afterABreak = broken; + broken |= line != null && !line.oneLine(); + if (line == null || afterABreak) { unread++; } else if (side.get(part) == 2 || !line.oneLine() || prefixed || Math.abs(line.baseline() - baseline) > ZONE_PLACE_CLEARANCE) { @@ -1921,7 +2143,11 @@ private void reportZoneLine(int zoneIndex, boolean header, DocumentNode content) /** * What a paragraph of a page zone loses of its own on the zone's line: its direction, its - * prefix's letters, the size its text is fitted to, its markdown marks, and its outline entry. + * prefix's letters, the size its text is fitted to, its markdown marks, its outline entry, + * its anchor's bookmark, and where it sets a picture off the baseline. A zone is written into + * a part each kind of page repeats, which is no one place in the document a bookmark could + * mark; on the page, a link to it lands on the last page drawing it, and a page reference + * does not find it. * * @param lines the lines the page laid it out in, empty where they are not read */ @@ -1945,6 +2171,12 @@ private List zoneParagraphLost(ParagraphNode node, List lines, boolean header) { - ZoneLine tallest = null; - double edge = header ? Double.NEGATIVE_INFINITY : Double.POSITIVE_INFINITY; - for (ZoneLine line : lines) { - if (tallest == null || line.above() + line.below() > tallest.above() + tallest.below()) { - tallest = line; - } - edge = header ? Math.max(edge, line.top()) : Math.min(edge, line.bottom()); - } - if (tallest == null) { - return Double.NaN; - } - return header ? edge - tallest.above() : edge + tallest.below(); + text.lines().size()); } private void appendZonePart(XWPFParagraph para, DocumentNode part) { @@ -8615,9 +8831,10 @@ private static void setTextBrokenAtLines(XWPFRun run, String text) { * paragraph, so every line of it is then at least the picture's reach, and otherwise the * editor's own measure of its text, which in LibreOffice is taller than the page's — * where the page makes only the line holding the picture taller. A paragraph written with - * no exact height — a page zone's, one with no layout — grows to its pictures on its own - * and is left alone. A paragraph of one line is held to the page's line instead where it - * can be ({@link #holdPicturesInTheLine}).

+ * no exact height — one with no layout, a page zone's the layout does not measure — grows + * to its pictures on its own and is left alone; a page zone's exact line is already as tall + * as its pictures need ({@link #zonePlacement}). A paragraph of one line is held to the + * page's line instead where it can be ({@link #holdPicturesInTheLine}).

*/ private static void makeRoomForPictures(XWPFParagraph para, PictureReach pictures) { CTPPr properties = para.getCTP().getPPr(); diff --git a/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxTextBands.java b/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxTextBands.java index ab4181f04..f707380bd 100644 --- a/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxTextBands.java +++ b/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxTextBands.java @@ -106,11 +106,24 @@ static double lineHeight(DocumentHeaderFooter band) { static double distanceFromEdge(DocumentHeaderFooter band) { double line = lineHeight(band); if (band.getZone() == DocumentHeaderFooterZone.HEADER) { - double baseline = band.getHeight() - band.getFontSize() / 2.0; - return Math.max(0, baseline - line * BASELINE_SHARE); + return distanceFromEdge(true, band.getHeight() - band.getFontSize() / 2.0, line); } - double baseline = band.getHeight() - band.getFontSize(); - return Math.max(0, baseline - line * (1 - BASELINE_SHARE)); + return distanceFromEdge(false, band.getHeight() - band.getFontSize(), line); + } + + /** + * How far from its page edge an exact line stands to set its baseline where the page has + * it, in points: from the top of the page to the top of a header's line, or from the bottom + * of the page to the bottom of a footer's. A line that would stand past the edge stops at + * it, its baseline as far in as that leaves. + * + * @param header whether the line is a header's + * @param baselineFromEdge how far the page sets the baseline from the edge: down from the + * page's top for a header, up from its foot for a footer + * @param line the exact line's height + */ + static double distanceFromEdge(boolean header, double baselineFromEdge, double line) { + return Math.max(0, baselineFromEdge - line * (header ? BASELINE_SHARE : 1 - BASELINE_SHARE)); } /** diff --git a/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxZoneParts.java b/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxZoneParts.java new file mode 100644 index 000000000..8441465b9 --- /dev/null +++ b/render-docx/src/main/java/com/demcha/compose/document/backend/semantic/docx/DocxZoneParts.java @@ -0,0 +1,173 @@ +package com.demcha.compose.document.backend.semantic.docx; + +import com.demcha.compose.document.node.DocumentNode; +import com.demcha.compose.document.node.InlineHighlightRun; +import com.demcha.compose.document.node.InlineImageAlignment; +import com.demcha.compose.document.node.InlineImageRun; +import com.demcha.compose.document.node.InlineRun; +import com.demcha.compose.document.node.InlineShapeRun; +import com.demcha.compose.document.node.InlineSvgRun; +import com.demcha.compose.document.node.InlineTextRun; +import com.demcha.compose.document.node.PageFieldNode; +import com.demcha.compose.document.node.ParagraphNode; +import com.demcha.compose.document.node.RowNode; +import com.demcha.compose.document.style.DocumentTextStyle; + +import java.util.List; +import java.util.Objects; + +/** + * The parts of a page zone's line, as a Word header or footer writes them: what they are, what + * the line they stand on needs, and whether the page drew them as the file holds them. + */ +final class DocxZoneParts { + + private DocxZoneParts() { + } + + /** The parts a zone's line is written from, in order: a row's children, or the zone's one node. */ + static List of(DocumentNode content) { + return content instanceof RowNode row ? row.children() : List.of(content); + } + + /** + * Whether content built for a page reads as the content written, as far as a zone's line is + * set by it: its parts one by one, each paragraph's text and face and each of its runs' — the + * letters, their face, and a picture's size and place — and each page field's kind and face. + * Colour is left out: it sets nothing of the line, and a colour built afresh is not equal to + * itself. + * + * @param written the content as written + * @param drawn the content as built for the page, or {@code null} where it is not known + */ + static boolean readAlike(DocumentNode written, DocumentNode drawn) { + if (drawn == null) { + return false; + } + List parts = of(written); + List others = of(drawn); + if (parts.size() != others.size()) { + return false; + } + for (int index = 0; index < parts.size(); index++) { + if (!partAlike(parts.get(index), others.get(index))) { + return false; + } + } + return true; + } + + private static boolean partAlike(DocumentNode part, DocumentNode other) { + if (part instanceof ParagraphNode paragraph) { + return other instanceof ParagraphNode drawn + && Objects.equals(paragraph.text(), drawn.text()) + && sameFace(paragraph.textStyle(), drawn.textStyle()) + && Objects.equals(paragraph.autoSize(), drawn.autoSize()) + && runsAlike(paragraph.inlineRuns(), drawn.inlineRuns()); + } + if (part instanceof PageFieldNode field) { + return other instanceof PageFieldNode drawn && field.kind() == drawn.kind() + && sameFace(field.textStyle(), drawn.textStyle()); + } + return part.getClass() == other.getClass(); + } + + private static boolean runsAlike(List runs, List others) { + if (runs.size() != others.size()) { + return false; + } + for (int index = 0; index < runs.size(); index++) { + if (!runAlike(runs.get(index), others.get(index))) { + return false; + } + } + return true; + } + + private static boolean runAlike(InlineRun run, InlineRun other) { + if (run.getClass() != other.getClass()) { + return false; + } + if (run instanceof InlineTextRun text) { + InlineTextRun drawn = (InlineTextRun) other; + return Objects.equals(text.text(), drawn.text()) && sameFace(text.textStyle(), drawn.textStyle()); + } + if (run instanceof InlineHighlightRun text) { + InlineHighlightRun drawn = (InlineHighlightRun) other; + return Objects.equals(text.text(), drawn.text()) && sameFace(text.textStyle(), drawn.textStyle()); + } + if (run instanceof InlineImageRun image) { + InlineImageRun drawn = (InlineImageRun) other; + return samePlace(image.width(), image.height(), image.alignment(), image.baselineOffset(), + drawn.width(), drawn.height(), drawn.alignment(), drawn.baselineOffset()); + } + if (run instanceof InlineSvgRun svg) { + InlineSvgRun drawn = (InlineSvgRun) other; + return samePlace(svg.width(), svg.height(), svg.alignment(), svg.baselineOffset(), + drawn.width(), drawn.height(), drawn.alignment(), drawn.baselineOffset()); + } + InlineShapeRun shape = (InlineShapeRun) run; + InlineShapeRun drawn = (InlineShapeRun) other; + return samePlace(shape.width(), shape.height(), shape.alignment(), shape.baselineOffset(), + drawn.width(), drawn.height(), drawn.alignment(), drawn.baselineOffset()); + } + + /** Whether two styles set letters alike: face, size and weight or slant, whatever their colour. */ + private static boolean sameFace(DocumentTextStyle style, DocumentTextStyle other) { + if (style == null || other == null) { + return style == other; + } + return Objects.equals(style.fontName(), other.fontName()) + && Double.compare(style.size(), other.size()) == 0 + && style.decoration() == other.decoration(); + } + + private static boolean samePlace(double width, double height, InlineImageAlignment alignment, double offset, + double otherWidth, double otherHeight, InlineImageAlignment otherAlignment, + double otherOffset) { + return Double.compare(width, otherWidth) == 0 && Double.compare(height, otherHeight) == 0 + && alignment == otherAlignment && Double.compare(offset, otherOffset) == 0; + } + + /** + * How tall the tallest picture among a zone paragraph's runs stands above its baseline in + * Word: the paragraph's lines are not the body's, so a picture is written on the baseline, + * as tall as it is drawn. + */ + static double tallestPicture(ParagraphNode paragraph) { + double tallest = 0; + for (InlineRun run : paragraph.inlineRuns()) { + if (run instanceof InlineImageRun image) { + tallest = Math.max(tallest, image.height()); + } else if (run instanceof InlineSvgRun svg) { + tallest = Math.max(tallest, svg.height()); + } else if (run instanceof InlineShapeRun shape) { + // Drawn with its stroke's reach and the empty margin a shape keeps round its ink. + tallest = Math.max(tallest, DocxShapePictures.of(shape).height()); + } + } + return tallest; + } + + /** + * Whether the page sets an inline picture anywhere but on its line's baseline: Word stands a + * zone paragraph's on it, its lines not being the body's. + */ + static boolean setOffTheBaseline(InlineRun run) { + InlineImageAlignment alignment; + double offset; + if (run instanceof InlineImageRun image) { + alignment = image.alignment(); + offset = image.baselineOffset(); + } else if (run instanceof InlineSvgRun svg) { + alignment = svg.alignment(); + offset = svg.baselineOffset(); + } else if (run instanceof InlineShapeRun shape) { + alignment = shape.alignment(); + offset = shape.baselineOffset(); + } else { + return false; + } + return alignment != InlineImageAlignment.BASELINE || offset != 0; + } +} diff --git a/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxNodeFieldLedgerTest.java b/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxNodeFieldLedgerTest.java index 85369311a..a5dafdbde 100644 --- a/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxNodeFieldLedgerTest.java +++ b/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxNodeFieldLedgerTest.java @@ -47,8 +47,8 @@ /** * What the DOCX export does with every field of every node, and with every output option: - * writes it, reports it when the author set it, has nothing for Word to carry, or leaves a - * known gap that names what is lost. + * writes it, reports it when the author set it, or has nothing for Word to carry. There is no + * fourth answer: a field the export neither writes nor names is not one this ledger can hold. * *

The export's report promises a caller that what the file does not carry is said, not * approximated in silence. That promise is only as good as the fields someone thought about: @@ -59,32 +59,32 @@ * *

A node's entries are about the body. A page zone is written as one line of its * paragraphs' runs, page fields and tabs, which loses more of a paragraph and a row than the - * body does: where its parts stand and what its paragraphs lose of their own is reported, and - * what is not is recorded once, as the {@code zones} output option's gap.

+ * body does: where its parts stand and what its paragraphs lose of their own is reported, once, + * as the {@code zones} output option's entry.

* *

The entries are claims about the export, not proofs of it: a {@code WRITTEN} field is - * proved by the export's own tests, a {@code REPORTED} one by a report test, and a {@code GAP} - * is a loss the report does not yet name, each with what is lost. Closing a gap moves its entry, - * in the same change that makes the export write or report it.

+ * proved by the export's own tests, a {@code REPORTED} one by a report test. A field added later + * takes one of the three fates in the change that adds it.

*/ class DocxNodeFieldLedgerTest { private enum Fate { /** Reaches the Word file, directly or through the geometry the layout resolved. */ WRITTEN, - /** Not in the Word file, and named in the export's report when the author set it. */ + /** + * Not in the Word file in some or all cases, and named in the export's report there when + * the author set it; the entry says which, where it is not every case. + */ REPORTED, /** * Nothing for a Word file to carry: a name the report addresses the node by, an * identity the layout resolves a wrapper's place by, or a field the page itself does * not apply. */ - INERT, - /** Lost, in some or all cases, without a report note yet; the entry says what. */ - GAP + INERT } - /** A field's fate, and for a gap what is lost, for a report what names it. */ + /** A field's fate, and a note on it: for a report the cases it names, for another fate why. */ private record Entry(Fate fate, String note) { } @@ -108,11 +108,14 @@ private record Entry(Fate fate, String note) { "headersAndFooters:REPORTED:a band's translucent separator, flattened against white; any other is " + "written, or reported where Word's parts cannot hold it", // The node entries below are the body's. A zone is written as one line of its - // paragraphs' runs, page fields and tabs; what of that the report does not name is - // a gap of its own. - "zones:GAP:in a page zone, its line's height and the room its parts hold above and below their " - + "text, which Word sets its own way, and a paragraph's anchor; where its parts stand across the " - + "line and on its baseline, and what else a zone holds, is reported"); + // paragraphs' runs, page fields and tabs, and has its own entry here. + "zones:REPORTED:where its parts stand across the line and off the baseline Word sets it on, what " + + "its paragraphs lose of their own — an anchor, a picture's place among them — what else a zone " + + "holds, lines reaching past the page margin, a line the layout does not measure, and a zone " + + "built as nothing for no page in particular; its line's height, its lines, its tallest part's " + + "baseline and the room that part holds above and below its text are written, as exact lines " + + "placed from the edge, in a frame at its height where a measured zone shares its kind with " + + "another page zone"); static { node(AlignNode.class, "name:INERT", "child:WRITTEN", "align:WRITTEN", "margin:WRITTEN"); @@ -181,11 +184,11 @@ private record Entry(Fate fate, String note) { "align:INERT:the page sets a field in a box a point wider than its number, which its " + "alignment moves it no further within", "padding:REPORTED:at its sides, and above and below beside other parts, where they stand the " - + "field off the place Word sets it on the zone's line; a lone field's above and below, in the " - + "zones option's gap", + + "field off the place Word sets it on the zone's line; a lone or tallest field's above and " + + "below are written, where its line stands", "margin:REPORTED:at its sides, and above and below beside other parts, where they stand the " - + "field off the place Word sets it on the zone's line; a lone field's above and below, in the " - + "zones option's gap"); + + "field off the place Word sets it on the zone's line; a lone or tallest field's above and " + + "below are written, where its line stands"); node(PageReferenceNode.class, "name:INERT", "anchor:REPORTED:where the anchor has no bookmark, its number is written as text", "textStyle:WRITTEN", "align:WRITTEN", "placeholderText:WRITTEN", @@ -292,7 +295,7 @@ void everyFieldOfEveryNodeHasAnEntry() { } }); assertThat(undecided).as("fields with no entry: decide whether the DOCX export writes them, " - + "reports them, or leaves a gap, and say which here").isEmpty(); + + "reports them, or has nothing to carry, and say which here").isEmpty(); assertThat(stale).as("entries for fields the node no longer has").isEmpty(); } @@ -303,27 +306,11 @@ void everyOutputOptionHasAnEntry() { .containsExactlyInAnyOrderElementsOf(OUTPUT_OPTIONS.keySet()); } - @Test - void everyGapSaysWhatIsLost() { - List unexplained = new ArrayList<>(); - NODES.forEach((kind, entries) -> entries.forEach((field, entry) -> { - if (entry.fate() == Fate.GAP && entry.note().isBlank()) { - unexplained.add(kind.getSimpleName() + "." + field); - } - })); - OUTPUT_OPTIONS.forEach((option, entry) -> { - if (entry.fate() == Fate.GAP && entry.note().isBlank()) { - unexplained.add("DocumentOutputOptions." + option); - } - }); - assertThat(unexplained).as("a gap names what the Word file loses").isEmpty(); - } - private static void node(Class kind, String... specs) { NODES.put(kind, fields(specs)); } - /** {@code field:FATE}, or {@code field:GAP:what is lost}. */ + /** {@code field:FATE}, or {@code field:FATE:a note}, as {@code field:REPORTED:the cases the report names}. */ private static Map fields(String... specs) { Map entries = new LinkedHashMap<>(); for (String spec : specs) { diff --git a/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxPageZonePositionTest.java b/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxPageZonePositionTest.java index 84510a668..9a7da03eb 100644 --- a/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxPageZonePositionTest.java +++ b/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxPageZonePositionTest.java @@ -25,8 +25,9 @@ *

Nothing was written, so Word used its own distance — 36pt — and the probe's footer sat * 14.5pt higher than the page draws it, on every page. The engine does not state the * distance either: a zone is a band of a given height against the edge, with its content - * laid out inside it from the top. So the distance is read from where the content landed - * in the resolved layout.

+ * laid out inside it from the top. So the distance is read from the resolved layout — from the + * baseline the zone's text stands on there (DocxZoneLineTest), or, with none, from where its + * content landed.

* * @author Artem Demchyshyn */ @@ -49,12 +50,15 @@ void aFootersDistanceFollowsItsBandRatherThanWordsDefault() throws Exception { } @Test - void aHeadersDistanceIsItsContentsTopFromThePageTop() throws Exception { - // A header's content starts at the band's top, so its padding is exactly the gap - // between the page's top edge and the content. - long distance = headerDistance(zone(DocumentHeaderFooterZone.HEADER, 40, new DocumentInsets(9, 0, 0, 0))); - - assertThat(distance).isEqualTo(Math.round(9 * TWIPS_PER_POINT)); + void aHeadersDistanceFollowsItsPadding() throws Exception { + // A header's content starts at the band's top, its padding down: 10pt more padding stands + // it 10pt further from the edge. Where its baseline stands is DocxZoneLineTest's. + long shallow = headerDistance(zone(DocumentHeaderFooterZone.HEADER, 40, new DocumentInsets(9, 0, 0, 0))); + long deep = headerDistance(zone(DocumentHeaderFooterZone.HEADER, 40, new DocumentInsets(19, 0, 0, 0))); + + assertThat(deep - shallow).isEqualTo(Math.round(10 * TWIPS_PER_POINT)); + assertThat(shallow).as("inside the band, and not Word's 720-twip default") + .isGreaterThanOrEqualTo(0L).isLessThan(Math.round(9 * TWIPS_PER_POINT)); } @Test diff --git a/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxTextBandsTest.java b/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxTextBandsTest.java index d5c688acd..6c81cda06 100644 --- a/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxTextBandsTest.java +++ b/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxTextBandsTest.java @@ -65,6 +65,15 @@ void aFootersBaselineStandsWhereThePageSetsIt() { .isCloseTo(19, within(1e-9)); } + @Test + void anExactLineStandsItsBaselineWhereThePageSetsItAndStopsAtTheEdge() { + // A header's 10pt line puts its baseline 8pt below its top, a footer's 2pt above its foot. + assertThat(DocxTextBands.distanceFromEdge(true, 20, 10)).isCloseTo(12, within(1e-9)); + assertThat(DocxTextBands.distanceFromEdge(false, 6, 10)).isCloseTo(4, within(1e-9)); + assertThat(DocxTextBands.distanceFromEdge(true, 5, 10)).as("past the top edge").isZero(); + assertThat(DocxTextBands.distanceFromEdge(false, 1, 10)).as("past the foot").isZero(); + } + @Test void aBandTooLowForItsLineStartsAtTheEdge() { DocumentHeaderFooter header = DocumentHeaderFooter.builder() diff --git a/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxZoneLineTest.java b/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxZoneLineTest.java new file mode 100644 index 000000000..713379cbe --- /dev/null +++ b/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxZoneLineTest.java @@ -0,0 +1,644 @@ +package com.demcha.compose.document.backend.semantic.docx; + +import com.demcha.compose.GraphCompose; +import com.demcha.compose.document.api.DocumentSession; +import com.demcha.compose.document.dsl.ParagraphBuilder; +import com.demcha.compose.document.dsl.RowBuilder; +import com.demcha.compose.document.output.DocumentHeaderFooterZone; +import com.demcha.compose.document.output.DocumentPageZone; +import com.demcha.compose.document.style.DocumentInsets; +import com.demcha.compose.document.style.DocumentTextStyle; +import org.apache.pdfbox.Loader; +import org.apache.pdfbox.pdmodel.PDDocument; +import org.apache.pdfbox.text.PDFTextStripper; +import org.apache.pdfbox.text.TextPosition; +import org.apache.poi.xwpf.usermodel.XWPFDocument; +import org.apache.poi.xwpf.usermodel.XWPFHeaderFooter; +import org.apache.poi.xwpf.usermodel.XWPFParagraph; +import org.junit.jupiter.api.Test; +import org.openxmlformats.schemas.wordprocessingml.x2006.main.CTFramePr; +import org.openxmlformats.schemas.wordprocessingml.x2006.main.CTPageMar; +import org.openxmlformats.schemas.wordprocessingml.x2006.main.CTSpacing; +import org.openxmlformats.schemas.wordprocessingml.x2006.main.STLineSpacingRule; + +import java.io.ByteArrayInputStream; +import java.io.IOException; +import java.util.ArrayList; +import java.util.List; +import java.util.concurrent.atomic.AtomicReference; +import java.util.function.Consumer; + +import static org.assertj.core.api.Assertions.assertThat; +import static org.assertj.core.api.Assertions.within; + +/** + * A page zone's line stands its text on the baseline the page sets it on. + * + *

A zone's line is an exact line, as tall as its tallest part's line on the page and as a + * picture in it needs, standing as far from its edge as puts that part's baseline where the page + * has it — both editors stand an exact line's baseline four fifths down it, as measured + * ({@link DocxTextBands#BASELINE_SHARE}). So the room a lone part's padding and margin hold above + * and below its text is in that distance. Placed by its content's edges, Word's own line stood a + * header's text its padding high. A zone that shares its kind with another page zone stands in a + * frame at its own height, as a band does.

+ * + *

The page's baseline is read from the PDF the engine draws; Word's from the file: the + * distance from the edge, or the frame's place, and the exact line.

+ */ +class DocxZoneLineTest { + + private static final double PAGE_HEIGHT = 600; + private static final DocumentTextStyle CHROME = DocumentTextStyle.DEFAULT.withSize(8); + + @Test + void aHeadersTextStandsOnThePagesBaseline() throws Exception { + Exported exported = export(zone(DocumentHeaderFooterZone.HEADER, 40, new DocumentInsets(9, 0, 0, 0), + "Chrome", CHROME)); + + assertThat(fromTheTop(exported.margin().getHeader()) + share(exported.headerLine())) + .as("Word's baseline, from the page's top") + .isCloseTo(exported.baseline("Chrome"), within(0.1)); + } + + @Test + void aFootersTextStandsOnThePagesBaseline() throws Exception { + Exported exported = export(zone(DocumentHeaderFooterZone.FOOTER, 40, new DocumentInsets(6, 0, 0, 0), + "Chrome", CHROME)); + double line = lineOf(exported.footerLine()); + + assertThat(PAGE_HEIGHT - fromTheTop(exported.margin().getFooter()) - (line - share(exported.footerLine()))) + .isCloseTo(exported.baseline("Chrome"), within(0.1)); + } + + @Test + void theRoomALonePartHoldsAboveItsTextIsInTheDistance() throws Exception { + // A page number padded 6pt down stands 6pt lower on the page, and in Word. + Exported plain = export(fieldZone(DocumentInsets.zero())); + Exported padded = export(fieldZone(new DocumentInsets(6, 0, 0, 0))); + + assertThat(fromTheTop(padded.margin().getHeader()) - fromTheTop(plain.margin().getHeader())) + .isCloseTo(padded.baseline("1") - plain.baseline("1"), within(0.1)) + .isCloseTo(6, within(0.1)); + } + + @Test + void aLonePageFieldStandsOnThePagesBaselineAndItsMarginIsInTheDistanceToo() throws Exception { + Exported header = export(fieldZone(DocumentHeaderFooterZone.HEADER, new DocumentInsets(6, 0, 0, 0), + DocumentInsets.zero())); + Exported footer = export(fieldZone(DocumentHeaderFooterZone.FOOTER, DocumentInsets.zero(), + new DocumentInsets(0, 0, 6, 0))); + Exported plain = export(fieldZone(DocumentHeaderFooterZone.HEADER, DocumentInsets.zero(), DocumentInsets.zero())); + Exported spaced = export(fieldZone(DocumentHeaderFooterZone.HEADER, DocumentInsets.zero(), + new DocumentInsets(6, 0, 0, 0))); + + assertThat(fromTheTop(header.margin().getHeader()) + share(header.headerLine())) + .isCloseTo(header.baseline("1"), within(0.1)); + assertThat(PAGE_HEIGHT - fromTheTop(footer.margin().getFooter()) + - (lineOf(footer.footerLine()) - share(footer.footerLine()))) + .isCloseTo(footer.baseline("1"), within(0.1)); + assertThat(fromTheTop(spaced.margin().getHeader()) - fromTheTop(plain.margin().getHeader())) + .as("its margin above, as its padding") + .isCloseTo(spaced.baseline("1") - plain.baseline("1"), within(0.1)) + .isCloseTo(6, within(0.1)); + } + + @Test + void aFootersLineOfPartsStandsOnItsTallestPartsBaseline() throws Exception { + Exported exported = export(DocumentPageZone.builder().zone(DocumentHeaderFooterZone.FOOTER).height(40) + .padding(new DocumentInsets(4, 0, 0, 0)) + .content(page -> new RowBuilder().name("Line") + .addParagraph(p -> p.text("v2.4").textStyle(CHROME)) + .flexSpacer() + .addParagraph(p -> p.text("Acme").textStyle(DocumentTextStyle.DEFAULT.withSize(18))) + .build()) + .build()); + XWPFParagraph line = exported.footerLine(); + + assertThat(PAGE_HEIGHT - fromTheTop(exported.margin().getFooter()) - (lineOf(line) - share(line))) + .isCloseTo(exported.baseline("Acme"), within(0.1)); + assertThat(lineOf(line)).as("as tall as the 18pt part's line, not the 8pt one's").isBetween(16.0, 18.0); + } + + @Test + void aFootersParagraphOfTwoLinesStandsBothOnThePagesBaselines() throws Exception { + // Word grows a footer up from its distance: the paragraph stands a line further from the + // edge than one line would, its first line on the page's first baseline. + Exported exported = export(zone(DocumentHeaderFooterZone.FOOTER, 40, new DocumentInsets(4, 0, 0, 0), + "First\nSecond", CHROME)); + XWPFParagraph line = exported.footerLine(); + double foot = PAGE_HEIGHT - fromTheTop(exported.margin().getFooter()); + + assertThat(foot - (lineOf(line) - share(line))).as("the second line's baseline") + .isCloseTo(exported.baseline("Second"), within(0.1)); + assertThat(foot - (lineOf(line) - share(line)) - lineOf(line)).as("the first line's baseline") + .isCloseTo(exported.baseline("First"), within(0.1)); + } + + @Test + void aFootersRowBesideAPartOfTwoLinesStandsItsFirstLineOnThePagesBaseline() throws Exception { + // The part of two lines makes Word's paragraph two lines tall, whichever part is tallest: + // the part before its break stands on the first, where the page sets it. + Exported exported = export(footerRow(p -> p.text("Acme").textStyle(CHROME), + p -> p.text("First\nSecond").textStyle(CHROME))); + XWPFParagraph line = exported.footerLine(); + + assertThat(PAGE_HEIGHT - fromTheTop(exported.margin().getFooter()) - (lineOf(line) - share(line)) + - lineOf(line)).as("the first line's baseline").isCloseTo(exported.baseline("Acme"), within(0.1)); + assertThat(exported.report().bySubject().get("page zone")).extracting(DocxExportReport.Note::detail) + .containsExactly("a footer written as one line of Word's footer; 1 of its 2 parts stands off where " + + "the page sets them"); + } + + @Test + void wherePartsAfterAPartOfTwoLinesStandIsNotMeasured() throws Exception { + // Word sets what follows the break on its second line, whatever the page sets it on. + Exported exported = export(footerRow(p -> p.text("First\nSecond").textStyle(CHROME), + p -> p.text("Acme").textStyle(CHROME))); + + assertThat(exported.report().bySubject().get("page zone")).extracting(DocxExportReport.Note::detail) + .containsExactly("a footer written as one line of Word's footer; 1 of its 2 parts stands off where " + + "the page sets them; where 1 of its 2 parts stands is not measured"); + } + + @Test + void aHeaderOfLinesReachingPastTheMarginHoldsTheBodyAtIt() throws Exception { + // Three 18pt lines of a zone the body runs under reach some 55pt down a page whose margin + // is 36pt, where one line would not reach the margin. + Exported exported = export(DocumentPageZone.builder().zone(DocumentHeaderFooterZone.HEADER).height(60) + .reserveSpace(false).padding(new DocumentInsets(4, 0, 0, 0)) + .content(page -> new ParagraphBuilder().text("One\nTwo\nThree") + .textStyle(DocumentTextStyle.DEFAULT.withSize(18)).build()) + .build(), 36); + + assertThat(fromTheTop(exported.margin().getHeader()) + lineOf(exported.headerLine())) + .as("one line would stay inside the margin").isLessThan(36); + assertThat(DocxTwips.of(exported.margin().getTop())).as("written negative").isNegative(); + assertThat(exported.report().bySubject().get("page zone")).extracting(DocxExportReport.Note::detail) + .anySatisfy(note -> assertThat(note).startsWith("a header whose line reaches ")); + } + + @Test + void aPictureInAZoneHasRoomAboveTheBaselineAndItsPlaceIsNamed() throws Exception { + // Word stands a zone's picture on the baseline, and an exact line cuts what passes its top: + // the line is tall enough for it, the text still on the page's baseline. + Exported exported = export(DocumentPageZone.builder().zone(DocumentHeaderFooterZone.HEADER).height(60) + .padding(new DocumentInsets(4, 0, 0, 0)) + .content(page -> new ParagraphBuilder() + .inlineImage(com.demcha.compose.document.image.DocumentImageData.fromBytes(png()), 24, 24) + .inlineText(" Acme", CHROME).build()) + .build()); + XWPFParagraph line = exported.headerLine(); + + assertThat(share(line)).as("room above the baseline for the 24pt picture").isGreaterThanOrEqualTo(24); + assertThat(fromTheTop(exported.margin().getHeader()) + share(line)) + .isCloseTo(exported.baseline("Acme"), within(0.1)); + assertThat(exported.report().bySubject().get("page zone")).extracting(DocxExportReport.Note::detail) + .containsExactly("a header written as one line of Word's header; a paragraph's picture stands on " + + "the line's baseline, not where the page sets it"); + } + + @Test + void aLineReachingPastTheMarginHoldsTheBodyAtItAndIsNamed() throws Exception { + // A 30pt picture at the foot of a 36pt header band: the line it needs reaches past the + // margin, where Word would move the body down. + Exported exported = export(DocumentPageZone.builder().zone(DocumentHeaderFooterZone.HEADER).height(36) + .padding(new DocumentInsets(4, 0, 0, 0)) + .content(page -> new ParagraphBuilder() + .inlineImage(com.demcha.compose.document.image.DocumentImageData.fromBytes(png()), 30, 30, + com.demcha.compose.document.node.InlineImageAlignment.BASELINE) + .inlineText(" Acme", CHROME).build()) + .build(), 36); + + assertThat(DocxTwips.of(exported.margin().getTop())).as("written negative").isNegative(); + // The picture stands on the baseline, where the page sets it too: the reach is all there is. + assertThat(exported.report().bySubject().get("page zone")).extracting(DocxExportReport.Note::detail) + .singleElement().satisfies(note -> assertThat(note).startsWith("a header whose line reaches ") + .endsWith("past the page margin, which is written negative so that Word holds the body at the " + + "margin, as the page does; LibreOffice moves the body clear of it")); + } + + @Test + void aPartFittedSmallerIsGivenItsStylesLine() throws Exception { + // Word writes it at its style's 18pt, where the page fits it smaller: in a line the page's + // height, its letters' tops would be cut. + Exported exported = export(DocumentPageZone.builder().zone(DocumentHeaderFooterZone.HEADER).height(40) + .padding(new DocumentInsets(4, 0, 0, 0)) + .content(page -> new ParagraphBuilder() + .text("A running header far too long to set at its eighteen points across this page") + .textStyle(DocumentTextStyle.DEFAULT.withSize(18)).autoSize(18, 6).build()) + .build()); + Exported unfitted = export(zone(DocumentHeaderFooterZone.HEADER, 40, new DocumentInsets(4, 0, 0, 0), "Acme", + DocumentTextStyle.DEFAULT.withSize(18))); + + assertThat(lineOf(exported.headerLine())).as("the 18pt style's line") + .isCloseTo(lineOf(unfitted.headerLine()), within(0.05)); + } + + @Test + void aZoneWhoseFaceDiffersOnItsFirstPageIsNotMeasured() throws Exception { + // Written at 18pt for no page in particular, it is set at 8pt on page 1: a line measured by + // that would cut the written letters' tops. + Exported exported = export(session -> session.chrome().zone(DocumentPageZone.footer(40, + page -> new ParagraphBuilder().text("Acme") + .textStyle(page.isLast() ? DocumentTextStyle.DEFAULT.withSize(18) : CHROME).build())), + true); + CTSpacing spacing = exported.footerLine().getCTP().getPPr().getSpacing(); + + assertThat(spacing.isSetLineRule()).as("the line left Word's").isFalse(); + assertThat(exported.report().bySubject().get("page zone")).extracting(DocxExportReport.Note::detail) + .containsExactly("a footer written as one line of Word's footer; whether its text stands where the " + + "page sets it is not measured"); + } + + @Test + void aZoneThatBuildsNothingForNoPageInParticularIsNamed() throws Exception { + Exported exported = export(session -> session.chrome().zone(DocumentPageZone.footer(40, + page -> page.isPaginated() ? new ParagraphBuilder().text("Drawn").textStyle(CHROME).build() : null)), + false); + + assertThat(exported.document().getFooterList()).as("nothing to write").isEmpty(); + assertThat(exported.report().bySubject().get("page zone")).extracting(DocxExportReport.Note::detail) + .containsExactly("a footer is not written: built for no page in particular, its content is none; " + + "a zone absent from some pages says so through appliesTo"); + } + + @Test + void aZoneThatReadsOtherwiseOnItsFirstPageIsNotMeasured() throws Exception { + // Written for no page in particular, it reads "End"; the page set "Continued" on page 1. + Exported exported = export(session -> session.chrome().zone(DocumentPageZone.footer(40, + page -> new ParagraphBuilder().text(page.isLast() ? "End" : "Continued").textStyle(CHROME).build())), + true); + CTSpacing spacing = exported.footerLine().getCTP().getPPr().getSpacing(); + + assertThat(spacing.isSetLineRule()).as("the line left Word's").isFalse(); + assertThat(exported.report().bySubject().get("page zone")).extracting(DocxExportReport.Note::detail) + .containsExactly("a footer written as one line of Word's footer; whether its text stands where the " + + "page sets it is not measured"); + } + + @Test + void withoutALayoutTheLineIsWords() throws Exception { + byte[] docx; + try (DocumentSession session = GraphCompose.document().pageSize(400, PAGE_HEIGHT) + .margin(DocumentInsets.of(72)).create()) { + session.chrome().zone(zone(DocumentHeaderFooterZone.HEADER, 40, new DocumentInsets(9, 0, 0, 0), + "Chrome", CHROME)); + session.pageFlow(page -> page.addParagraph("Body")); + CapturingBackend captured = new CapturingBackend(); + session.export(captured); + docx = new DocxSemanticBackend().export(captured.graph, + new com.demcha.compose.document.backend.semantic.SemanticExportContext(captured.canvas, + List.of(), null, captured.options)); + } + try (XWPFDocument document = new XWPFDocument(new ByteArrayInputStream(docx))) { + CTSpacing spacing = document.getHeaderList().get(0).getParagraphs().get(0).getCTP().getPPr().getSpacing(); + assertThat(spacing.isSetLineRule()).as("no exact line with nothing to measure it by").isFalse(); + } + } + + @Test + void aFramedZoneLeavesTheDistanceOfAZoneInTheFlow() throws Exception { + // A zone the layout does not measure stays in the flow, at the distance from the edge its + // content gives; a framed zone of its kind, placed after it, leaves that distance alone. + DocumentPageZone unmeasured = DocumentPageZone.builder().zone(DocumentHeaderFooterZone.HEADER).height(40) + .padding(new DocumentInsets(20, 0, 0, 0)) + .content(page -> new ParagraphBuilder().text(page.isLast() ? "End" : "Continued").textStyle(CHROME) + .build()) + .build(); + Exported alone = export(session -> session.chrome().zone(unmeasured), true); + Exported withAFramedOne = export(session -> { + session.chrome().zone(unmeasured); + session.chrome().zone(zone(DocumentHeaderFooterZone.HEADER, 40, new DocumentInsets(4, 0, 0, 0), + "Framed", CHROME)); + }, true); + + assertThat(DocxTwips.of(withAFramedOne.margin().getHeader())) + .isEqualTo(DocxTwips.of(alone.margin().getHeader())); + assertThat(headerParagraph(withAFramedOne.document(), "Framed").getCTP().getPPr().isSetFramePr()).isTrue(); + } + + @Test + void twoZonesAtOneHeightAreTwoFramesNotOne() throws Exception { + // Word takes adjacent paragraphs with the same frame for one frame: a hairline parts them. + Exported exported = export(session -> { + session.chrome().zone(zone(DocumentHeaderFooterZone.HEADER, 40, new DocumentInsets(8, 0, 0, 0), + "Left", CHROME)); + session.chrome().zone(zone(DocumentHeaderFooterZone.HEADER, 40, new DocumentInsets(8, 0, 0, 0), + "Again", CHROME)); + }, false); + List paragraphs = exported.document().getHeaderList().get(0).getParagraphs(); + + assertThat(paragraphs).extracting(XWPFParagraph::getText).containsSubsequence("Left", "", "Again"); + for (String text : List.of("Left", "Again")) { + assertThat(headerParagraph(exported.document(), text).getCTP().getPPr().isSetFramePr()) + .as("%s in a frame", text).isTrue(); + } + assertThat(paragraphs.get(paragraphs.indexOf(headerParagraph(exported.document(), "Left")) + 1) + .getCTP().getPPr().isSetFramePr()).as("the hairline between stands in the flow").isFalse(); + } + + @Test + void aLineOfPartsStandsOnItsTallestPartsBaseline() throws Exception { + Exported exported = export(DocumentPageZone.builder().zone(DocumentHeaderFooterZone.HEADER).height(40) + .padding(new DocumentInsets(4, 0, 0, 0)) + .content(page -> new RowBuilder().name("Line") + .addParagraph(p -> p.text("v2.4").textStyle(CHROME)) + .flexSpacer() + .addParagraph(p -> p.text("Acme").textStyle(DocumentTextStyle.DEFAULT.withSize(18))) + .build()) + .build()); + + assertThat(fromTheTop(exported.margin().getHeader()) + share(exported.headerLine())) + .isCloseTo(exported.baseline("Acme"), within(0.1)); + assertThat(lineOf(exported.headerLine())).as("as tall as the 18pt part's line, not the 8pt one's") + .isBetween(16.0, 18.0); + } + + @Test + void zonesThatShareAKindEachStandInAFrameAtTheirOwnHeight() throws Exception { + // A cover's header on the first page and the running header on the rest: Word holds one + // distance from the edge for both, so each stands in a frame where the page sets it. + Exported exported = export(session -> { + session.chrome().zone(DocumentPageZone.builder().zone(DocumentHeaderFooterZone.HEADER).height(60) + .padding(new DocumentInsets(24, 0, 0, 0)).appliesTo(page -> page.isFirst()) + .content(page -> new ParagraphBuilder().text("Cover").textStyle(CHROME).build()).build()); + session.chrome().zone(DocumentPageZone.builder().zone(DocumentHeaderFooterZone.HEADER).height(60) + .padding(new DocumentInsets(8, 0, 0, 0)).appliesTo(page -> !page.isFirst()) + .content(page -> new ParagraphBuilder().text("Running").textStyle(CHROME).build()).build()); + }, true); + + for (String text : List.of("Cover", "Running")) { + XWPFParagraph line = headerParagraph(exported.document(), text); + CTFramePr frame = line.getCTP().getPPr().getFramePr(); + assertThat(frame).as("%s in a frame", text).isNotNull(); + assertThat(DocxTwips.of(frame.getY()) / 20.0 + share(line)).as("%s's baseline", text) + .isCloseTo(exported.baseline(text), within(0.1)); + } + } + + @Test + void zonesOfOneKindOnTheSamePagesEachStandAtTheirOwnHeightInOnePart() throws Exception { + // Both on every page, in one part, a header's or a footer's: written in the flow, the + // second line stood under the first, whatever the page set it at. + for (DocumentHeaderFooterZone kind : DocumentHeaderFooterZone.values()) { + Exported exported = export(session -> { + session.chrome().zone(zone(kind, 60, new DocumentInsets(30, 0, 0, 0), "Lower", CHROME)); + session.chrome().zone(zone(kind, 60, new DocumentInsets(8, 0, 0, 0), "Upper", CHROME)); + }, false); + + for (String text : List.of("Lower", "Upper")) { + XWPFParagraph line = headerParagraph(exported.document(), text); + CTFramePr frame = line.getCTP().getPPr().getFramePr(); + assertThat(frame).as("%s's %s in a frame", kind, text).isNotNull(); + assertThat(frame.getHRule()).as("at least the line tall, a part of more lines going on below") + .isEqualTo(org.openxmlformats.schemas.wordprocessingml.x2006.main.STHeightRule.AT_LEAST); + assertThat(DocxTwips.of(frame.getH()) / 20.0).as("%s's %s's frame, its line tall", kind, text) + .isCloseTo(lineOf(line), within(0.05)); + assertThat(DocxTwips.of(frame.getY()) / 20.0 + share(line)).as("%s's %s's baseline", kind, text) + .isCloseTo(exported.baseline(text), within(0.1)); + } + } + } + + @Test + void aFramedFooterOfTwoLinesStandsItsFirstOnThePagesBaselineInAFrameOfBoth() throws Exception { + Exported exported = export(session -> { + session.chrome().zone(zone(DocumentHeaderFooterZone.FOOTER, 60, new DocumentInsets(30, 0, 0, 0), + "Lower", CHROME)); + session.chrome().zone(zone(DocumentHeaderFooterZone.FOOTER, 60, new DocumentInsets(4, 0, 0, 0), + "First\nSecond", CHROME)); + }, false); + XWPFParagraph line = headerParagraph(exported.document(), "First\nSecond"); + CTFramePr frame = line.getCTP().getPPr().getFramePr(); + + assertThat(DocxTwips.of(frame.getY()) / 20.0 + share(line)).as("the first line's baseline") + .isCloseTo(exported.baseline("First"), within(0.1)); + assertThat(DocxTwips.of(frame.getH()) / 20.0).as("a frame of both lines") + .isCloseTo(2 * lineOf(line), within(0.05)); + } + + @Test + void aZoneParagraphsAnchorIsNamed() throws Exception { + Exported exported = export(zone(DocumentHeaderFooterZone.FOOTER, 40, DocumentInsets.zero(), + "Chrome", CHROME, "chrome-anchor")); + + assertThat(exported.report().bySubject().get("page zone")).extracting(DocxExportReport.Note::detail) + .containsExactly("a footer written as one line of Word's footer; a paragraph's anchor has no " + + "bookmark in the Word file: a link to it points at none"); + } + + @Test + void aLineThePageSetsAtTheEdgeStandsAtIt() throws Exception { + // An exact line's baseline is four fifths down it; text set against the page's top edge + // stands a little higher than that, and the line stops at the edge. + Exported exported = export(zone(DocumentHeaderFooterZone.HEADER, 40, DocumentInsets.zero(), "Acme", + DocumentTextStyle.DEFAULT.withSize(18))); + + assertThat(fromTheTop(exported.margin().getHeader())).isZero(); + assertThat(share(exported.headerLine()) - exported.baseline("Acme")).as("lower than the page's, by a hair") + .isBetween(0.0, 1.0); + assertThat(exported.report().bySubject()).as("within the place a part keeps").doesNotContainKey("page zone"); + } + + @Test + void aLineStoppedAtTheEdgeFurtherThanAPartKeepsItsPlaceIsNamed() throws Exception { + // At 80pt the line's top would stand 1.8pt past the page's edge: stopped there, the text + // stands that much lower than the page sets it. + Exported exported = export(zone(DocumentHeaderFooterZone.HEADER, 120, DocumentInsets.zero(), "Acme", + DocumentTextStyle.DEFAULT.withSize(80))); + + assertThat(share(exported.headerLine()) - exported.baseline("Acme")).isGreaterThan(1.5); + assertThat(exported.report().bySubject().get("page zone")).extracting(DocxExportReport.Note::detail) + .containsExactly("a header written as one line of Word's header; its text stands off where the page " + + "sets it"); + } + + private static DocumentPageZone zone(DocumentHeaderFooterZone kind, double height, DocumentInsets padding, + String text, DocumentTextStyle style) { + return zone(kind, height, padding, text, style, null); + } + + private static DocumentPageZone zone(DocumentHeaderFooterZone kind, double height, DocumentInsets padding, + String text, DocumentTextStyle style, String anchor) { + return DocumentPageZone.builder().zone(kind).height(height).padding(padding) + .content(page -> { + ParagraphBuilder paragraph = new ParagraphBuilder().text(text).textStyle(style); + if (anchor != null) { + paragraph.anchor(anchor); + } + return paragraph.build(); + }) + .build(); + } + + /** A footer's row of two paragraphs, the second against the right margin. */ + private static DocumentPageZone footerRow(Consumer left, Consumer right) { + return DocumentPageZone.builder().zone(DocumentHeaderFooterZone.FOOTER).height(40) + .padding(new DocumentInsets(4, 0, 0, 0)) + .content(page -> new RowBuilder().name("Line").addParagraph(left).flexSpacer().addParagraph(right) + .build()) + .build(); + } + + /** A header holding the page number alone, 4pt down, its own padding round it. */ + private static DocumentPageZone fieldZone(DocumentInsets padding) { + return fieldZone(DocumentHeaderFooterZone.HEADER, padding, DocumentInsets.zero()); + } + + /** A zone holding the page number alone, 4pt in from its page edge, its own sides round it. */ + private static DocumentPageZone fieldZone(DocumentHeaderFooterZone kind, DocumentInsets padding, + DocumentInsets margin) { + return DocumentPageZone.builder().zone(kind).height(40) + .padding(kind == DocumentHeaderFooterZone.HEADER ? new DocumentInsets(4, 0, 0, 0) + : new DocumentInsets(0, 0, 4, 0)) + .content(page -> new RowBuilder().name("Line") + .add(new com.demcha.compose.document.node.PageFieldNode("", + com.demcha.compose.document.node.PageFieldKind.NUMBER, CHROME, + com.demcha.compose.document.node.TextAlign.LEFT, padding, margin)) + .build()) + .build(); + } + + private static byte[] png() { + java.awt.image.BufferedImage image = new java.awt.image.BufferedImage(8, 8, + java.awt.image.BufferedImage.TYPE_INT_RGB); + try (java.io.ByteArrayOutputStream out = new java.io.ByteArrayOutputStream()) { + javax.imageio.ImageIO.write(image, "png", out); + return out.toByteArray(); + } catch (IOException failure) { + throw new java.io.UncheckedIOException(failure); + } + } + + /** Takes the graph, the canvas and the options a session hands any backend. */ + private static final class CapturingBackend + implements com.demcha.compose.document.backend.semantic.SemanticBackend { + + private com.demcha.compose.document.layout.DocumentGraph graph; + private com.demcha.compose.document.layout.LayoutCanvas canvas; + private com.demcha.compose.document.output.DocumentOutputOptions options; + + @Override + public String name() { + return "capture"; + } + + @Override + public byte[] export(com.demcha.compose.document.layout.DocumentGraph documentGraph, + com.demcha.compose.document.backend.semantic.SemanticExportContext context) { + this.graph = documentGraph; + this.canvas = context.canvas(); + this.options = context.outputOptions(); + return new byte[0]; + } + } + + /** + * How far down its exact line an exact line's baseline stands, in points: four fifths, as + * measured in Word and LibreOffice (DocxTextBands.BASELINE_SHARE). + */ + private static double share(XWPFParagraph paragraph) { + return 0.8 * lineOf(paragraph); + } + + private static double lineOf(XWPFParagraph paragraph) { + CTSpacing spacing = paragraph.getCTP().getPPr().getSpacing(); + assertThat(spacing.getLineRule()).as("the zone's line is exact").isEqualTo(STLineSpacingRule.EXACT); + return DocxTwips.of(spacing.getLine()) / 20.0; + } + + private static double fromTheTop(Object twips) { + return DocxTwips.of(twips) / 20.0; + } + + /** A header's or a footer's paragraph reading the text. */ + private static XWPFParagraph headerParagraph(XWPFDocument document, String text) { + List parts = new ArrayList<>(document.getHeaderList()); + parts.addAll(document.getFooterList()); + for (XWPFHeaderFooter part : parts) { + for (XWPFParagraph paragraph : part.getParagraphs()) { + if (paragraph.getText().equals(text)) { + return paragraph; + } + } + } + throw new AssertionError("no header or footer paragraph reading " + text); + } + + private record Exported(XWPFDocument document, DocxExportReport report, List text) { + + CTPageMar margin() { + return document.getDocument().getBody().getSectPr().getPgMar(); + } + + XWPFParagraph headerLine() { + return document.getHeaderList().get(0).getParagraphs().get(0); + } + + XWPFParagraph footerLine() { + return document.getFooterList().get(0).getParagraphs().get(0); + } + + /** Where the page sets a word's baseline, from its top, on the first page it draws it. */ + double baseline(String word) { + StringBuilder letters = new StringBuilder(); + for (int start = 0; start < text.size(); start++) { + letters.setLength(0); + for (int index = start; index < text.size() && letters.length() < word.length(); index++) { + letters.append(text.get(index).getUnicode()); + } + if (letters.toString().equals(word)) { + return text.get(start).getYDirAdj(); + } + } + throw new AssertionError("the page draws no " + word); + } + } + + private static Exported export(DocumentPageZone zone) throws Exception { + return export(session -> session.chrome().zone(zone), false); + } + + private static Exported export(DocumentPageZone zone, double margin) throws Exception { + return export(session -> session.chrome().zone(zone), false, margin); + } + + private static Exported export(Consumer chrome, boolean twoPages) throws Exception { + return export(chrome, twoPages, 72); + } + + private static Exported export(Consumer chrome, boolean twoPages, double margin) throws Exception { + AtomicReference report = new AtomicReference<>(); + try (DocumentSession session = GraphCompose.document() + .pageSize(400, PAGE_HEIGHT) + .margin(DocumentInsets.of(margin)) + .create()) { + chrome.accept(session); + session.pageFlow(page -> { + page.addParagraph(p -> p.text("Body")); + if (twoPages) { + page.addPageBreak(pageBreak -> { }); + page.addParagraph(p -> p.text("More")); + } + }); + List text = textOf(session.toPdfBytes()); + byte[] docx = session.export(new DocxSemanticBackend(report::set)); + return new Exported(new XWPFDocument(new ByteArrayInputStream(docx)), report.get(), text); + } + } + + /** Every letter the page draws, in the order it draws them, page by page. */ + private static List textOf(byte[] pdf) throws IOException { + List letters = new ArrayList<>(); + try (PDDocument document = Loader.loadPDF(pdf)) { + PDFTextStripper stripper = new PDFTextStripper() { + @Override + protected void processTextPosition(TextPosition text) { + letters.add(text); + } + }; + stripper.getText(document); + } + return letters; + } +} diff --git a/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxZonePartsTest.java b/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxZonePartsTest.java new file mode 100644 index 000000000..e0d9a02db --- /dev/null +++ b/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxZonePartsTest.java @@ -0,0 +1,119 @@ +package com.demcha.compose.document.backend.semantic.docx; + +import com.demcha.compose.document.dsl.ParagraphBuilder; +import com.demcha.compose.document.dsl.RowBuilder; +import com.demcha.compose.document.image.DocumentImageData; +import com.demcha.compose.document.node.DocumentNode; +import com.demcha.compose.document.node.InlineImageAlignment; +import com.demcha.compose.document.node.InlineRun; +import com.demcha.compose.document.node.PageFieldKind; +import com.demcha.compose.document.node.PageFieldNode; +import com.demcha.compose.document.node.ParagraphNode; +import com.demcha.compose.document.style.DocumentColor; +import com.demcha.compose.document.style.DocumentTextStyle; +import org.junit.jupiter.api.Test; + +import java.io.ByteArrayOutputStream; +import java.io.IOException; +import java.io.UncheckedIOException; +import java.util.function.Function; + +import static org.assertj.core.api.Assertions.assertThat; + +/** + * A page zone's parts: whether content built for a page reads as the content written, as far as + * the zone's line is set by it, and what of a paragraph's pictures the line holds. + */ +class DocxZonePartsTest { + + private static final DocumentImageData PICTURE = DocumentImageData.fromBytes(png()); + + @Test + void contentBuiltAfreshReadsAlikeWhateverItsColourIsMadeOf() { + // A colour built afresh is not equal to itself: it sets nothing of the line either. + Function zone = text -> new RowBuilder().name("Line") + .addParagraph(p -> p.text(text).textStyle(style(8))) + .flexSpacer() + .add(new PageFieldNode(PageFieldKind.NUMBER, style(8))) + .build(); + + assertThat(DocxZoneParts.readAlike(zone.apply("Acme"), zone.apply("Acme"))).isTrue(); + assertThat(DocxZoneParts.readAlike(zone.apply("Acme"), zone.apply("Other"))).as("other text").isFalse(); + } + + @Test + void anotherFaceSizeOrFieldDoesNotReadAlike() { + assertThat(DocxZoneParts.readAlike(paragraph("Acme", style(18)), paragraph("Acme", style(8)))) + .as("another size").isFalse(); + assertThat(DocxZoneParts.readAlike(paragraph("Acme", style(8)), + paragraph("Acme", DocumentTextStyle.builder().size(8).decoration( + com.demcha.compose.document.style.DocumentTextDecoration.BOLD).build()))) + .as("another weight").isFalse(); + assertThat(DocxZoneParts.readAlike(new PageFieldNode(PageFieldKind.NUMBER, style(8)), + new PageFieldNode(PageFieldKind.TOTAL, style(8)))).as("another field").isFalse(); + assertThat(DocxZoneParts.readAlike(paragraph("Acme", style(8)), null)).as("nothing built").isFalse(); + assertThat(DocxZoneParts.readAlike(paragraph("Acme", style(8)), + new RowBuilder().name("Line").addParagraph(p -> p.text("Acme").textStyle(style(8))) + .flexSpacer().build())).as("another count of parts").isFalse(); + } + + @Test + void aRunsFaceOrAPicturesSizeOrPlaceMakesItReadOtherwise() { + assertThat(DocxZoneParts.readAlike(withPicture(24, InlineImageAlignment.BASELINE, style(8)), + withPicture(24, InlineImageAlignment.BASELINE, style(8)))).isTrue(); + assertThat(DocxZoneParts.readAlike(withPicture(24, InlineImageAlignment.BASELINE, style(8)), + withPicture(12, InlineImageAlignment.BASELINE, style(8)))).as("another size").isFalse(); + assertThat(DocxZoneParts.readAlike(withPicture(24, InlineImageAlignment.BASELINE, style(8)), + withPicture(24, InlineImageAlignment.CENTER, style(8)))).as("another place").isFalse(); + assertThat(DocxZoneParts.readAlike(withPicture(24, InlineImageAlignment.BASELINE, style(8)), + withPicture(24, InlineImageAlignment.BASELINE, style(18)))).as("a run's other face").isFalse(); + } + + @Test + void theTallestPictureIsTheLinesAndOnlyAPictureOffTheBaselineStandsOffIt() { + ParagraphNode paragraph = (ParagraphNode) new ParagraphBuilder() + .inlineImage(PICTURE, 12, 12) + .inlineImage(PICTURE, 30, 30, InlineImageAlignment.BASELINE) + .inlineText(" Acme", style(8)).build(); + + assertThat(DocxZoneParts.tallestPicture(paragraph)).isEqualTo(30); + assertThat(DocxZoneParts.tallestPicture(paragraph("Acme", style(8)))).as("no picture").isZero(); + InlineRun onTheBaseline = runOf(withPicture(24, InlineImageAlignment.BASELINE, style(8))); + InlineRun centred = runOf(withPicture(24, InlineImageAlignment.CENTER, style(8))); + InlineRun raised = runOf((ParagraphNode) new ParagraphBuilder() + .inlineImage(PICTURE, 24, 24, InlineImageAlignment.BASELINE, 2, null).build()); + assertThat(DocxZoneParts.setOffTheBaseline(onTheBaseline)).isFalse(); + assertThat(DocxZoneParts.setOffTheBaseline(centred)).isTrue(); + assertThat(DocxZoneParts.setOffTheBaseline(raised)).as("raised off it").isTrue(); + assertThat(DocxZoneParts.setOffTheBaseline(new com.demcha.compose.document.node.InlineTextRun("Acme"))) + .as("letters").isFalse(); + } + + private static DocumentTextStyle style(double size) { + return DocumentTextStyle.builder().size(size).color(DocumentColor.rgb(40, 60, 90)).build(); + } + + private static ParagraphNode paragraph(String text, DocumentTextStyle style) { + return (ParagraphNode) new ParagraphBuilder().text(text).textStyle(style).build(); + } + + private static ParagraphNode withPicture(double size, InlineImageAlignment alignment, DocumentTextStyle style) { + return (ParagraphNode) new ParagraphBuilder().inlineImage(PICTURE, size, size, alignment) + .inlineText(" Acme", style).build(); + } + + private static InlineRun runOf(ParagraphNode paragraph) { + return paragraph.inlineRuns().get(0); + } + + private static byte[] png() { + java.awt.image.BufferedImage image = new java.awt.image.BufferedImage(8, 8, + java.awt.image.BufferedImage.TYPE_INT_RGB); + try (ByteArrayOutputStream out = new ByteArrayOutputStream()) { + javax.imageio.ImageIO.write(image, "png", out); + return out.toByteArray(); + } catch (IOException failure) { + throw new UncheckedIOException(failure); + } + } +} diff --git a/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxZoneReportTest.java b/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxZoneReportTest.java index 01b89987b..1b567b732 100644 --- a/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxZoneReportTest.java +++ b/render-docx/src/test/java/com/demcha/compose/document/backend/semantic/docx/DocxZoneReportTest.java @@ -159,14 +159,14 @@ void aPartOnABaselineOfItsOwnOrOfMoreLinesThanOneIsCounted() throws Exception { .build()))) .as("in a header too").containsExactly("a header written as one line of Word's header; 2 of its 3 " + "parts stand off where the page sets them"); - // Word stands the line at the footer's foot, which a smaller part set lower holds: the - // larger part moves down onto it, and the smaller one onto the larger one's baseline. + // Word's line stands on its tallest part's baseline, where the page sets that part: a + // smaller part the page sets lower moves up onto it, and the larger one stays. assertThat(zoneNotes(DocumentPageZone.footer(60, page -> new RowBuilder().name("Line") .addParagraph(p -> p.text("Acme").textStyle(DocumentTextStyle.DEFAULT.withSize(18))) .flexSpacer() .addParagraph(p -> p.text("v2.4").textStyle(CHROME).margin(new DocumentInsets(30, 0, 0, 0))) .build()))) - .containsExactly(FOOTER + "2 of its 2 parts stand off where the page sets them"); + .containsExactly(FOOTER + "1 of its 2 parts stands off where the page sets them"); // A part the page seats off its baseline is set on Word's. assertThat(zoneNotes(DocumentPageZone.footer(40, page -> new RowBuilder().name("Line") .addParagraph(p -> p.text("Confidential").textStyle(DocumentTextStyle.DEFAULT.withSize(18)))