Skip to content

fix(docx): name in the report a paragraph or list item the page reads as markdown - #864

Merged
DemchaAV merged 2 commits into
2.5-devfrom
fix/docx-report-session-markdown
Oct 6, 2026
Merged

DemchaAV merged 2 commits into
2.5-devfrom
fix/docx-report-session-markdown

Conversation

@DemchaAV

@DemchaAV DemchaAV commented Oct 6, 2026 •

Copy link
Copy Markdown
Owner

Why

A session reads markdown unless it is told not to: GraphCompose.document() defaults to markdown(true).

  • What the page does. It reads a paragraph, or a list item, of plain text holding a mark of emphasis or code (*, _, `) as markdown, through flexmark (MarkDownParser):
    • it sets the text the emphasis marks style bold or italic, and a heading bold, the first three levels larger;
    • it drops the marks its parser reads as syntax, a code span's backticks and a link's address among them.
  • What the file did. The DOCX export writes the text as authored, so Word showed **bold** with its asterisks and none of the bold. The report named none of it, and the markdown flag never reaches a semantic backend.
  • The corpus has one such paragraph. TimelineMinimal's open-source project line ends in *(Open source)*: the page sets it in italic, and Word showed the asterisks.

What changed

  • markdownLost names a paragraph the page read as markdown. It is added to the paragraph's note on every body path (paragraphLost: the body, a text box, a side of a pair, a badge) and to a zone paragraph's, on its page zone note (zoneParagraphLost). The phrase: "its markdown marks are written as letters, where the page sets the text they mark and drops them".
  • itemsMarkdownLost names a list's items, on the list's note: "its items' markdown marks are written as letters, …".
  • How it is read (marksDropped, shared by both).
    • The page's own trigger comes first (holdsAMarkdownTrigger, readsAsMarkdown): plain text, no runs, holding *, _ or `. Only then are the laid-out lines read.
    • The lines hold fewer of the characters markdown syntax is made of (*_`\#[]()>~) than the text, as the page lays it out, where the page read it. Both sides count the same characters, so one the page keeps counts on both.
    • The text is counted as the page lays it out:
      • a paragraph's prefix before the first line has its marks left out;
      • a list item has a marker typed before it taken off (ListMarker.normalizeItemText), as both the page and the file take it off;
      • a list's lines are its items' (itemLines), without a marker laid out on its own.
    • Lines that hold no text, as a paragraph given no width lays out, are not counted.
    • So text the page sets as authored is not named:
      • markdown off;
      • an underscore inside a word, which markdown keeps;
      • a marker typed before a list item;
      • runs, which the page never reads.
  • Where the lines are not read, it says so — if the page's parser would drop a mark (parserDropsAMark runs MarkDownParser on the text). That happens with no layout, or composed in a table cell: a paragraph there is matched to its lines by its text, and a list is not matched at all. The phrase then ends "— whether the page reads them is not measured". A cell holding node_js is not named, since the parser keeps that underscore.
  • The file is not changed. It still writes the text as authored, and the docs say so.
  • Ledger. The entries that now name the loss:
    • a paragraph's text moves from WRITTEN to REPORTED;
    • a list's nestedItems moves from WRITTEN to REPORTED;
    • a list's items, already REPORTED, names it too.
  • Docs. These say what is named, that markdown is the default, and that the file does not yet write the text as the page sets it:
    • the recipe's Paragraphs row and its list-report paragraph;
    • the capability matrix's paragraph and list rows;
    • render-docx/README.md;
    • the CHANGELOG.

Verification

  • ./mvnw -B -ntp install -pl :graph-compose-render-docx → BUILD SUCCESS: 1096 tests, 0 failures, 1 skipped (the property-gated fidelity probe).
  • DocxMarkdownReportTest is new, with 5 tests.
    • Named:
      • a paragraph read as markdown;
      • a heading;
      • a code span alone;
      • one behind a prefix with a mark of its own, and with as many marks as the text drops;
      • a list whose item is read: flat, with markers in a column of their own, and in a tree of items;
      • a zone paragraph, on the zone's note;
      • not measured:
        • a paragraph with no layout;
        • a paragraph composed in a table cell;
        • a list composed in a table cell.
    • Not named:
      • markdown off, for a paragraph, a list and a zone;
      • no mark, with a layout and without;
      • an underscore inside a word, in a paragraph and in a table cell's paragraph and list;
      • runs;
      • a prefix's mark with markdown off;
      • a paragraph given no width;
      • a marker typed before a list item, with markdown on and off.
  • The tests fail without the code they cover. 11 sabotages were run, each breaking one thing, and each made the tests that cover it fail:
    • the body, the zone and the list phrase;
    • the comparison, never and at equal counts;
    • the not-measured phrase, never and always;
    • the prefix's marks;
    • a typed marker kept;
    • a tree of items left unread;
    • lines with no text counted.
  • The DOCX bytes do not change. The 62 corpus documents exported deterministically are byte-identical to the export before the change: DocxFidelityCorpusTest -Dgraphcompose.docxFidelity=export, SHA-256 per file, 0 of 62 differ.
  • Across the corpus the report names one paragraph more, TimelineMinimal's project line: 951 notes, from 950.
  • Documentation guards:
    • core: -pl :graph-compose-core -Dtest='com.demcha.documentation.**' → 166 tests, 0 failures;
    • qa: the documentation guards plus DocxPageZoneTest, DocxTransparentWrapperTest, TimelineRailAcrossBackendsTest and RtlAcrossBackendsTest → 50 tests, 0 failures.
  • The full reactor gate was not run; no public signature, POM or workflow file changed.

Known limits

  • A list marker holding marks can hide the loss. A legacy-layout list sets its marker in each item's line, so ListMarker.custom("*") or "(a)" adds marks to the lines.
  • Composed in a table cell, a paragraph or a list is "not measured" wherever the parser would drop a mark, markdown on or off, since its lines are not read there.
  • Some of what markdown changes is not counted. A 1. line read as a list item, a --- rule and an entity like & hold none of the counted characters.
  • The fix — writing the runs the page sets — is not made here. The export still writes the text as authored.

Lane: render-docx backend (report only, no change to what is written) plus tests and docs.

… as markdown

A session reads markdown unless told not to: the page sets the text a
paragraph's or a plain list item's emphasis marks style and drops the
marks, while the DOCX export writes the text as authored, marks and all.
The paragraph's, the zone's and the list's notes now name it, read from
the laid-out lines, which hold fewer marks than the text where the page
read it, and say it is not measured where the lines are not read. The
written bytes do not change.
…ame unread lines only where the parser drops a mark

A marker typed before a list item is taken off by both the page and the
file, so it is no markdown mark lost: items are counted as the page lays
them out, against their own lines without a marker laid out apart. Where
no lines are read, the note says markdown is not measured only where the
page's parser would drop a mark, so a cell holding node_js is not named.
Lines holding no text are not counted, and the docs say the file does
not yet write the text as the page sets it.
@DemchaAV
DemchaAV merged commit 74a192a into 2.5-dev Oct 6, 2026
13 checks passed
@DemchaAV
DemchaAV deleted the fix/docx-report-session-markdown branch October 6, 2026 20:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant