Skip to content

0.26.0: HWP → HTML on our own reader (replaces pyhwp) + HWP reader fixes - #10

Merged
CocoRoF merged 1 commit into
mainfrom
feat/hwp-html
Sep 23, 2026
Merged

CocoRoF merged 1 commit into
mainfrom
feat/hwp-html

Conversation

@CocoRoF

@CocoRoF CocoRoF commented Sep 23, 2026

Copy link
Copy Markdown
Contributor

Why

XGEN's document source viewer rendered .hwp through pyhwp's hwp5html (AGPL-3.0). XGEN keeps no AGPL dependency, so this adds an HTML renderer on our own HWP reader. Comparing it against hwp5html on real documents also exposed reader bugs that already affected workflow's HWP preview (DOCX path).

New: documents/legacy/hwp_html.hwp_to_html(content) -> str

One self-contained HTML page built directly from the reader's model (no DOCX detour, no subprocess):

  • fonts, sizes, colors, alignment, margins, hanging indents, line spacing
  • merged cells, column widths, cell and table border/fill
  • pictures, headers/footers, master pages, footnotes
  • inline objects kept at their position in the line

Safe to inject unsanitized (the frontend does):

  • text and attributes are escaped
  • font names are stripped to [\w .-]
  • every CSS rule is scoped under .hwp-doc
  • only png/jpeg/gif/bmp are embedded (bmp/tif are converted to png), with a 64 MB image cap

Reader fixes (shared with hwp_to_docx)

  • Pictures dropped: BinData compression is per item (spec table 18). png/jpg stored raw were inflated, failed, and were silently skipped (11 of 15 pictures in one real press release).
  • Wrong strikethrough: per table 35, strikethrough is bits 18-20. Underline kind 2 was read as strike.
  • Misplaced styles: char-shape positions are raw word positions (controls count as 8 words), so styles after section or field controls were misplaced.
  • Missing text:
    • hyphen, 묶음 빈칸 and 고정폭 빈칸 are now text
    • auto numbers are rendered (표 1)
    • 덧말 and 글자 겹침 text is kept
    • footnotes and endnotes are kept
    • inline (treat-as-char) objects are placed in the line
    • nested paragraph lists and text-box attachments are kept
  • Layout:
    • paragraph margins and indent
    • distribute alignment
    • Korean line breaking by 어절
    • landscape pages
    • table border/fill and caption side
    • per-section headers
  • Distribution documents (배포용) open per Hancom's published spec, and copy/print protection carries over. DRM and certificate-encrypted files are refused by name. Decompression is capped at 256 MB.

Verification

  • 31 HWP files: 8 format fixtures and 23 real documents (government reports, a 10 MB handbook, exams, the HWP spec itself).
    • hwp5html fails on 5 of them; ours fails on none.
    • Every file converts on both the HTML and DOCX paths.
    • Coverage of doc2chunk's chunk text (what citation highlighting searches) is ≥ hwp5html on every file.
    • Image and table counts are ≥ hwp5html.
  • Tests:
    • new tests/unit/test_hwp_html.py (25 tests, fixtures built per spec)
    • the rich fixture's strike encoding is corrected to table 35
    • full suite: 957 passed, 1 skipped

본 제품은 한글과컴퓨터의 글 문서 파일(.hwp) 공개 문서를 참고하여 개발하였습니다.

🤖 Generated with Claude Code

…ader fixes

New documents/legacy/hwp_html.hwp_to_html: one self-contained HTML page
straight from HwpFile's model — no DOCX detour, no subprocess. Fonts/sizes/
colors, alignment/margins/hanging indents/line spacing, merges/column
widths/cell and table border-fill, pictures, headers/footers, master pages,
footnotes, inline objects in place. Safe to inject unsanitized: text and
attributes escaped, font names stripped to [\w .-], every CSS rule under
.hwp-doc, only png/jpeg/gif/bmp embedded (bmp/tif → png), 64 MB image cap.

Reader (shared with hwp_to_docx, so workflow previews get these too):
- HwpFile loader (header checks, DocInfo, sections, BinData) split out.
- Char-shape positions are raw word positions (controls count 8): mapped,
  so styles after section/field controls land where the file says.
- Hyphen / 묶음 빈칸 / 고정폭 빈칸 are text; surrogate pairs kept whole.
- Extended controls paired with their CTRL_HEADER by ctrl id: auto numbers
  (표 142, number shapes 표 134), 덧말/글자 겹침 text in place, tbl/gso get
  inline (treat-as-char, 표 70 bit 0) + position, footnotes/endnotes kept.
- BinData compression per item (표 18 bits 4-5): png/jpg stored raw were
  being inflated, failing and silently dropped (11 of 15 pictures in a
  real press release).
- Char shape per 표 35: strike is bits 18-20, underline 1/overline 3,
  super/subscript — underline kind 2 was read as strike.
- Para shape: left/right margin + indent (표 43), distribute align,
  Korean line break by 어절/글자 (표 44 bit 7). Landscape pages (표 132).
- Table border-fill (표 74), caption side (표 73), nested paragraph lists
  hanging off a paragraph, text-box attachments, per-section headers.
- Distribution documents (배포용) opened per Hancom's published spec
  (seed → MSVC rand → XOR → AES-128 ECB); copy/print protection kept.
  DRM / certificate-encrypted files refused by name. 256 MB decompression cap.

Checked on 31 HWP files (8 format fixtures + 23 real documents incl.
government reports, a 10 MB handbook, exams, the HWP spec itself):
hwp5html fails on 5 of them, ours on none; text coverage of doc2chunk's
chunk text in the HTML is ≥ hwp5html on every file; images/tables ≥.

본 제품은 한글과컴퓨터의 글 문서 파일(.hwp) 공개 문서를 참고하여 개발하였습니다.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@CocoRoF
CocoRoF merged commit 06dbb9c into main Sep 23, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant