Conversation
The "ill-formed: wrong fourth byte" SECTIONs in unit-unicode3.cpp, unit-unicode4.cpp, and unit-unicode5.cpp guarded their loop with a check on byte3 instead of byte4. Since the enclosing loop already restricts byte3 to its valid range, the guard was always true and the section's "continue" fired unconditionally, so check_utf8string()/check_utf8dump() were never actually invoked for a malformed fourth byte. Fixing the guard naively (byte3 -> byte4) would also have swept the full byte2 x byte3 combinatorics for every byte4 value, adding millions of redundant iterations: the lexer validates continuation bytes strictly in sequence with early exit (see next_byte_in_range() in lexer.hpp), so once byte2/byte3 are within their valid range, the byte4 outcome does not depend on which valid byte2/byte3 values were chosen. Instead, byte2 and byte3 are now held to a small hedge of representative valid prefixes (range corners plus a midpoint) while byte4 is still swept exhaustively over its full 0x00-0xFF range, since that is the actual property under test. Also fixed the garbled "skip fourth second byte" comment in unit-unicode3.cpp. Verified offline: before the fix, the "wrong fourth byte" subcase executes 0 assertions in all three files (proving it was dead code); after the fix, it executes 11520 (unicode3), 34560 (unicode4), and 11520 (unicode5) assertions, and a deliberately reintroduced bug in the lexer's byte4 range check causes it to fail (proving it is now meaningful). Total per-file assertion counts grow by the same small amounts, not by millions, and all other sections in these files still pass unchanged. Fixes #5416 Signed-off-by: Niels Lohmann <mail@nlohmann.me>
5 tasks
The maintainer wants exhaustive coverage of every byte combination here rather than the representative-prefix reduction, matching the style of the sibling "wrong second/third byte" sections in the same files. Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fixes #5416.
The
SECTION("ill-formed: wrong fourth byte")blocks intests/src/unit-unicode3.cpp,tests/src/unit-unicode4.cpp, andtests/src/unit-unicode5.cppguarded theirbyte4sweep with a check onbyte3instead ofbyte4:Because the enclosing loop already restricts
byte3to0x80..0xBF, this guard was always true, socontinuefired unconditionally andcheck_utf8string()/check_utf8dump()were never actually called — a malformed fourth byte in a 4-byte UTF-8 sequence has never been tested.unit-unicode3.cppalso had a garbled comment ("skip fourth second byte"), fixed here as well.Fix
Fixed the guard to check
byte4(and fixed the garbled comment inunit-unicode3.cpp), and kept the fullbyte2 x byte3combinatorics — matching the style of the sibling "wrong second/third byte" sections in the same files, which also sweep every combination of the bytes they don't directly target. This is a deliberate choice: a smaller, representative-prefix version of this fix was considered (and briefly implemented) on the reasoning that the lexer'snext_byte_in_range()validates continuation bytes strictly in sequence with early-exit, sobyte4's accept/reject outcome doesn't depend on which specific validbyte2/byte3values were used to reach it — but the maintainer prefers exhaustive coverage of every byte combination here for a stronger coverage claim, consistent with how the rest of the file already tests this property, so the full sweep is kept.The companion issue #5418 (further runtime-reduction ideas: the wrong-2nd/3rd-byte sections' combinatorics, the binary-format 4x-parse issue, 16-bit integer sweeps, etc.) is intentionally left out of scope for a separate PR/decision.
Verification (done offline, without CMake/network)
For each file, compiled and ran with:
Proved the section was dead code: on the pre-fix code, running just the "wrong fourth byte" subcase (via doctest's
--subcasefilter) executes 0 assertions in all three files.Proved the fix is meaningful: deliberately reintroducing a bug in the lexer's
byte4range check for the0xF0case (temporarily widening it to accept any byte) causes the fixedunit-unicode3.cpptest to fail — confirming the test now actually catches a byte4 validation regression. The bug was reverted immediately after confirming the failure.Full-suite assertion counts, before -> after (with the full combinatorial sweep restored):
(unicode4's +28.3M matches the issue's own estimate of ~2.36M added iterations for a full-sweep fix.)
No regressions: all other sections in all three files still pass unchanged, and all three files run clean (0 failed) with
--no-skip.Breaking change?
No. This is a test-only change (no
include/changes) and does not affect the public API in any way.— opened by Claude Code on behalf of @nlohmann